Biological information sequence generation method and device, equipment, storage medium and product
By acquiring biological information fragment sequences and descriptive tags, encoding them using coding networks and splicing them, the problem that amino acid sequence generation in the prior art cannot control structure or function is solved, and the efficient generation of biological information sequences with designated functions or structures is achieved.
Patent Information
- Application Number
- CN202410012058.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
The method of generating amino acid sequences in the prior art cannot effectively control the structure or function of the amino acid sequence, and it is difficult to meet the specific needs of users.
By acquiring the biological information fragment sequence and the description tag, the first and second encoding networks are used to encode them, and then splicing them and decoding them through the sequence generation network to generate a biological information sequence with a specified function or structure.
It improves the efficiency of the generation of biological information sequences, ensures that the generated sequence has the expected functions or structures and meets user needs.
Smart Images

Figure CN120260666A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of bioinformatics technology, and particularly to a method, device, equipment, storage medium and product for generating bioinformatics sequences. Background Art
[0002] A bioinformatics sequence refers to a sequence with bioinformatics. Taking the amino acid sequence as an example of the bioinformatics sequence, one or more amino acid sequences constitute a protein.
[0003] Taking the generation of amino acid sequences as an example, in the related art, amino acid fragment sequences are encoded; the sequence generation network decodes the encoded amino acid fragment sequences to generate amino acid sequences.
[0004] However, the construction of amino acid sequences involves not only amino acid fragment sequences but also at least one of the functions, structures, and species descriptions of amino acid sequences. After encoding the amino acid fragment sequences by the method in the related art and generating amino acid sequences, the structure or function of the obtained amino acid sequences is uncontrollable, and the generated amino acid sequences are difficult to meet the needs of users. Summary of the Invention
[0005] The present application provides a method, device, equipment, storage medium and product for generating bioinformatics sequences, and the technical solutions are as follows:
[0006] According to one aspect of the present application, a method for generating a bioinformatics sequence is provided, and the method includes:
[0007] Obtain a bioinformatics fragment sequence and a description tag, where the bioinformatics fragment sequence includes bioinformatics fragments required for constructing the bioinformatics sequence, and the description tag is based on natural language to indicate the inherent properties of the bioinformatics sequence;
[0008] Perform a first encoding operation on the bioinformatics fragment sequence through a first encoding network to obtain a bioinformatics fragment sequence identifier; perform a second encoding operation on the description tag through the second encoding network to obtain a description tag identifier;
[0009] Concatenate the bioinformatics fragment sequence identifier and the description tag identifier to obtain a bioinformatics identifier sequence;
[0010] Decode the bioinformatics identifier sequence through a sequence generation network to obtain the bioinformatics sequence having at least one of the functions and the structures.
[0011] According to one aspect of the present application, a device for generating a bioinformatics sequence is provided, and the device includes:
[0012] An acquisition module, configured to acquire a biological information fragment sequence and a description tag, where the biological information fragment sequence includes biological information fragments required for constructing the biological information sequence, and the description tag indicates the inherent properties of the biological information sequence based on natural language;
[0013] An encoding module, configured to perform a first encoding operation on the biological information fragment sequence through a first encoding network to obtain a biological information fragment sequence identifier; perform a second encoding operation on the description tag through the second encoding network to obtain a description tag identifier; splice the biological information fragment sequence identifier and the description tag identifier to obtain a biological information identifier sequence;
[0014] A decoding module, configured to decode the biological information identifier sequence through a sequence generation network to obtain the biological information sequence having at least one of the functions and the structures.
[0015] According to another aspect of the present application, there is provided a computer device, which includes: a processor and a memory, and at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method for generating a biological information sequence as described in the above aspect.
[0016] According to another aspect of the present application, there is provided a computer storage medium, and at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by the processor to implement the method for generating a biological information sequence as described in the above aspect.
[0017] According to another aspect of the present application, there is provided a computer program product, the computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium; the computer program is read and executed by a processor of a computer device, so that the computer device executes the method for generating a biological information sequence as described in the above aspect.
[0018] The beneficial effects brought by the technical solution provided by the present application at least include:
[0019] By acquiring a biological information fragment sequence and a description tag; performing respective corresponding encoding operations on the biological information fragment sequence and the description tag through an encoding network, and splicing the obtained biological information fragment sequence identifier and description tag identifier to obtain a biological information identifier sequence; decoding the biological information identifier sequence through a sequence generation network to obtain a biological information sequence. By adding a description tag and performing respective corresponding encoding operations on the biological information fragment sequence and the description tag, the present application generates a biological information sequence with specified functions or structures, thereby improving the generation efficiency of the biological information sequence. Description of the Drawings
[0020] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a schematic diagram of a method for generating a biological information sequence provided by an exemplary embodiment of the present application;
[0022] Figure 2 It is a schematic diagram of the architecture of a computer system provided by an exemplary embodiment of the present application;
[0023] Figure 3 It is a flowchart of a method for generating a biological information sequence provided by an exemplary embodiment of the present application;
[0024] Figure 4 It is a flowchart of a method for generating a biological information sequence provided by an exemplary embodiment of the present application;
[0025] Figure 5 It is a schematic diagram of encoding a network to encode a biological information fragment sequence and a description tag provided by an exemplary embodiment of the present application;
[0026] Figure 6 It is a schematic diagram of encoding a network to encode a biological information fragment sequence and a description tag provided by an exemplary embodiment of the present application;
[0027] Figure 7 It is a schematic diagram of the structural comparison between a natural enzyme and a synthetic enzyme provided by an exemplary embodiment of the present application;
[0028] Figure 8 It is a schematic diagram of a method for generating a biological information sequence provided by an exemplary embodiment of the present application;
[0029] Figure 9 It is a framework diagram of generating a training system for a biological information sequence generation model and training the biological information sequence generation model provided by an exemplary embodiment of the present application;
[0030] Figure 10 It is a flowchart of a method for training a biological information sequence generation model provided by an exemplary embodiment of the present application;
[0031] Figure 11 It is a schematic diagram of training an encoding network provided by an exemplary embodiment of the present application;
[0032] Figure 12It is a block diagram of a device for generating a biological information sequence provided by an exemplary embodiment of the present application;
[0033] Figure 13 It is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0034] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings. Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0035] The terms used in this disclosure are only for the purpose of describing specific embodiments and are not intended to limit the disclosure. The singular forms "a", "the", and "said" used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0036] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other.
[0037] For ease of understanding, the following explains several terms related to the present application.
[0038] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0039] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0040] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or a local area network to achieve the calculation, storage, processing, and sharing of data.
[0041] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of the technical network system require a large amount of computing and storage resources, such as video websites, picture-based websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.
[0042] Cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage.
[0043] As a basic capability provider of cloud computing, a cloud computing resource pool will be established, abbreviated as the cloud platform, generally called the IaaS (Infrastructure as a Service) platform. Various types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (virtual machines containing operating systems), storage devices, and network devices.
[0044] According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. Alternatively, the SaaS can be directly deployed on the IaaS. PaaS is the platform for software operation, such as databases, Web (World Wide Web) containers, etc. SaaS is various business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are the upper layers relative to IaaS.
[0045] Computer Vision (CV) is a science that studies how to make machines "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition and measurement in machine vision, and further performing graphics processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision research related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, MAE, etc. can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies.
[0046] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration. Pre-trained models are the latest development results of deep learning, integrating the above technologies.
[0047] An embodiment of this application provides a schematic diagram of a method for generating a biological information sequence, such as Figure 1As shown, this method can be executed by a computer device, which can be a terminal or a server.
[0048] Exemplarily, the computer device obtains a biological information fragment sequence 20 and a description tag 10; the computer device performs respective corresponding encoding operations on the biological information fragment sequence 20 and the description tag 10 through an encoding network 40 to obtain a biological information identification sequence 50; the computer device decodes the biological information identification sequence 50 through a sequence generation network 60 to obtain a biological information sequence 70 with at least one of function, structure, species, and task.
[0049] The biological information sequence 70 refers to a sequence with biological information.
[0050] Optionally, the biological information sequence 70 includes at least one of an amino acid sequence, a deoxyribonucleic acid (DNA) sequence, and a ribonucleic acid (RNA) sequence, but is not limited thereto, and the embodiments of the present application do not make specific limitations in this regard.
[0051] The biological information fragment sequence 20 includes biological information fragments required to construct the biological information sequence 70.
[0052] Taking the biological information sequence 70 as an amino acid sequence as an example, the biological information fragment sequence 20 includes amino acids required to construct the amino acid sequence, or amino acid fragments required to construct the amino acid sequence. Taking the biological information sequence 70 as a DNA sequence as an example, the biological information fragment sequence 20 includes at least one of the bases adenine, cytosine, guanine, and thymine required to construct the DNA sequence. Taking the biological information sequence 70 as an RNA sequence as an example, the biological information fragment sequence 20 includes at least one of the bases adenine, cytosine, guanine, and uracil required to construct the RNA sequence.
[0053] The description tag 10 indicates the inherent properties of the biological information sequence 70 based on natural language.
[0054] Optionally, the description tag 10 includes at least one of a species tag, a task tag, a function description tag, and a structure description tag, but is not limited thereto, and the embodiments of the present application do not make specific limitations in this regard.
[0055] The species tag is used to indicate the biological species to which the biological information sequence 70 belongs.
[0056] The task tag is used to indicate to the model the generation method of the biological information sequence 70.
[0057] The function description tag is used to indicate the functions of the generated bioinformatics sequence 70, or to indicate the protein functions in the organism to which the bioinformatics sequence 70 belongs. For example, the generated bioinformatics sequence has hydrophilicity, and the generated bioinformatics sequence has reducibility.
[0058] It should be noted that the above examples of the description tag 10 are only exemplary and not specifically limited.
[0059] The bioinformatics fragment tag sequence 30 refers to a sequence obtained by splicing the bioinformatics fragment sequence 20 and the description tag 10.
[0060] Optionally, the way of splicing the description tag 10 and the bioinformatics fragment sequence 20 includes directly splicing the description tag 10 and the bioinformatics fragment sequence 20, or disassembling the description tag 10 and inserting it into the bioinformatics fragment sequence 20, but not limited to this. The embodiments of the present application do not make specific limitations on this.
[0061] The bioinformatics identification sequence 50 refers to a sequence that represents the bioinformatics fragment sequence 20 and the description tag 10 with an identifier, or the bioinformatics identification sequence 50 refers to a sequence obtained by performing respective corresponding encoding operations on the bioinformatics fragment sequence 20 and the description tag 10, or the bioinformatics identification sequence 50 refers to an identification sequence obtained by performing an encoding operation on the bioinformatics fragment sequence 20 and the description tag 10.
[0062] Optionally, the bioinformatics identification sequence 50 includes a bioinformatics fragment sequence identifier corresponding to the bioinformatics fragment sequence 20 and a description tag identifier corresponding to the description tag 10.
[0063] The encoding network 40 is used to perform encoding operations on the bioinformatics fragment sequence 20 and the description tag 10.
[0064] The encoding network 40 includes a first encoding network and a second encoding network. The first encoding network encodes the bioinformatics fragment sequence 20, and the second encoding network encodes the description tag 10.
[0065] The sequence generation network 60 is used to generate the bioinformatics sequence 70 based on the bioinformatics identification sequence 50. It should be noted that the first encoding network and the second encoding network can be the same network.
[0066] Such as Figure 1 shown, the computer device obtains the bioinformatics fragment sequence 20 and the description tag 10, and the description tag 10 can be expressed as: <protein1> 、 <bind> 、 <protein2> 、 <yes>, wherein, <protein1>For indicating the first biological information fragment sequence, <protein2>For indicating a second biological information fragment sequence, <bind> <yes>For indicating the affinity of a biological information sequence generated from two biological information fragment sequences; the biological information fragment sequence 20 can be expressed as SAGIENEYFYEYDSMK, SAGIENEYF; the computer device splices the description tag 10 with the biological information fragment sequence 20 to obtain a biological information fragment tag sequence 30, and the biological information fragment tag sequence 30 can be expressed as: <protein1>SAGIENEYFYEYDS MK <bind> <protein2>SAGIENEYF <yes>; The computer device inputs the biological information fragment tag sequence 30 into the encoding network 40 for encoding. The first encoding network in the encoding network 40 performs a first encoding operation on the biological information fragment sequence 20, and the obtained biological information fragment sequence identifier can be expressed as: [*, 11791, 31137, 354, 2373, 8400, *, *, 11791, 258, 2438, *]. The second encoding network in the encoding network 40 performs a second encoding operation on the description tag 10, and the obtained description tag identifier can be expressed as: [3, *, *, *, *, *, 5, 4, *, *, *, 7]. Then, the finally obtained biological information identifier sequence 50 can be expressed as: [3, 11791, 31137, 354, 2373, 8400, 5, 4, 11791, 258, 2438, 7]; The computer device inputs the biological information identifier sequence 50 into the sequence generation network 60 for decoding to obtain the biological information sequence 70.
[0067] It should be noted that the encoding method of the encoding network for the biological information fragment sequence and the description tag includes any one of the following methods, but is not limited to this; Method 1: Mix and splice the biological information fragment sequence 20 and the description tag 10 to obtain the biological information fragment tag sequence 30. The encoding network 40 performs the corresponding encoding operations on the biological information fragment sequence and the description tag in the biological information fragment tag sequence 30 to obtain the biological information fragment sequence identifier corresponding to the biological information fragment sequence 20 and the description tag identifier corresponding to the description tag 10. Among them, the arrangement positions of the biological information fragment sequence identifier and the description tag identifier in the biological information identifier sequence 50 are the same as the arrangement positions of the biological information fragment sequence 20 and the description tag 10 in the biological information fragment tag sequence 30. Figure 1 This method is adopted in []. Method 2: Input the biological information fragment sequence 20 and the description tag 10 into the encoding network 40 respectively. The encoding network 40 performs the corresponding encoding operations on the biological information fragment sequence and the description tag to obtain the biological information fragment sequence identifier corresponding to the biological information fragment sequence 20 and the description tag identifier corresponding to the description tag 10. The computer device mixes and splices the biological information fragment sequence identifier and the description tag identifier to obtain the biological information identifier sequence 50.
[0068] In summary, the method provided in this embodiment obtains the biological information fragment sequence and the description tag; performs the corresponding encoding operations on the biological information fragment sequence and the description tag through the encoding network to obtain the biological information identifier sequence; decodes the biological information identifier sequence through the sequence generation network to obtain the biological information sequence. By adding the description tag and performing the corresponding encoding operations on the biological information fragment sequence and the description tag respectively, the present application generates a biological information sequence with a specified function or structure, thereby improving the generation efficiency of the biological information sequence.
[0069] Figure 2 The figure shows a schematic architecture diagram of a computer system provided by an embodiment of the present application. The computer system may include: a terminal 100 and a server 200.
[0070] The terminal 100 may be an electronic device such as a mobile phone, a tablet computer, an in-vehicle terminal (car computer), a wearable device, a personal computer (PC), an in-vehicle terminal, an aircraft, a vending terminal, etc. A client for running a target application may be installed in the terminal 100. The target application may be an application generated with reference to a biological information sequence, or other applications provided with a function of generating a biological information sequence. The present application does not make any limitation thereto. In addition, the present application does not make any limitation to the form of the target application, including but not limited to an application (App) installed in the terminal 100, a mini-program, etc., and may also be in the form of a web page.
[0071] The server 200 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, a cloud database, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, a content delivery network (CDN), and a cloud server providing basic cloud computing services such as a big data and artificial palm image recognition platform. The server 200 may be the background server of the above target application, and is used to provide background services for the client of the target application.
[0072] Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve the calculation, storage, processing, and sharing of data. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. applied based on the cloud computing business model, and can form a resource pool, which can be used on demand, flexibly and conveniently. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.
[0073] In some embodiments, the above-mentioned server can also be implemented as a node in a blockchain system. Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Essentially, a blockchain is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer.
[0074] Communication can be carried out between the terminal 100 and the server 200 through a network, such as a wired or wireless network.
[0075] In the method for generating a biological information sequence provided in the embodiments of the present application, the execution subject of each step can be a computer device, and the computer device refers to an electronic device with data computing, processing, and storage capabilities. Taking Figure 2 the implementation environment of the shown solution as an example, the method for generating a biological information sequence can be executed by the terminal 100 (for example, the client of the target application installed and running in the terminal 100 executes the method for generating a biological information sequence), or can be executed by the server 200, or can be executed by the interaction and cooperation of the terminal 100 and the server 200. The present application does not make any limitations in this regard.
[0076] Figure 3 is a flowchart of the method for generating a biological information sequence provided in an exemplary embodiment of the present application. This method can be executed by a computer device, and the computer device can be a terminal or a server. This method includes:
[0077] Step 302: Obtain a biological information fragment sequence and a description tag.
[0078] A biological information sequence refers to a sequence with biological information.
[0079] Optionally, the biological information sequence includes at least one of an amino acid sequence, a deoxyribonucleic acid sequence, and a ribonucleic acid sequence, but is not limited thereto. The embodiments of the present application do not make specific limitations in this regard.
[0080] The biological information fragment sequence includes biological information fragments required for constructing a biological information sequence.
[0081] Taking the biological information sequence as an amino acid sequence as an example, the biological information fragment sequence includes amino acids required for constructing an amino acid sequence, or amino acid fragments required for constructing an amino acid sequence.
[0082] Taking the biological information sequence as a deoxyribonucleic acid sequence as an example, the biological information fragment sequence includes at least one base among adenine, cytosine, guanine, and thymine required to construct the deoxyribonucleic acid sequence.
[0083] Taking the biological information sequence as a sequence as an example, the biological information fragment sequence includes at least one base among adenine, cytosine, guanine, and uracil required to construct the ribonucleic acid sequence.
[0084] The description label indicates the inherent properties of the biological information sequence based on natural language.
[0085] Optionally, the description label includes at least one of a function description label and a structure description label, but is not limited thereto, and the embodiments of the present application do not make specific limitations thereon.
[0086] The function description label is used to indicate the functions possessed by the generated biological information sequence.
[0087] The structure description label is used to indicate the structures possessed by the generated biological information sequence.
[0088] Step 304: Perform a first encoding operation on the biological information fragment sequence through a first encoding network to obtain a biological information fragment sequence identifier.
[0089] The encoding network is used to perform an encoding operation on the biological information fragment sequence and the description label.
[0090] The first encoding operation refers to an operation of converting the biological information fragment sequence into a biological information fragment sequence identifier.
[0091] The biological information fragment sequence identifier refers to the identifier corresponding to the biological information fragment sequence, and the biological information fragment sequence identifier is the unique identifier corresponding to the biological information fragment sequence. For example, if the biological information fragment sequence is: SAGIENEYF, the biological information fragment sequence identifier corresponding to the biological information fragment sequence can be expressed as: 11791, 258, 2438.
[0092] Step 306: Perform a second encoding operation on the description label through a second encoding network to obtain a description label identifier.
[0093] The description label identifier refers to the identifier corresponding to the description label, and the description label identifier is the unique identifier corresponding to the description label. For example, the description label is: <protein1> 、 <bind> 、 <protein2>, the description tag identifier can be represented as: 3, 5, 4.
[0094] The second encoding operation refers to the operation of converting the description tag into a description tag identifier.
[0095] Step 308: Concatenate the biological information fragment sequence identifier and the description tag identifier to obtain a biological information identifier sequence.
[0096] The biological information identifier sequence refers to a sequence that uses identifiers to represent the biological information fragment sequence and the description tag.
[0097] The biological information identifier sequence includes the biological information fragment sequence identifier corresponding to the biological information fragment sequence and the description tag identifier corresponding to the description tag.
[0098] Optionally, the way of concatenating the biological information fragment sequence identifier and the description tag identifier includes directly concatenating the biological information fragment sequence identifier and the description tag identifier, or inserting the disassembled description tag identifier into the biological information fragment sequence identifier, but is not limited thereto, and the embodiments of the present application do not make specific limitations on this.
[0099] Step 310: Decode the biological information identifier sequence through a sequence generation network to obtain a biological information sequence.
[0100] The sequence generation network is used to generate a biological information sequence based on the biological information identifier sequence.
[0101] The decoding operation refers to the operation of restoring the identifier to the information it represents.
[0102] Exemplarily, the computer device decodes the biological information identifier sequence through a sequence generation network to obtain a biological information sequence with at least one of the functions, structures, tasks, and species specified by the description tag.
[0103] Optionally, the sequence generation network can be any one of a pre-trained generative transformer (GPT), a convolutional neural network (CNN), a transformer neural network (Transformer), or a conformer network enhanced by convolution, or a combination of at least two, but is not limited thereto.
[0104] In summary, the method provided in this embodiment obtains a biological information fragment sequence and a description tag; performs respective corresponding encoding operations on the biological information fragment sequence and the description tag through an encoding network, and splices the obtained biological information fragment sequence identifier and description tag identifier to obtain a biological information identifier sequence; decodes the biological information identifier sequence through a sequence generation network to obtain a biological information sequence having at least one of function, structure, task, and species. By adding a description tag and performing respective corresponding encoding operations on the biological information fragment sequence and the description tag, the present application generates a biological information sequence with a specified function or structure, improving the generation efficiency of the biological information sequence.
[0105] Figure 4 FIG. 4 is a flowchart of a method for generating a biological information sequence provided by an exemplary embodiment of the present application. This method can be executed by a computer device, and the computer device can be a terminal or a server. The method includes:
[0106] Step 402: Obtain a biological information fragment sequence and a description tag.
[0107] A biological information sequence refers to a sequence with biological information.
[0108] Optionally, the biological information sequence includes at least one of an amino acid sequence, a deoxyribonucleic acid sequence, and a ribonucleic acid sequence, but is not limited thereto, and the embodiments of the present application do not make specific limitations on this.
[0109] The biological information fragment sequence includes biological information fragments required to construct the biological information sequence.
[0110] The description tag indicates the inherent properties of the biological information sequence based on natural language.
[0111] Optionally, the description tag includes at least one of a function description tag and a structure description tag, but is not limited thereto, and the embodiments of the present application do not make specific limitations on this.
[0112] The function description tag is used to indicate the function possessed by the generated biological information sequence.
[0113] The structure description tag is used to indicate the structure possessed by the generated biological information sequence.
[0114] Step 404: Perform a first encoding operation on the biological information fragment sequence based on a biological information vocabulary through a first encoding network to obtain a biological information fragment sequence identifier.
[0115] The biological information fragment sequence identifier refers to the identifier corresponding to the biological information fragment sequence, and the biological information fragment sequence identifier is the unique identifier corresponding to the biological information fragment sequence.
[0116] The encoding network is used to perform an encoding operation on the biological information fragment sequence and the description tag.
[0117] The first encoding operation refers to the operation of converting the biological information fragment sequence into a biological information fragment sequence identifier.
[0118] The implementation manner of the first encoding operation includes: the computer device obtains the biological information vocabulary; through the first encoding network, based on the biological information vocabulary, the computer device performs the first encoding operation on the biological information fragment sequence to obtain the biological information fragment sequence identifier.
[0119] The biological information vocabulary includes the correspondence between the biological information fragment and the biological information fragment identifier.
[0120] Exemplarily, the first encoding network in the computer device assigns values to the biological information fragment sequence based on the biological information fragment identifiers in the biological information vocabulary to obtain the biological information fragment sequence identifier; that is, the first encoding network searches for the biological information fragment identifier in the biological information vocabulary based on the biological information fragment sequence, so as to determine the biological information fragment identifier corresponding to the biological information fragment sequence.
[0121] In some embodiments, the obtaining manner of the biological information vocabulary includes: the computer device obtains the sample biological information fragment; the computer device splits the sample biological information fragment into sample biological information units, where the sample biological information unit refers to the sub-fragment with a unit size obtained by splitting the sample biological information fragment; the computer device constructs an initial biological information vocabulary based on the sample biological information units, where the initial biological information vocabulary refers to the vocabulary composed of the sample biological information units; the computer device counts the repetition frequency of the sample biological information unit groups in the initial biological information vocabulary, and updates the initial biological information vocabulary according to the repetition frequency; repeat the above step until the size of the initial biological information vocabulary reaches the expected size or the repetition frequency of the sample biological information unit groups is 1; the computer device assigns values to at least one of the sample biological information units and the sample biological information unit groups in the initial biological information vocabulary to obtain the biological information vocabulary.
[0122] The sample biological information unit group refers to a combination composed of at least two consecutive sample biological information units.
[0123] The repetition frequency refers to the number of repetitions of the sample biological information unit or the sample biological information unit group in the initial biological information vocabulary.
[0124] Exemplarily, the computer device counts the repetition frequency of the sample biological information unit groups in the initial biological information vocabulary, saves the sample biological information unit group with the highest repetition frequency as the new sample biological information unit to the initial biological information vocabulary; the computer device updates the initial biological information vocabulary based on the new sample biological information unit.
[0125] The new sample biological information unit refers to taking the sample biological information unit group as a new sample biological information unit.
[0126] In some embodiments, the method for updating the initial biological information vocabulary includes: the computer device deletes the sample biological information units that make up the new sample biological information unit based on the new sample biological information unit, so as to obtain the updated initial biological information vocabulary.
[0127] For example, the sample biological information fragments obtained by the computer device are: SAG, SAGI, SAGIE, NEGIE. The computer device splits the sample biological information fragments into sample biological information units. The initial biological information vocabulary composed of the sample biological information units can be expressed as: (S, A, G, I, E, N). The computer device counts the repetition frequency of the sample biological information unit groups in the initial biological information vocabulary and finds that the repetition frequency of SAG is the highest, which is 3 times. Therefore, SAG is taken as the new sample biological information unit, and the sample biological information units that make up the new sample biological information unit are deleted. That is, the updated initial biological information vocabulary can be expressed as: (SAG, I, E, N); repeating the statistics of the repetition frequency, it is found that the repetition frequency of IE is the highest, which is 2 times. Therefore, IE is taken as the new sample biological information unit, and the sample biological information units that make up the new sample biological information unit are deleted. Since E is also used elsewhere, E is retained. Therefore, the updated initial biological information vocabulary can be expressed as: (SAG, IE, E, N); since the repetition frequency of the sample biological information unit groups in the initial biological information vocabulary is 1, the updated initial biological information vocabulary is assigned values to obtain the biological information vocabulary.
[0128] Step 406: Perform a second encoding operation on the description label based on the label encoding mapping table through the second encoding network to obtain a description label identifier.
[0129] The description label identifier refers to the identifier corresponding to the description label, and the description label identifier is the unique identifier corresponding to the description label.
[0130] The second encoding operation refers to the operation of converting the description label into a description label identifier.
[0131] In some embodiments, the computer device obtains a label encoding mapping table, where the label encoding mapping table includes the correspondence between the description label and the description label identifier; perform a second encoding operation on the description label based on the label encoding mapping table through the second encoding network to obtain a description label identifier.
[0132] Exemplarily, the second encoding network assigns values to the description tags based on the tag encoding mapping table to obtain the description tag identifiers. The second encoding network in the computer device assigns values to the description tags based on the description tag identifiers in the tag encoding mapping table to obtain the description tag identifiers; that is, the second encoding network looks up the description tag identifiers in the tag encoding mapping table based on the description tags, so as to determine the description tag identifiers corresponding to the description tags.
[0133] In some embodiments, during the generation of the bioinformatics sequence, for example, the sequence composed of the bioinformatics fragment sequence and the description tags is: <eukaryota> <metazoa> <chordata> <craniata><Vertebrata> <euteleostomi> <actinopterygii> <neopterygii> <teleostei> <ostariophysi> <cypriniformes> <nemacheilidae> <triplophysa><:>MEEITQIKKRLSQTVRLEGKE DLLSKKDSITNLKTEEHVSVKKMVISEPKPEKKEDIQLKK, where <:> is the prompt end symbol, the front is the description label, and the back is the bioinformatics fragment sequence. Since the number of description labels is relatively large, the method of remapping the description labels can be used to map the description labels to label tokens Token, and then convert the label tokens into description label identifiers. For example, the mapping relationship between the description labels and the label tokens is shown in Table 1.
[0134] Table 1 Mapping relationship between description labels and label tokens
[0135]
[0136]
[0137] When the number of description labels is less than 200, use 1-bit label token Token encoding. When the number of description labels is in the range of [200, 40000], use 2-bit label token Token encoding. When the number of description labels is in the range of [40000, 8000000], use 3-bit label token Token encoding.
[0138] For example, taking the task of judging the affinity of proteins as an example, <protein1>SAGIENEYFYEYDSM K <bind> <protein2>SAGIENEYF <yes>Pseudo-natural language expression of a protein, wherein the description label is: <protein1> 、 <bind> 、 <protein2> 、 <yes>, the biological information fragment sequences are SAGIE NEYFYEYDSMK, SAGIENEYF; the result after remapping the description tags is: <r0>SAGI ENEYFYEYDSMK <r2> <r1>SAGIENEYF <r4>;The further encoded numerical values, i.e., the obtained biological information identification sequence is: [3, 11791, 31137, 354, 2373, 8400, 5, 4, 11791, 258, 2438, 7].
[0139] Step 408: Concatenate the biological information fragment sequence identifier and the description tag identifier to obtain the biological information identification sequence.
[0140] The biological information identification sequence refers to the sequence that represents the biological information fragment sequence and the description tag with identifiers.
[0141] Exemplarily, after obtaining the biological information fragment sequence identifier and the description tag identifier, the biological information fragment sequence identifier and the description tag identifier are mixed and concatenated to obtain the biological information identification sequence.
[0142] Optionally, the way of concatenating the biological information fragment sequence identifier and the description tag identifier includes directly concatenating the biological information fragment sequence identifier and the description tag identifier, or disassembling the description tag identifier and inserting it into the biological information fragment sequence identifier, but is not limited thereto. The embodiments of the present application do not make specific limitations on this.
[0143] In some embodiments, the way of encoding the biological information fragment sequence and the description tag by the encoding network includes: As Figure 5 shown in the schematic diagram of encoding the biological information fragment sequence and the description tag by the encoding network, the computer device inputs the biological information fragment sequence 501 and the description tag 502 into the encoding network 503 respectively. The encoding network 503 performs the respective corresponding encoding operations on the biological information fragment sequence 501 and the description tag 502, obtains the biological information fragment sequence identifier 504 corresponding to the biological information fragment sequence 501 and the description tag identifier 505 corresponding to the description tag 502. The computer device mixes and concatenates the biological information fragment sequence identifier 504 and the description tag identifier 505 to obtain the biological information identification sequence 506.
[0144] In some embodiments, the way of encoding the biological information fragment sequence and the description tag by the encoding network further includes: As Figure 6 The schematic diagram showing the encoded network encoding the biological information fragment sequence and the description label. The computer device mixes and splices the biological information fragment sequence 601 and the description label 602 to obtain the biological information fragment label sequence 603. The encoding network 604 performs respective corresponding encoding operations on the biological information fragment sequence 601 and the description label 602 in the biological information fragment label sequence 603 to obtain the biological information fragment sequence identifier corresponding to the biological information fragment sequence 602 and the description label identifier corresponding to the description label 601. Among them, the arrangement positions of the biological information fragment sequence identifier and the description label identifier in the biological information identifier sequence 605 are the same as the arrangement positions of the biological information fragment sequence 602 and the description label 601 in the biological information fragment label sequence 603.
[0145] Step 410: Decode the biological information identifier sequence through the sequence generation network to obtain the biological information sequence.
[0146] The sequence generation network is used to generate the biological information sequence based on the biological information identifier sequence.
[0147] The decoding operation refers to the operation of restoring the identifier to the information it represents.
[0148] Exemplarily, the computer device decodes the biological information identifier sequence through the sequence generation network to obtain the biological information sequence with at least one of the functions, structures, species, and tasks specified by the description label.
[0149] Optionally, the sequence generation network can be any one of a pre-trained generative transformer (GPT), a convolutional neural network (CNN), a transformer neural network (Transformer), or a conformer network enhanced by convolution, or a combination of at least two, but not limited to this.
[0150] Exemplarily, the computer device inputs the biological information identifier sequence into the sequence generation network and obtains the biological information sequence in an iterative output manner.
[0151] Among them, the sequence generation network generates the (i + 1)-th output result based on the biological information identifier sequence and the first i output results, and stops outputting until the (i + 1)-th output result meets the preset termination result, where i is a positive integer.
[0152] To verify the generation effect of the method for generating a biological information sequence proposed in the embodiments of the present application, five description tags are used, and the five description tags can be respectively expressed as: <Phage_lysozyme><:>, <Glyco_hydro_108><:>, <Glucosaminidase><:>, <transglycosylas><:>, <Pesticin>, by comparing the structure of the natural enzyme and the structure of the synthetic enzyme obtained by the method of the embodiments of the present application, so as to determine the generation effect of the biological information sequence. As Figure 7 As shown in the schematic diagram of the structure comparison between the natural enzyme and the synthetic enzyme, it can be seen from the figure that the structure of the synthetic enzyme is basically the same as that of the natural enzyme, indicating that the generation method of the biological information sequence proposed in the embodiments of the present application has a good generation effect.
[0153] In summary, the method provided in this embodiment includes obtaining a biological information fragment sequence and a description tag; performing respective corresponding encoding operations on the biological information fragment sequence and the description tag through an encoding network, and splicing the obtained biological information fragment sequence identifier and description tag identifier to obtain a biological information identifier sequence; decoding the biological information identifier sequence through a sequence generation network to obtain a biological information sequence. By adding a description tag and performing respective corresponding encoding operations on the biological information fragment sequence and the description tag, the present application generates a biological information sequence with a specified function or structure, thereby improving the generation efficiency of the biological information sequence.
[0154] In the method provided in this embodiment, by encoding the biological information fragment sequence and the description tag separately, that is, performing respective corresponding encoding operations on the biological information fragment sequence and the description tag. On the one hand, since the languages of the biological information fragment sequence and the description tag are different, if the same encoding operation is performed on the biological information fragment sequence and the description tag, the same vocabulary needs to be used to encode the biological information fragment sequence and the description tag, which will weaken the encoding ability of the encoding network and thus lead to a decrease in the operation speed of the model; on the other hand, if the biological information fragment sequence identifier corresponding to the biological information fragment sequence and the description tag identifier corresponding to the description tag are aggregated together, the vocabulary size in the vocabulary will be too large, resulting in memory overflow.
[0155] In the method provided in this embodiment, by encoding the description tag, the description tag can also be converted into an identifier that the model can understand, thereby generating a biological information sequence with a specified function or structure, and improving the generation efficiency of the biological information sequence.
[0156] Figure 8 FIG. is a schematic diagram of a method for generating a biological information sequence provided by an exemplary embodiment of the present application. This method can be executed by a computer device, and the computer device can be a terminal or a server. The method includes:
[0157] The computer device displays a biological information fragment input box 801. By inputting the biological information fragments required to construct the biological information sequence 807 into the biological information fragment input box 801, the biological information fragments input in the biological information fragment input box 801 form a biological information fragment sequence 802. On the other hand, the required description label 804 is selected from the description label list 803.
[0158] The computer device inputs the biological information fragment sequence 802 and the description label 804 into the encoding network 805 for encoding to obtain a biological information identification sequence. The biological information identification sequence includes the biological information fragment sequence identification corresponding to the biological information fragment sequence 802 and the description label identification corresponding to the description label 804. The encoding network 805 is used to perform encoding operations on the biological information fragment sequence 802 and the description label 804.
[0159] The computer device inputs the biological information identification sequence into the sequence generation network 806 for decoding to obtain a biological information sequence 807 with at least one of function, structure, species, and task. The sequence generation network 806 is used to generate the biological information sequence 807 based on the biological information identification sequence.
[0160] In summary, the method provided in this embodiment obtains a biological information fragment sequence and a description label; performs respective corresponding encoding operations on the biological information fragment sequence and the description label through an encoding network to obtain a biological information identification sequence; and decodes the biological information identification sequence through a sequence generation network to obtain a biological information sequence. By adding a description label and performing respective corresponding encoding operations on the biological information fragment sequence and the description label, the present application generates a biological information sequence with a specified function or structure, thereby improving the generation efficiency of the biological information sequence.
[0161] The above embodiment describes the method for generating a biological information sequence. Next, the training method of the biological information sequence generation model will be described.
[0162] The training method of the biological information sequence generation model involved in the present application can be implemented based on the training system of the biological information sequence generation model. This solution includes the generation stage of the training system of the biological information sequence generation model and the training stage of the biological information sequence generation model. Figure 9 is a framework diagram showing the generation of a training system of a biological information sequence generation model and the training of a biological information sequence generation model, as Figure 9 As shown in the figure, in the generation stage of the training system of the biological information sequence generation model, after the training system generation device 910 of the biological information sequence generation model obtains the training system of the biological information sequence generation model through a preset training sample data set, the training result of the biological information sequence generation model is generated based on the training system of the biological information sequence generation model. In the training stage of the biological information sequence generation model, the training device 920 of the biological information sequence generation model processes the input biological information fragment sequence and description label based on the training system of the biological information sequence generation model to obtain the training result of the biological information sequence generation model.
[0163] Among them, the training system generation device 910 of the biological information sequence generation model and the training device 920 of the biological information sequence generation model can be computer devices. For example, the computer device can be a fixed computer device such as a personal computer or a server, or the computer device can also be a mobile computer device such as a tablet computer or an e-reader.
[0164] Optionally, the training system generation device 910 of the biological information sequence generation model and the training device 920 of the biological information sequence generation model can be the same device, or the training system generation device 910 of the biological information sequence generation model and the training device 920 of the biological information sequence generation model can also be different devices. Moreover, when the training system generation device 910 of the biological information sequence generation model and the training device 920 of the biological information sequence generation model are different devices, the training system generation device 910 of the biological information sequence generation model and the training device 920 of the biological information sequence generation model can be the same type of device. For example, the training system generation device 910 of the biological information sequence generation model and the training device 920 of the biological information sequence generation model can both be servers; or the training system generation device 910 of the biological information sequence generation model and the training device 920 of the biological information sequence generation model can also be different types of devices. For example, the training device 920 of the biological information sequence generation model can be a personal computer or a terminal, while the training system generation device 910 of the biological information sequence generation model can be a server, etc. The embodiments of the present application do not limit the specific types of the training system generation device 910 of the biological information sequence generation model and the training device 920 of the biological information sequence generation model.
[0165] Figure 10 It is a flowchart of a training method for a biological information sequence generation model provided by an exemplary embodiment of the present application. This method can be executed by a computer device, which can be a terminal or a server. A biological information sequence generation model is set in the computer device, and the biological information sequence generation model includes an encoding network and a sequence generation network. This method includes:
[0166] Step 1002: Obtain the sample biological information fragment sequence, the sample description label, and the sample biological information sequence.
[0167] The sample biological information sequence refers to a sequence with biological information.
[0168] The sample biological information sequence refers to a sequence composed of the sample biological information fragment sequences that conforms to the sample description label. This sample biological information sequence is used as a reference sequence to compare the error between the sample biological information sequence and the predicted biological information sequence.
[0169] Optionally, the sample biological information sequence includes at least one of an amino acid sequence, a deoxyribonucleic acid sequence, and a ribonucleic acid sequence, but is not limited thereto. The embodiments of the present application do not make specific limitations in this regard.
[0170] The sample biological information fragment sequence includes the biological information fragments required to construct the sample biological information sequence.
[0171] Taking the sample biological information sequence as an amino acid sequence as an example, the sample biological information fragment sequence includes the amino acids required to construct the amino acid sequence, or the amino acid fragments required to construct the amino acid sequence.
[0172] Taking the sample biological information sequence as a deoxyribonucleic acid sequence as an example, the sample biological information fragment sequence includes at least one of the bases adenine, cytosine, guanine, and thymine required to construct the deoxyribonucleic acid sequence.
[0173] Taking the sample biological information sequence as a sequence as an example, the sample biological information fragment sequence includes at least one of the bases adenine, cytosine, guanine, and uracil required to construct the ribonucleic acid sequence.
[0174] The sample description label indicates the inherent properties of the sample biological information sequence based on natural language.
[0175] Optionally, the sample description label includes at least one of a function description label and a structure description label, but is not limited thereto. The embodiments of the present application do not make specific limitations in this regard.
[0176] The function description label is used to indicate the function possessed by the generated biological information sequence.
[0177] The structure description label is used to indicate the structure possessed by the generated biological information sequence.
[0178] Step 1004: Perform a first encoding operation on the sample biological information fragment sequence through a first encoding network to obtain a sample biological information fragment sequence identifier.
[0179] The sample biological information fragment sequence identifier refers to the identifier corresponding to the sample biological information fragment sequence, and the sample biological information fragment sequence identifier is the unique identifier corresponding to the sample biological information fragment sequence. For example, if the sample biological information fragment sequence is: SAGIENEYF, the sample biological information fragment sequence identifier corresponding to the sample biological information fragment sequence can be expressed as: 11791, 258, 2438.
[0180] The encoding network is used to perform an encoding operation on the sample biological information fragment sequence and the sample description tag.
[0181] The first encoding network is used to perform an encoding operation on the sample biological information fragment sequence.
[0182] The encoding operation refers to the operation of converting the information in the sample biological information fragment sequence and the sample description tag into an identifier.
[0183] The first encoding operation refers to the operation of converting the sample biological information fragment sequence into the sample biological information fragment sequence identifier.
[0184] The implementation method of the first encoding operation includes: the computer device obtains the biological information vocabulary; the first encoding network performs the first encoding operation on the sample biological information fragment sequence based on the biological information vocabulary to obtain the sample biological information fragment sequence identifier.
[0185] The biological information vocabulary includes the corresponding relationship between the sample biological information fragment and the sample biological information fragment identifier.
[0186] Exemplarily, the encoding network in the computer device assigns a value to the sample biological information fragment sequence based on the sample biological information fragment identifier in the sample biological information vocabulary to obtain the sample biological information fragment sequence identifier; that is, the encoding network searches for the sample biological information fragment identifier in the biological information vocabulary based on the sample biological information fragment sequence, so as to determine the sample biological information fragment identifier corresponding to the sample biological information fragment sequence.
[0187] Exemplarily, the acquisition method of the biological information vocabulary includes: the computer device acquires sample biological information fragments; the computer device splits the sample biological information fragments into sample biological information units, where the sample biological information unit refers to a sub-fragment of a unit size obtained by splitting the sample biological information fragments; the computer device constructs an initial biological information vocabulary based on the sample biological information units, where the initial biological information vocabulary refers to a vocabulary composed of sample biological information units; the computer device counts the repetition frequency of the sample biological information unit groups in the initial biological information vocabulary and updates the initial biological information vocabulary according to the repetition frequency; repeat the above step until the size of the initial biological information vocabulary reaches the expected size or the repetition frequency of the sample biological information unit groups is 1; the computer device assigns values to at least one of the sample biological information units and the sample biological information unit groups in the initial biological information vocabulary to obtain the biological information vocabulary.
[0188] The sample biological information unit group refers to a combination composed of at least two consecutive sample biological information units.
[0189] The repetition frequency refers to the number of repetitions of the sample biological information unit or the sample biological information unit group in the initial biological information vocabulary.
[0190] Exemplarily, the computer device counts the repetition frequency of the sample biological information unit groups in the initial biological information vocabulary, saves the sample biological information unit group with the highest repetition frequency as a new sample biological information unit to the initial biological information vocabulary; the computer device updates the initial biological information vocabulary based on the new sample biological information unit.
[0191] The new sample biological information unit refers to taking the sample biological information unit group as a new sample biological information unit.
[0192] Exemplarily, the method for updating the initial biological information vocabulary includes: the computer device deletes the sample biological information units that make up the new sample biological information unit based on the new sample biological information unit to obtain the updated initial biological information vocabulary.
[0193] Step 1006: Perform a second encoding operation on the sample description tag through the second encoding network to obtain a sample description tag identifier.
[0194] The sample description tag identifier refers to the identifier corresponding to the sample description tag, and the sample description tag identifier is the unique identifier corresponding to the sample description tag. For example, the sample description tag is: <protein1> 、 <bind> 、 <protein2>, the sample description label identification can be represented as: 3, 5, 4.
[0195] The second encoding operation refers to the operation of converting the sample description label into the sample description label identification.
[0196] Exemplarily, the computer device obtains a label encoding mapping table, which includes the correspondence between the sample description label and the sample description label identification; the second encoding network performs a second encoding operation on the sample description label based on the label encoding mapping table to obtain the sample description label identification.
[0197] Exemplarily, the encoding network assigns a value to the sample description label based on the label encoding mapping table to obtain the sample description label identification. The encoding network in the computer device assigns a value to the description label based on the sample description label identification in the label encoding mapping table to obtain the sample description label identification; that is, the encoding network looks up the sample description label identification in the label encoding mapping table based on the sample description label, so as to determine the sample description label identification corresponding to the sample description label.
[0198] Step 1008: Concatenate the sample biological information fragment sequence identification and the sample description label identification to obtain the sample biological information identification sequence.
[0199] The sample biological information identification sequence refers to the identification sequence obtained by performing an encoding operation on the sample biological information fragment sequence and the sample description label.
[0200] The sample biological information identification sequence includes the sample biological information fragment sequence identification corresponding to the sample biological information fragment sequence and the sample description label identification corresponding to the sample description label.
[0201] Exemplarily, after obtaining the biological information fragment sequence identification and the description label identification, the biological information fragment sequence identification and the description label identification are mixed and concatenated to obtain the biological information identification sequence.
[0202] Optionally, the way of concatenating the biological information fragment sequence identification and the description label identification includes directly concatenating the biological information fragment sequence identification and the description label identification, or inserting the disassembled description label identification into the biological information fragment sequence identification, but is not limited thereto, and the embodiments of the present application do not make specific limitations on this.
[0203] Step 1010: Decode the sample biological information identification sequence through a sequence generation network to obtain the predicted biological information sequence.
[0204] The prediction sequence generation network is used to generate a predicted biological information sequence based on the sample biological information identification sequence.
[0205] The decoding operation refers to the operation of restoring the identification to the information it represents.
[0206] Exemplarily, the computer device decodes the sample biological information identification sequence through a sequence generation network to obtain a predicted biological information sequence.
[0207] Optionally, the sequence generation network can be any one of a pre-trained Generative Pre-trained Transformer (GPT), a Convolutional Neural Network (CNN), a Transformer neural network, or a Conformer network with convolution enhancement, or a combination of at least two, but not limited to this.
[0208] Exemplarily, the computer device inputs the sample biological information identification sequence into the sequence generation network and obtains the predicted biological information sequence by means of iterative output.
[0209] Among them, the sequence generation network generates the (i + 1)-th output result based on the sample biological information identification sequence and the first i output results, and stops outputting until the (i + 1)-th output result conforms to a preset termination result, where i is a positive integer.
[0210] Step 1012: Calculate a training loss based on the sample biological information sequence and the predicted biological information sequence.
[0211] Exemplarily, the computer device calculates the training loss of the biological information sequence generation model based on the sample biological information sequence and the predicted biological information sequence.
[0212] The training loss refers to the difference value between the input and output of the biological information sequence generation model, and the performance of the biological information sequence generation model is measured by the training loss.
[0213] Step 1014: Update the model parameters of at least one of the first encoding network, the second encoding network, and the sequence generation network according to the training loss.
[0214] Exemplarily, the computer device updates the model parameters of at least one of the first encoding network, the second encoding network, and the sequence generation network in the biological information sequence generation model according to the training loss.
[0215] The update of the model parameters refers to updating the network parameters in the biological information sequence generation model, or updating the network parameters of each network module in the model, or updating the network parameters of each network layer in the model, but not limited to this, and the embodiments of the present application do not make limitations in this regard.
[0216] Based on the loss function value, use the loss function value as a training metric to update the model parameters of at least one of the first encoding network, the second encoding network, and the sequence generation network in the biological information sequence generation model until the loss function value converges, so as to obtain a trained biological information sequence generation model.
[0217] The convergence of the loss function value means that the loss function value no longer changes, or the error difference between two adjacent iterations during the training of the biological information sequence generation model is less than a preset value, or the number of training times of the biological information sequence generation model reaches at least one of the preset number of times, but is not limited to this. The embodiments of the present application do not make limitations in this regard.
[0218] Optionally, the target condition satisfied by the training can be that the number of training iterations of the initial model reaches the target number, and those skilled in the art can preset the number of training iterations in advance. Or, the target condition satisfied by the training can be that the loss value meets the target threshold condition, such as the loss value is less than 0.00001, but is not limited to this. The embodiments of the present application do not make limitations in this regard.
[0219] In summary, the method provided in this embodiment obtains a sample biological information fragment sequence, a sample description label, and a sample biological information sequence; performs respective corresponding encoding operations on the sample biological information fragment sequence and the sample description label through an encoding network to obtain a sample biological information identification sequence; decodes the sample biological information identification sequence through a sequence generation network to obtain a predicted biological information sequence; calculates a training loss based on the sample biological information sequence and the predicted biological information sequence; and updates the model parameters of at least one of the first encoding network, the second encoding network, and the sequence generation network according to the training loss. By training the biological information sequence generation model, the present application enables the biological information sequence generation model to generate a biological information sequence with a specified function or structure according to a description label, improving the generation effect of the biological information sequence generation model.
[0220] The embodiments of the present application provide a method for generating a biological information sequence, which can be executed by a computer device, and the computer device can be a terminal or a server.
[0221] The bioinformatics sequence generation model provided by the embodiments of the present application is mainly used for protein generation. During use, the generation of proteins is controlled by using functional description tags and structural description tags of proteins. The description tags are selected from the given tag list. The description tags and the bioinformatics fragment sequence form a complete prompt sentence (which can also be called a bioinformatics fragment tag sequence), and the prompt sentence is input into an encoding network - sequence generation network, such as models like GPT, to generate an amino acid sequence. The amino acid sequence can predict its structure through a structure sequence model, such as AlphaFold, ESMFold, etc. After generating the corresponding structure, the structure of the protein can be three - dimensionally reconstructed for easy understanding.
[0222] To make the encoding network have the highest amino acid encoding efficiency, optionally, the encoding network can adopt at least one of Byte Pair Encoding (BPE) tokenizer, Unigram language model, and Subword algorithm for turning words into smaller units. The embodiments of the present application adopt the BPE tokenizer and train the bioinformatics vocabulary of the BPE tokenizer by using the amino acid sequence. To enable the tokenizer to handle any number of description tags, we reserve 200 tag tokens for the encoding of tag descriptions, and the values of the 200 tag tokens are <r0> 、 <r1> 、 <r2>Up to <r199>By using the Sentencepiece tool to train the BPE tokenizer, the character coverage rate is 100%. Before training, we need to prepare the dataset, which is also the biological information vocabulary used for training.
[0223] As Figure 11 Shown in the schematic diagram of training the encoding network, constructing a corpus is the most crucial step in training the encoding network. We use the multi-species reference information group 1101 as the data source. By randomly sampling a certain number of amino acid sequences, we obtain the biological information fragment 1102 and construct the corpus 1103 based on the biological information fragment 1102. The first encoding network in the encoding network 1105 is trained with the corpus in the corpus 1103. After training, the first encoding network in the encoding network 1105 can only encode the biological information fragment 1102, such as amino acid sequences. For the second encoding network in the encoding network 1105 to encode the description tags, the identification list 1104 is required for implementation. The label token Token in the reserved identification list 1104 in the encoding network 1105 is used for encoding the description tags.
[0224] There are various related tasks for proteins, such as generating proteins with specific functions, protein classification, and expression level regression. In these tasks, the data not only contains the biological information fragment sequences but also the description tags of the biological information fragment sequences. If the encoding network cannot encode the description tags of proteins, a lot of information will be lost. In the task of annotating proteins, concise description tags are the most common. Common protein annotation libraries include Pfam and GO databases, and each database has an independent tag annotation system. Each protein sequence usually includes multiple description tag annotations. For example, the species to which the sequence belongs and what functions it has. In addition, when constructing downstream tasks, certain tags are also required to construct pseudo-natural language. Such as: <eukaryota> <metazoa> <chordata> <craniata> <vertebrata> <euteleostomi><Actino pterygii> <neopterygii> <teleostei> <ostariophysi> <cypriniformes> <nemacheilidae> <triplophysa><:>MEEITQIKKRLSQTVRLEGKEDLLSKKDSITNLKTEEHVSVKKMVISEPKPEKKEDIQLKK, where <:> is the prompt end symbol, and the subsequent is the biological information fragment sequence corresponding to the description tag.
[0225] Figure 12 The structural schematic diagram of the generating device for the biological information sequence provided by an exemplary embodiment of the present application is shown. This device can be implemented as all or part of a computer device through software, hardware, or a combination of both. The device includes:
[0226] An acquisition module 1201, configured to acquire a biological information fragment sequence and a description tag, where the biological information fragment sequence includes the biological information fragments required to construct the biological information sequence, and the description tag is based on natural language to indicate the inherent properties of the biological information sequence;
[0227] An encoding module 1202, configured to perform a first encoding operation on the biological information fragment sequence through a first encoding network to obtain a biological information fragment sequence identifier; perform a second encoding operation on the description tag through a second encoding network to obtain a description tag identifier; splice the biological information fragment sequence identifier and the description tag identifier to obtain a biological information identifier sequence;
[0228] A decoding module 1203, configured to decode the biological information identifier sequence through a sequence generation network to obtain the biological information sequence.
[0229] In some embodiments, the acquisition module 1201 is further configured to acquire a biological information vocabulary, where the biological information vocabulary includes the correspondence between biological information fragments and biological information fragment identifiers.
[0230] In some embodiments, the encoding module 1202 is further configured to perform the first encoding operation on the biological information fragment sequence by the first encoding network based on the biological information vocabulary to obtain the biological information fragment sequence identifier.
[0231] In some embodiments, the encoding module 1202 is further configured to assign values to the biological information fragment sequence by the first encoding network based on the biological information fragment identifiers in the biological information vocabulary to obtain the biological information fragment sequence identifier.
[0232] In some embodiments, the acquisition module 1201 is further configured to acquire sample biological information fragments; split the sample biological information fragments into sample biological information units, where the sample biological information unit refers to a sub-fragment of unit size obtained by splitting the sample biological information fragments; construct an initial biological information vocabulary based on the sample biological information units, where the initial biological information vocabulary refers to a vocabulary composed of the sample biological information units; count the repetition frequency of the sample biological information unit groups in the initial biological information vocabulary, and update the initial biological information vocabulary according to the repetition frequency, where the sample biological information unit group refers to a combination composed of at least two consecutive sample biological information units; repeat the above step until the size of the initial biological information vocabulary reaches the desired size or the repetition frequency of the sample biological information unit group is 1; assign values to at least one of the sample biological information units and the sample biological information unit groups in the initial biological information vocabulary to obtain the biological information vocabulary.
[0233] In some embodiments, the device further includes a calculation module 1204. The calculation module 1204 is configured to count the repetition frequency of the sample biological information unit groups in the initial biological information vocabulary, and save the sample biological information unit group with the highest repetition frequency as a new sample biological information unit to the initial biological information vocabulary, where the new sample biological information unit refers to taking the sample biological information unit group as a new sample biological information unit.
[0234] In some embodiments, the device further includes an update module 1205. The update module 1205 is configured to update the initial biological information vocabulary based on the new sample biological information unit.
[0235] In some embodiments, the update module 1205 is further configured to delete the sample biological information units that make up the new sample biological information unit based on the new sample biological information unit to obtain the updated initial biological information vocabulary.
[0236] In some embodiments, the acquisition module 1201 is further configured to acquire a tag coding mapping table, where the tag coding mapping table includes the correspondence between the description tags and the description tag identifiers.
[0237] In some embodiments, the encoding module 1202 is further configured to perform the second encoding operation on the description tags based on the tag coding mapping table through a second encoding network to obtain the description tag identifiers.
[0238] In some embodiments, the encoding module 1202 is further configured to assign values to the description tags based on the tag coding mapping table through the second encoding network to obtain the description tag identifiers.
[0239] In some embodiments, the decoding module 1203 is further configured to input the biological information identification sequence into the sequence generation network, and obtain the biological information sequence by means of iterative output.
[0240] Wherein, the sequence generation network generates the (i + 1)-th output result based on the biological information identification sequence and the first i output results, and stops outputting until the (i + 1)-th output result meets a preset termination result, where i is a positive integer.
[0241] In some embodiments, the obtaining module 1201 is further configured to obtain a sample biological information fragment sequence, a sample description tag, and a sample biological information sequence, where the sample biological information fragment sequence includes biological information fragments required for constructing the sample biological information sequence, and the sample description tag indicates the inherent properties of the sample biological information sequence based on natural language.
[0242] In some embodiments, the encoding module 1202 is further configured to perform the first encoding operation on the sample biological information fragment sequence through the first encoding network to obtain a sample biological information fragment sequence identifier; perform the second encoding operation on the sample description tag through the first encoding network to obtain a sample description tag identifier; and splice the sample biological information fragment sequence identifier and the sample description tag identifier to obtain a sample biological information identification sequence.
[0243] In some embodiments, the decoding module 1203 is further configured to decode the sample biological information identification sequence through the sequence generation network to obtain a predicted biological information sequence.
[0244] In some embodiments, the calculation module 1204 is further configured to calculate a training loss based on the sample biological information sequence and the predicted biological information sequence.
[0245] In some embodiments, the updating module 1205 is further configured to update model parameters of at least one of the first encoding network, the second encoding network, and the sequence generation network according to the training loss.
[0246] Figure 13 FIG. 0 shows a block diagram of a computer device 1300 shown in an exemplary embodiment of the present application. The computer device may be implemented as the server in the above solution of the present application. The computer device 1300 includes a central processing unit (CPU) 1301, a system memory 1304 including a random access memory (RAM) 1302 and a read-only memory (ROM) 1303, and a system bus 1305 connecting the system memory 1304 and the central processing unit 1301. The computer device 1300 also includes a mass storage device 1306 for storing an operating system 1309, application programs 1310, and other program modules 1311.
[0247] The mass storage device 1306 is connected to the central processing unit 1301 through a mass storage controller (not shown) connected to the system bus 1305. The mass storage device 1306 and its associated computer-readable medium provide non-volatile storage for the computer device 1300. That is to say, the mass storage device 1306 may include a computer-readable medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.
[0248] Without loss of generality, the computer-readable medium may include a computer storage medium and a communication medium. The computer storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer storage medium includes RAM, erasable programmable read-only registers (EPROM), electrically-erasable programmable read-only memories (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile discs (DVD) or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage medium is not limited to the above several. The above system memory 1304 and mass storage device 1306 may be collectively referred to as a memory.
[0249] According to various embodiments of the present disclosure, the computer device 1300 may also operate by connecting to a remote computer on a network such as the Internet. That is, the computer device 1300 may be connected to the network 1308 through the network interface unit 1307 connected to the system bus 1305. Or rather, the network interface unit 1307 may also be used to connect to other types of networks or remote computer systems (not shown).
[0250] The memory further includes at least one segment of computer program. The at least one segment of computer program is stored in the memory, and the central processing unit 1301 implements all or part of the steps in the method for generating a biological information sequence or the method for training a biological information sequence generation model shown in the above various embodiments by executing the at least one segment of program.
[0251] An embodiment of the present application further provides a computer device, which includes a processor and a memory. At least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the method for generating a biological information sequence or the method for training a biological information sequence generation model provided in the above method embodiments.
[0252] An embodiment of the present application further provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is loaded and executed by the processor to implement the method for generating a biological information sequence or the method for training a biological information sequence generation model provided in the above method embodiments.
[0253] An embodiment of the present application further provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The computer program is read and executed by the processor of the computer device, so that the computer device executes to implement the method for generating a biological information sequence or the method for training a biological information sequence generation model provided in the above method embodiments.
[0254] It can be understood that in the specific implementation of the present application, for data, historical data, and portraits and other user data processing related to user identity or characteristics, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.
[0255] It should be noted that, unless otherwise clearly defined herein, all terms used in the claims are to be interpreted according to their ordinary meanings in the technical field. Unless otherwise explicitly stated, all references to "an element, device, component, equipment, step, etc." are to be construed openly as referring to at least one instance of the element, device, component, equipment, step, etc. Unless explicitly stated, the steps of any method disclosed herein are not necessarily to be performed in the exact order disclosed.
[0256] It should be understood that "a plurality of" as mentioned herein means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0257] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.
[0258] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.< / triplophysa> < / nemacheilidae> < / cypriniformes> < / ostariophysi> < / teleostei> < / neopterygii> < / euteleostomi> < / vertebrata> < / craniata> < / chordata> < / metazoa> < / eukaryota> < / r1> < / r0> < / bind> < / protein1> < / transglycosylas> < / r2> < / yes> < / protein2> < / bind> < / protein1> < / yes> < / bind> < / triplophysa> < / nemacheilidae> < / cypriniformes> < / ostariophysi> < / teleostei> < / neopterygii> < / actinopterygii> < / euteleostomi> < / craniata> < / chordata> < / metazoa> < / eukaryota> < / bind> < / protein1> < / yes> < / bind> < / yes> < / bind> < / yes> < / protein2> < / bind> < / protein1>
Claims
1. A method for generating a biological information sequence, characterized in that, The method includes: Obtaining a biological information fragment sequence and a description tag, where the biological information fragment sequence includes the biological information fragments required to construct the biological information sequence, and the description tag is based on natural language to indicate the inherent properties of the biological information sequence; Performing a first encoding operation on the biological information fragment sequence through a first encoding network to obtain a biological information fragment sequence identifier; performing a second encoding operation on the description tag through a second encoding network to obtain a description tag identifier; Concatenating the biological information fragment sequence identifier and the description tag identifier to obtain a biological information identifier sequence, where the biological information identifier sequence refers to a sequence representing the biological information fragment sequence and the description tag with identifiers; Decoding the biological information identifier sequence through a sequence generation network to obtain the biological information sequence.
2. The method according to claim 1, characterized in that, The performing a first encoding operation on the biological information fragment sequence through the first encoding network to obtain a biological information fragment sequence identifier includes: Obtaining a biological information vocabulary, where the biological information vocabulary includes the correspondence between biological information fragments and biological information fragment identifiers; Performing the first encoding operation on the biological information fragment sequence through the first encoding network based on the biological information vocabulary to obtain the biological information fragment sequence identifier.
3. The method according to claim 2, wherein The performing the first encoding operation on the biological information fragment sequence through the first encoding network based on the biological information vocabulary to obtain the biological information fragment sequence identifier includes: Assigning values to the biological information fragment sequence through the first encoding network based on the biological information fragment identifiers in the biological information vocabulary to obtain the biological information fragment sequence identifier.
4. The method according to claim 3, characterized in that, Before obtaining the biological information vocabulary, the method further includes: Obtaining sample biological information fragments; Splitting the sample biological information fragments into sample biological information units, where the sample biological information unit refers to a sub-fragment of unit size obtained by splitting; Constructing an initial biological information vocabulary based on the sample biological information units; Counting the repetition frequency of sample biological information unit groups in the initial biological information vocabulary, and updating the initial biological information vocabulary according to the repetition frequency, where the sample biological information unit group refers to a combination composed of at least two consecutive sample biological information units; Repeating the previous step in this claim until the size of the initial biological information vocabulary reaches the expected size or the repetition frequency of the sample biological information unit group is 1; Assigning values to at least one of the sample biological information units and the sample biological information unit groups in the initial biological information vocabulary to obtain the biological information vocabulary.
5. The method according to claim 4, wherein The counting the repetition frequency of sample biological information unit groups in the initial biological information vocabulary and updating the initial biological information vocabulary according to the repetition frequency includes: Counting the repetition frequency of the sample biological information unit groups in the initial biological information vocabulary, and saving the sample biological information unit group with the highest repetition frequency as a new sample biological information unit to the initial biological information vocabulary; Update the initial biological information vocabulary based on the new sample biological information unit.
6. The method according to claim 5, wherein The updating of the initial biological information vocabulary based on the new sample biological information unit includes: Delete the sample biological information units that make up the new sample biological information unit based on the new sample biological information unit to obtain the updated initial biological information vocabulary.
7. The method according to claim 1, wherein The performing of the second encoding operation on the description tag by the second encoding network to obtain a description tag identifier includes: Obtain a tag encoding mapping table, where the tag encoding mapping table includes the correspondence between the description tag and the description tag identifier; Perform the second encoding operation on the description tag based on the tag encoding mapping table by the second encoding network to obtain the description tag identifier.
8. The method according to claim 7, characterized in that, The performing of the second encoding operation on the description tag based on the tag encoding mapping table by the second encoding network to obtain the description tag identifier includes: Assign a value to the description tag based on the tag encoding mapping table by the second encoding network to obtain the description tag identifier.
9. The method according to any one of claims 1 to 8, characterized in that, The decoding of the biological information identifier sequence by the sequence generation network to obtain the biological information sequence includes: Input the biological information identifier sequence into the sequence generation network and obtain the biological information sequence by means of iterative output; Wherein, the sequence generation network generates the (i + 1)-th output result based on the biological information identifier sequence and the first i output results, and stops outputting until the (i + 1)-th output result meets a preset termination result, and i is a positive integer.
10. The method according to any one of claims 1 to 8, characterized in that The method further includes: Obtain a sample biological information fragment sequence, a sample description tag, and a sample biological information sequence, where the sample biological information fragment sequence includes biological information fragments required to construct the sample biological information sequence, and the sample description tag is based on natural language to indicate the inherent properties of the sample biological information sequence; Perform the first encoding operation on the sample biological information fragment sequence by the first encoding network to obtain a sample biological information fragment sequence identifier; perform the second encoding operation on the sample description tag by the second encoding network to obtain a sample description tag identifier; Concatenate the sample biological information fragment sequence identifier and the sample description tag identifier to obtain a sample biological information identifier sequence; Decode the sample biological information identifier sequence by the sequence generation network to obtain a predicted biological information sequence; Calculate a training loss based on the sample biological information sequence and the predicted biological information sequence; Update the model parameters of at least one of the first encoding network, the second encoding network, and the sequence generation network according to the training loss.
11. A generating device for a biological information sequence, characterized in that, The device includes: An acquisition module, configured to acquire a biological information fragment sequence and a description tag, where the biological information fragment sequence includes biological information fragments required to construct the biological information sequence, and the description tag is based on natural language to indicate the inherent properties of the biological information sequence; An encoding module, configured to perform a first encoding operation on the biological information fragment sequence through a first encoding network to obtain a biological information fragment sequence identifier; perform a second encoding operation on the description tag through the second encoding network to obtain a description tag identifier; splice the biological information fragment sequence identifier and the description tag identifier to obtain a biological information identifier sequence; A decoding module, configured to decode the biological information identifier sequence through a sequence generation network to obtain the biological information sequence.
12. A computer device, characterized in that, The computer device includes: a processor and a memory, and at least one computer program is stored in the memory, and at least one of the computer programs is loaded and executed by the processor to implement the method for generating a biological information sequence according to any one of claims 1 to 10.
13. A computer storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and at least one computer program is loaded and executed by a processor to implement the method for generating a biological information sequence according to any one of claims 1 to 10.
14. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium; the computer program is read and executed by a processor of a computer device from the computer-readable storage medium, so that the computer device executes the method for generating a biological information sequence according to any one of claims 1 to 10.