Protein sequence processing method and device and electronic equipment

Through multi-resolution generator and discriminator network training, combined with generative adversarial networks, the problem of poor protein sequence generation in the prior art is solved, and efficient generation and accuracy of long sequences are achieved.

CN120432004APending Publication Date: 2025-08-05SHANGHAI SHUHUI BIOLOGICAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510105115.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-04
Filing Date
2025-01-22
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing protein sequence generation methods are not effective when generating sequences with lengths exceeding 1024. Traditional machine learning methods cannot fully capture the intrinsic mechanism of protein sequences. Neural network-based models have the challenge of improving training effects in long sequence processing.

Method used

By machine learning of the initial generator network based on multiple vectors, the target generator network is obtained, and the initial discriminator network is trained in combination with multi-resolution prediction and actual protein image sets to form a protein sequence generation model. Using the adversarial learning mechanism of the generative adversarial network and the discriminator network, the generator network is optimized to generate multi-resolution protein sequences.

Benefits of technology

It achieves efficient generation of protein sequences, improves generation accuracy and efficiency, and can generate a wider range of sequences, suitable for protein sequences with lengths exceeding 1024.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120432004A_ABST
    Figure CN120432004A_ABST
Patent Text Reader

Abstract

The invention discloses a protein sequence processing method and device and electronic equipment. The method relates to the field of bioinformatics and computational biology, and comprises the following steps: based on a plurality of vectors, performing machine learning on an initial generator network to obtain a target generator network and predicted protein image sets respectively corresponding to the plurality of vectors; performing machine learning on the initial discriminator network based on predicted protein image sets corresponding to the plurality of vectors and a plurality of actual protein image sets to obtain a target discriminator network and a discrimination result; obtaining a protein sequence generation model based on the target generator network and the discriminator network under the condition that the discrimination result meets a predetermined discrimination condition; a plurality of protein sequences are generated using a protein sequence generation model. According to the invention, the technical problem of poor protein sequence generation effect of a protein sequence generation method in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority of a Chinese patent application with the application number 202410161036.1 and the title "Protein Sequence Processing Method, Device and Electronic Device" filed with the Chinese Patent Office on February 4, 2024, the entire content of which is incorporated herein by reference. Technical field

[0003] The present invention relates to the fields of bioinformatics and computational biology, and in particular, to a protein sequence processing method, device and electronic device. Background technique

[0004] Traditional protein sequence design methods include point mutation and combination methods based on homologous proteins. However, these methods are usually limited by known mechanisms and it is difficult to create completely new protein sequences. Traditional machine learning methods, such as hidden Markov models, etc., utilize existing protein sequence data to discover potential laws in protein sequences and generate protein sequences with specific functions. However, traditional machine learning methods have some limitations. Taking the hidden Markov model as an example, its number of parameters is small and the model complexity is low, unable to fully capture the internal mechanism of protein sequences. Currently, existing sequence generation models based on neural networks, due to their high tunability and powerful representation learning ability, have to some extent solved the problems of traditional machine learning generation models. However, for sequences with a length exceeding 1024, the models in related technologies still face challenges in improving the training effect when processing sequences with a length exceeding 1024.

[0005] In view of the above problems, no effective solution has been proposed yet. Summary of the invention

[0006] Embodiments of the present invention provide a protein sequence processing method, device and electronic device to at least solve the technical problem of poor protein sequence generation effect existing in the protein sequence generation method in related technologies.

[0007] According to one aspect of an embodiment of the present invention, there is provided a method for processing protein sequences, including: performing machine learning on an initial generator network based on a plurality of vectors to obtain a target generator network, and a set of predicted protein images respectively corresponding to the plurality of vectors, wherein the set of predicted protein images includes a plurality of predicted protein sequence maps, and the plurality of predicted protein sequence maps correspond to different resolutions; performing machine learning on an initial discriminator network based on the set of predicted protein images respectively corresponding to the plurality of vectors, and a set of actual protein images to obtain a target discriminator network and a discrimination result, wherein the discrimination result is used to indicate the discrimination degree between the predicted protein sequence maps respectively corresponding to the plurality of vectors and the corresponding actual protein sequence maps, the set of actual protein images includes a plurality of actual protein sequence maps, and the plurality of actual protein sequence maps correspond to different resolutions; obtaining a protein sequence generation model based on the target generator network and the discriminator network when the discrimination result meets a predetermined discrimination condition; and generating a plurality of protein sequences by using the protein sequence generation model.

[0008] According to another aspect of an embodiment of the present invention, there is also provided a protein sequence processing device, including: a first machine learning module, configured to perform machine learning on an initial generator network based on a plurality of vectors to obtain a target generator network, and a set of predicted protein images respectively corresponding to the plurality of vectors, wherein the set of predicted protein images includes a plurality of predicted protein sequence maps, and the plurality of predicted protein sequence maps correspond to different resolutions; a second machine learning module, configured to perform machine learning on an initial discriminator network based on the set of predicted protein images respectively corresponding to the plurality of vectors, and a set of actual protein images to obtain a target discriminator network and a discrimination result, wherein the discrimination result is used to indicate the discrimination degree between the predicted protein sequence maps respectively corresponding to the plurality of vectors and the corresponding actual protein sequence maps, the set of actual protein images includes a plurality of actual protein sequence maps, and the plurality of actual protein sequence maps correspond to different resolutions; an obtaining module, configured to obtain a protein sequence generation model based on the target generator network and the discriminator network when the discrimination result meets a predetermined discrimination condition; and a generating module, configured to generate a plurality of protein sequences by using the protein sequence generation model.

[0009] According to another aspect of an embodiment of the present invention, there is also provided a non-volatile storage medium storing multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform any one of the protein sequence processing methods.

[0010] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including one or more processors and a memory, where the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the protein sequence processing method according to any one of the above.

[0011] In the embodiments of the present invention, by performing machine learning on an initial generator network based on multiple vectors, a target generator network is obtained, as well as a set of predicted protein images corresponding to the multiple vectors respectively, where the set of predicted protein images includes multiple predicted protein sequence maps, and the multiple predicted protein sequence maps correspond to different resolutions; based on the set of predicted protein images corresponding to the multiple vectors respectively, and a set of actual protein images, machine learning is performed on an initial discriminator network to obtain a target discriminator network and a discrimination result, where the discrimination result is used to indicate the discrimination degree between the predicted protein sequence maps corresponding to the multiple vectors respectively and the corresponding actual protein sequence maps, the set of actual protein images includes multiple actual protein sequence maps, and the multiple actual protein sequence maps correspond to different resolutions; when the discrimination result meets a predetermined discrimination condition, a protein sequence generation model is obtained based on the target generator network and the discriminator network; multiple protein sequences are generated using the protein sequence generation model, achieving the purpose of training a protein sequence generation model that integrates protein sequence features of multiple resolutions based on a generator network and a discriminator network for efficient generation of protein sequences, thereby realizing the technical effect of improving the accuracy and efficiency of protein sequence generation, and further solving the technical problem of poor protein sequence generation effect existing in the protein sequence generation method in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:

[0013] Figure 1 is a flowchart of a protein sequence processing method according to an embodiment of the present invention;

[0014] Figure 2 is a flowchart of an optional multi-resolution protein image acquisition according to an embodiment of the present invention;

[0015] Figure 3 is a schematic diagram of an optional protein sequence length distribution according to an embodiment of the present invention;

[0016] Figure 4An optional protein sequence processing flowchart according to an embodiment of the present invention;

[0017] Figure 5 An optional model effect diagram according to an embodiment of the present invention;

[0018] Figure 6 An optional protein sequence distribution comparison diagram according to an embodiment of the present invention;

[0019] Figure 7 An optional comparison schematic diagram of the predicted length and the actual length of the sequence according to an embodiment of the present invention;

[0020] Figure 8 An optional comparison schematic diagram of the consistency after the Cas9 protein sequence length is reduced according to an embodiment of the present invention;

[0021] Figure 9 An optional comparison schematic diagram of the consistency after the IscB protein sequence length is reduced according to an embodiment of the present invention;

[0022] Figure 10 A schematic diagram of a protein sequence processing device according to an embodiment of the present invention. Detailed implementation manners

[0023] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] First, for the convenience of understanding the embodiments of the present invention, some terms or nouns involved in the present invention will be explained below:

[0026] Blastp, a tool for aligning protein sequences, is part of the Basic Local Alignment Search Tool (BLAST) algorithm. Blastp can search for similar protein sequences in a database and generate alignment results and similarity scores. It is commonly used to determine the function of unknown proteins, find homologous genes, and perform protein sequence alignment and analysis.

[0027] t-SNE, a dimensionality reduction algorithm, whose full English name is t-distributed Stochastic Neighbor Embedding and Chinese full name is t distribution - random neighbor embedding.

[0028] Protein sequence generation is an important problem in the fields of bioinformatics and computational biology. It involves generating new protein sequences based on given conditions or data. In this problem, multiple known protein sequences, structural information, functional characteristics, and sequence tags are usually provided, and these are used to guide the training of the generation model. The task of the generation model is to generate a large number of protein sequences with specific properties. Although scientists have studied proteins for decades, designing proteins that can trigger specific chemical reactions has proven to be extremely challenging.

[0029] [[ID=IO]]Protein sequence generation technologies driven by artificial intelligence are gradually attracting more and more researchers' attention. In the field of synthetic biology, protein sequence generation is used to construct proteins with specific functions, such as industrial enzymes or biological components in synthetic biology. In related technologies, a generative adversarial network technology is proposed. Researchers use malate dehydrogenase as a template enzyme, showing that many sequences generated by the ProteinGAN model of protein generation adversarial network are soluble and exhibit the catalytic activity of methanol dehydrogenase (MDH); in related technologies, another Protein Sequence Generation ProGen model based on the Transformer is proposed. This model can be used to generate family sequences of lysozyme. The research team screened two artificial enzymes from the generated sequences, which can decompose bacterial cell walls with an activity comparable to that of the natural lysozyme HEWL. However, current sequence generation models are difficult to capture the long-range correlations of protein sequences and are not applicable to protein family sequences with a large length span.

[0030] It should be noted that in the original text, "[[ID=IO]]" seems to be a mislabeling, and it is translated as "[[ID=IO]]" here. You may want to check and correct it if necessary.Traditional protein sequence design methods include point mutations and combinatorial methods based on homologous proteins. However, these methods are usually limited by known mechanisms and it is difficult to create completely new protein sequences. Traditional machine learning methods, such as hidden Markov models, etc., utilize existing protein sequence data to discover potential patterns in protein sequences and generate protein sequences with specific functions. However, traditional machine learning methods have some limitations. Taking the hidden Markov model as an example, it has a small number of parameters and a low model complexity, and cannot fully capture the internal mechanism of protein sequences. Currently, existing sequence generation models based on neural networks have to some extent solved the problems of traditional machine learning generation models due to their high tunability and powerful representation learning ability. However, for sequences longer than 1024, these models still face challenges in improving the training effect when dealing with sequences longer than 1024. And protein sequence design based on artificial intelligence mainly focuses on generating specific family sequences, but how to make full use of deep generation models for more extensive sequence modification remains an issue to be studied.

[0031] In this application, the terms "polypeptide" and "protein" are used interchangeably herein and refer to polymers of amino acid residues. "Protein sequence" refers to the amino acid residue sequence of a protein and the nucleic acid sequence encoding the amino acid sequence of the protein.

[0032] The terms "nucleic acid", "polynucleotide", "nucleotide" are used interchangeably and refer to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) or polymers thereof in single-stranded or double-stranded form; nucleic acids include nucleic acids containing known nucleotide analogs or modified backbone residues or linkages, and nucleic acids can be synthetic, naturally occurring and non-naturally occurring, such as non-natural nucleic acids having binding properties similar to reference nucleic acids and being metabolized in a manner similar to reference nucleotides. Including but not limited to, phosphorothioates, phosphoroamidates, methylphosphonates, chiral-methylphosphonates, 2-O-methyl ribonucleotides, peptide-nucleic acids (PNAs) modified nucleic acids.

[0033] The term "amino acid" refers to naturally occurring amino acids, synthetic amino acids, as well as amino acid analogs and amino acid mimetics that act in a manner similar to naturally occurring amino acids. Naturally occurring amino acids include those encoded by the genetic code and their modified amino acids, such as hydroxyproline, γ-carboxyglutamic acid, and O-phosphoserine. Common naturally occurring amino acids are, for example: alanine (Ala; A), arginine (Arg; R), asparagine (Asn; N), aspartic acid (Asp; D), cysteine (Cys; C); glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G); histidine (His; H), isoleucine (Ile; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V). Amino acid analogs are compounds that have the same basic chemical structure as naturally occurring amino acids (i.e., an α-carbon bonded to a hydrogen, a carboxyl group, an amino group, and an R group), such as homoserine, norleucine, methionine sulfoxide, and methionine methyl sulfonium. Amino acid analogs typically have a modified R group (e.g., norleucine) or a modified peptide backbone, but retain the same basic chemical structure as naturally occurring amino acids. Amino acid mimetics are chemical compounds that have a structure different from the general chemical structure of amino acids but act in a manner similar to naturally occurring amino acids.

[0034] In an alternative embodiment, the protein sequence is an amino acid residue sequence. In an alternative embodiment, the protein sequence is an RNA sequence. In an alternative embodiment, the protein sequence is a DNA sequence.

[0035] According to an embodiment of the present invention, there is provided an embodiment of a method for processing a protein sequence. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0036] Figure 1 is a flowchart of a method for processing a protein sequence according to an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:

[0037] Step S102: Based on multiple vectors, perform machine learning on the initial generator network to obtain a target generator network and a set of predicted protein images corresponding to each of the multiple vectors. The set of predicted protein images includes multiple predicted protein sequence diagrams, and the multiple predicted protein sequence diagrams correspond to different resolutions.

[0038] Optionally, the initial generator network is the generator network before training. The above multiple vectors can be multiple vectors with a dimension of (128, 1), and the multiple vectors are drawn from a random distribution with a mean of 0 and a standard deviation of 1. Based on the multiple vectors, the initial generator network can be trained through machine learning to obtain the trained generator network, that is, the target generator network. The trained target generator network is used to generate protein images. At the same time, during the training process of the initial generator network, multiple protein sequence diagrams with different resolutions are output in each round of training, forming a set of predicted protein sequences. Through the above method, the initial generator network can learn the mapping relationship between vectors and protein sequence diagrams, which is used for the construction of subsequent protein sequence generation models.

[0039] In an optional embodiment, based on multiple vectors, performing machine learning on the initial generator network to obtain a target generator network and a set of predicted protein images corresponding to each of the multiple vectors includes: based on the multiple vectors, using the multi-layer perceptron in the initial generator network to obtain three-dimensional vectors corresponding to each of the multiple vectors, where the multiple vectors are two-dimensional vectors; based on the three-dimensional vectors corresponding to each of the multiple vectors, using multiple residual basic units in the initial generator network to obtain a set of predicted protein images corresponding to each of the multiple vectors, where the multiple residual basic units in the initial generator network correspond to different resolutions.

[0040] Optionally, the initial generator network includes a multi-layer perceptron and multiple residual basic units. The multiple vectors input to the initial generator network are two-dimensional vectors. During the training process of the initial generator network, first, the multi-layer perceptron in the initial generator network reshapes the two-dimensional multiple vectors into three-dimensional vectors, where c and w respectively represent the number of channels and the vector length, and "d" represents the discriminator. Represents the number of channels of the input of the first residual basic unit of the discriminator; then, through the processing of multiple residual basic units, multiple protein sequence maps of each vector at multiple different resolutions can be obtained. It should be noted that multiple residual basic units correspond to different image resolutions. The initial generator network can not only generate a single high-resolution (i.e., the predetermined resolution) protein image, but also generate protein images at multiple resolutions, including feature maps with resolutions reduced to one-half, one-fourth, one-eighth, and one-sixteenth. In this framework of multi-resolution generation, the initial generator network not only focuses on the generation of protein images at a single resolution, but can also process protein sequence information at different scales. The initial generator network is trained based on multiple different resolutions so that the initial resolution network can learn the features of protein sequences at different scale features to achieve a comprehensive capture of different scale features of protein sequences.

[0041] In an optional embodiment, based on the three-dimensional vectors respectively corresponding to multiple vectors, multiple residual basic units in the initial generator network are used to obtain a set of predicted protein images respectively corresponding to the multiple vectors, including: based on the three-dimensional vectors respectively corresponding to the multiple vectors, any one of the multiple vectors is obtained through the following method, and any one of the predicted protein sequence maps corresponding to any one vector: based on the three-dimensional vector of any one vector, the first transposed convolutional layer in the first residual basic unit is used to obtain a first feature map, where the first residual basic unit is any one of the multiple residual basic units in the initial generator network; based on the first feature map, the convolutional layer with a predetermined first dimension in the first residual basic unit is used to obtain a second feature map; based on the second feature map, the second transposed convolutional layer in the first residual basic unit is used to obtain a third feature map; a residual connection is performed on the second feature map and the third feature map to obtain any one of the predicted protein sequence maps.

[0042] Optionally, both the first transposed convolutional layer and the second transposed convolutional layer are transposed convolutional layers with a stride of 2, and the above-mentioned convolutional layer with a predetermined first dimension is a 1×1 convolutional layer. Each residual basic unit in the initial generator network takes a vector as input, first uses a transposed convolutional layer with a stride of 2, that is, the first transposed convolutional layer to amplify the vector width and obtain a first feature map Then use a 1×1 convolutional layer, that is, the convolutional layer with a predetermined first dimension to process the generated first feature map to capture higher-level features and obtain a second feature map; then, use another transposed convolutional layer with a stride of 2, that is, the second transposed convolutional layer to expand the width of the second feature map and obtain a third feature map, and perform a residual connection on the third feature map and the first feature map obtained in the previous step to obtain the predicted protein sequence map output by the corresponding layer residual basic unit Among them, any one of the predicted protein sequence diagrams is the predicted protein sequence diagram at the corresponding resolution. The third feature map output by the last residual basic unit can also be obtained through a 1×1 convolutional layer to get an output of (21, 1, L) (L represents the sequence length), that is, the predicted protein sequence diagram at the final standard resolution. In order to more comprehensively capture multi-scale features, the output of the generator network not only includes the (21, 1, L) of the last layer, but also includes the low-resolution protein sequence diagrams with dimensions obtained by passing the intermediate feature maps through 1×1 convolutional layers where x is the scaling factor. For example, the longest protein sequence length is 1600. After training the generator network, that is, the target generator, the lengths of the predicted protein sequence diagrams with low resolution are 100, 200, 400, and 800 respectively. Through the above method, the comprehensive capture of protein sequence features can be achieved.

[0043] Step S104: Based on the sets of predicted protein images corresponding to multiple vectors and the sets of actual protein images, perform machine learning on the initial discriminator network to obtain the target discriminator network and the discrimination result. The discrimination result is used to indicate the discrimination degree between the predicted protein sequence diagrams corresponding to multiple vectors and the corresponding actual protein sequence diagrams. The set of actual protein images includes multiple actual protein sequence diagrams, and the multiple actual protein sequence diagrams correspond to different resolutions.

[0044] Optionally, the initial discriminator network is the discriminator network before training. The discriminator network after training (i.e., the target discriminator network) needs to process protein images with multiple resolutions and accurately classify and discriminate images with different resolutions. Based on the feedback of the discriminator network, the generator network can first ensure consistency with the real protein sequence at low resolution, and then maintain consistency at high resolution, thereby improving the stability of the generator network. The discriminator network also has the ability to identify and distinguish images with multiple resolutions, and can first judge the consistency between the predicted protein sequence and the actual protein sequence from low resolution, and then judge the consistency at high resolution, reducing the difficulty of discrimination and improving the robustness and adaptability of the overall model.

[0045] In an alternative embodiment, before performing machine learning on the initial discriminator network based on the predicted protein image sets corresponding to multiple vectors and multiple actual protein image sets to obtain the target discriminator network and the discrimination result, the method further includes: obtaining multiple actual protein sequences; encoding the multiple actual protein sequences according to the amino acid types corresponding to the amino acids included in the multiple actual protein sequences to obtain the encoding results corresponding to the multiple actual protein sequences, wherein different types of amino acids correspond to different encoding identifiers; generating initial protein sequence maps corresponding to the multiple actual protein sequences based on the encoding results corresponding to the multiple actual protein sequences; and compressing the initial protein sequence maps corresponding to the multiple actual protein sequences according to multiple resolution rates to obtain multiple actual protein image sets, wherein the multiple actual protein sequences and the multiple actual protein image sets correspond to each other one by one.

[0046] Optionally, before performing machine learning, it is necessary to collect data, obtain multiple actual protein sequences, and after encoding each obtained actual protein sequence according to the amino acid type, obtain the initial protein sequence map corresponding to each actual protein sequence. Each initial protein sequence map is a protein sequence map with a predetermined resolution. Further, corresponding compression processing is performed on each initial protein sequence map according to multiple different resolutions to obtain protein sequence maps of each initial protein sequence map at multiple different resolutions, constituting the actual protein sequence set corresponding to each actual protein sequence.

[0047] Optionally, a ResNet architecture of a residual network used in computer vision can be adopted to encode the protein sequence, and each actual protein sequence is regarded as a highly abstract feature image to obtain the initial protein sequence map of each actual protein sequence. First, the quantization method of protein sequence data is described. Define an alphabet containing 21 elements to describe different types of amino acids (i.e., 20 basic amino acids and one letter "X" representing a space), and use integers between 1 and 21 to encode each type of amino acid. For each actual protein sequence S=(a1,a2,…,a n ), it is padded to reach the maximum length of 1600, that is, S=(a1,a2,…,a n ,a n+1 ,…,a 1600 ), where a i is the integer of the amino acid type at the i-th position, representing the residue at the i-th position. The protein sequence S is encoded using one-hot encoding to obtain a three-dimensional image of (21,1,1600) as the initial protein sequence map of each actual protein sequence, where S i,1,jIndicates whether the residue at the j-th position is of the i-th amino acid type (1 indicates yes, 0 indicates no). Figure 2 It is an optional flowchart for obtaining multi-resolution protein images according to an embodiment of the present invention, as Figure 2 shown. Taking the resolution reduction by half as an example, for a protein sequence image of (21, 1, 1600), by averaging the vector representations of adjacent residues, a protein sequence map of (21, 1, 800) was obtained as the initial protein sequence map. In Figure 2 the processing of training data, in order to obtain multi-scale features, not only the feature maps with the resolution reduced by half were obtained, but also the protein sequence images with the resolution reduced by one-fourth, one-eighth, and one-sixteenth were obtained. Specifically, the vector representations of four adjacent residues, eight adjacent residues, and sixteen adjacent residues were respectively averaged to achieve a comprehensive capture of features at different scales.

[0048] In an optional embodiment, obtaining multiple actual protein sequences includes: obtaining protein sequences of a first preset protein type from a protein sequence library to obtain a first number of actual protein sequences; determining, from the protein sequence library, protein sequences with a sequence identity greater than a preset threshold with respect to protein sequences of a second preset type to obtain a second number of actual protein sequences; and obtaining multiple actual protein sequences based on the first number of actual protein sequences and the second number of actual protein sequences.

[0049] Optionally, the protein sequences of the second preset protein type are protein sequences having a predetermined association relationship with the protein sequences of the first preset protein type. The protein sequences of the first preset protein type may be clustered regularly interspaced short palindromic repeat sequence CRISPR-associated protein 9 (Cas9) sequences, and the protein sequences of the second preset protein type may be class I receptor kinase-binding protein Iscb proteins. All Cas9 sequences can be screened on the protein UniProt website according to the gene name as the first number of actual protein sequences. Then, based on several classical Iscb proteins (Protein Data Bank PDB IDs: 7UTN, 8CTL, 8CSZ), sequence alignment searches in the protein library were performed using Blastp to screen out protein sequences with a sequence identity greater than the set threshold to obtain the second number of actual protein sequences. In this way, a large amount of sequence data of Cas9 and class Iscb proteins was obtained, that is, the first number of protein sequences and the second number of protein sequences were obtained.

[0050] Optionally, the first quantity of actual protein sequences and the second quantity of actual protein sequences can be directly used as the multiple actual protein sequences. It is also possible to perform screening processing on the first quantity of protein sequences and the second quantity of protein sequences based on the length of the protein sequences, and use the obtained protein sequences after screening as the multiple actual protein sequences for subsequent model training. Figure 3 is a schematic diagram of an optional protein sequence length distribution according to an embodiment of the present invention, as Figure 3 shown, which is the length distribution of the sequence dataset composed of multiple actual protein sequences. It can be seen that the length distribution of the protein sequences is quite extensive, ranging from 50 to 1750. Excluding extremely few sequences with a length exceeding 1600, a total of 8392 sequence data are finally obtained as multiple actual protein sequences.

[0051] Optionally, during the machine learning process, 8041 of the multiple actual protein sequences can be used for the training set, and 351 for the test set.

[0052] In an optional embodiment, based on the predicted protein image sets respectively corresponding to the multiple vectors and the multiple actual protein image sets, machine learning is performed on the initial discriminator network to obtain the target discriminator network and the discrimination result, including: using the predicted protein sequence maps in the predicted protein image sets respectively corresponding to the multiple vectors as the target predicted protein sequence maps, using the corresponding actual protein sequence maps as the target actual protein sequence maps, and obtaining the discrimination degree between the predicted protein sequence maps respectively corresponding to the multiple vectors and the corresponding actual protein sequence maps through the following method: based on the target predicted protein sequence map and the target actual protein sequence map, using the representation layer in the initial discriminator network to obtain the fourth feature map corresponding to the target predicted protein sequence map and the fifth feature map corresponding to the target actual protein sequence map; based on the fourth feature map and the fifth feature map, using the second residual basic unit to obtain the discrimination degree between the target predicted protein sequence map and the target actual protein sequence map, where the second residual basic unit is any one of the multiple residual basic units in the initial discriminator network; and obtaining the discrimination result based on the discrimination degrees between the predicted protein sequence maps respectively corresponding to the multiple vectors and the corresponding actual protein sequence maps.

[0053] Optionally, the input of the initial generator network is the sum of the predicted protein sequence maps and the actual protein sequence maps with multiple resolutions, and the corresponding size is Its size is where \(x=(0, 1, 2, 3, 4)\) is the resolution scaling factor. The initial discriminator network includes a representation layer and multiple residual basic units, and the multiple residual basic units correspond to different resolutions. To more effectively model the correlation relationships between different amino acids, the initial discriminator network first processes the predicted protein sequence map and the actual protein sequence at different resolutions through a representation layer to obtain the fourth feature map corresponding to the predicted protein sequence map and the fifth feature map corresponding to the actual protein sequence map. The representation layer is designed to learn a vector of dimension \((1, \text{emb_dim})\) for each type of amino acid, where \(\text{emb_dim}\) represents the embedding representation of the amino acid, so as to achieve the capture and learning of amino acid features. Subsequently, the obtained fourth feature map and fifth feature map are processed by the corresponding residual basic units to obtain the discrimination degree between the predicted protein sequence map and the actual protein sequence map at the corresponding resolution. According to the discrimination degrees between each predicted protein sequence map and the corresponding actual protein sequence map, the discrimination result is determined, where the discrimination result can be used to indicate the deviation degree between the predicted protein sequence map output by the trained generator network and the actual protein sequence map. Through the above method, the initial discriminator network can more fully understand the complex correlations between amino acids in the protein sequence and provide a richer and more meaningful feature representation for the subsequent discrimination process.

[0054] In an alternative embodiment, based on the fourth feature map and the fifth feature map, the second residual basic unit is used to obtain the discrimination degree between the target predicted protein sequence map and the target actual protein sequence map, including: using the filters in the second residual basic unit to perform feature extraction processing on the fourth feature map and the fifth feature map respectively to obtain the sixth feature map and the seventh feature map; splicing the fourth feature map and the sixth feature map to obtain the eighth feature map; and splicing the fifth feature map and the seventh feature map to obtain the ninth feature map; based on the discrimination degree between the eighth feature map and the ninth feature map, the discrimination degree between the target predicted protein sequence map and the target actual protein sequence map is determined.

[0055] Optionally, the target predicted protein sequence map is any one of the multiple predicted protein sequence maps, and the target actual protein sequence map is the actual protein sequence map corresponding to the target predicted protein sequence map among the multiple actual protein sequence maps. The vector dimensions corresponding to the fourth feature map and the fifth feature map are three-dimensional. For example, the vector representation forms corresponding to the fourth feature map and the fifth feature map are as follows: (emb_dim, 1, L). Each residual basic unit in this initial discriminator network includes a filter, and this filter can be a convolutional layer with a size of 3 and a stride of 2. Specifically, the fourth feature map corresponding to the target predicted protein sequence map and the fifth feature map corresponding to the target actual protein sequence map are respectively input into the filters in the residual basic units at the corresponding resolutions for feature extraction, and the sixth feature map and the seventh feature map are respectively obtained. The vector representation forms corresponding to the sixth feature map and the seventh feature map are as follows: "g" represents the generator, indicating the number of channels input to the first residual basic unit of the generator; the newly obtained feature maps are respectively concatenated with the corresponding filter inputs to obtain new feature maps. That is, the fourth feature map and the sixth feature map are concatenated to obtain the eighth feature map; the fifth feature map and the seventh feature map are concatenated to obtain the ninth feature map. The vector representation forms corresponding to the eighth feature map and the ninth feature map are as follows: Feature images at other resolutions, including where "emb_dim" represents the dimension of amino acid representation; they are also input into the initial discriminator network in the same way and corresponding processing is performed.

[0056] Step S106, when the discrimination result meets the predetermined discrimination condition, a protein sequence generation model is obtained based on the target generator network and the discriminator network.

[0057] Optionally, a generative adversarial network is used as the basic model. The generator network is used to generate protein images, while the discriminator network is responsible for distinguishing the generated images from the real images. Through this adversarial learning mechanism, the generator network continuously optimizes the generated images to make them closer to the real data distribution, while the discriminator network continuously improves its own discrimination ability to more accurately distinguish real images and generated images. This adversarial learning process can effectively improve the generation ability and robustness of the model, thereby achieving more accurate and higher-quality protein image generation.

[0058] In the process of machine learning, the initial generator network and the initial discriminator network are continuously trained and optimized. The corresponding loss function can adopt non-saturating loss and R1 regularization to ensure the stability of training. When the discrimination result output by the trained discriminator network meets the predetermined discrimination condition, the trained target generator network and the target discriminator network are output to form a protein sequence generation model; based on this protein sequence generation model, batch generation of protein sequences can be achieved. To further improve stability, spectral normalization is implemented in all layers of the generator network and the discriminator network.

[0059] Step S108, generating multiple protein sequences using the protein sequence generation model.

[0060] Optionally, after training the protein sequence model, based on the protein sequence model, using a two-dimensional vector as the input of the model, multiple protein sequences can be automatically generated.

[0061] In an optional embodiment, after generating multiple protein sequences using the protein sequence generation model, the method further includes: determining the vectors corresponding to the multiple protein sequences respectively, and the sequence lengths corresponding to the multiple protein sequences respectively; performing machine learning based on the vectors corresponding to the multiple protein sequences respectively, and the sequence lengths corresponding to the multiple protein sequences respectively, to obtain a protein length prediction model.

[0062] Optionally, the protein sequence generation model can learn the correspondence between any input vector and a protein sequence. For any known vector, a protein sequence corresponding to the vector can be generated accordingly, and the length of the corresponding generated protein sequence is also known. Based on the multiple protein sequences generated by the protein sequence generation model, and the vectors and sequence lengths corresponding to the multiple protein sequences respectively, using the vector as the input and the corresponding sequence length as the output, a protein length prediction model can be constructed through machine learning. Among them, in the process of machine learning, the protein length prediction model can learn the correspondence between the corresponding vector of the protein sequence and the protein sequence length. Based on this protein length prediction model, accurate prediction of the protein sequence length can be achieved, which is used for denoising of protein sequences and compression of sequence lengths.

[0063] In an alternative embodiment, after performing machine learning based on the vectors corresponding to multiple protein sequences and the sequence lengths corresponding to multiple protein sequences respectively to obtain a protein length prediction model, the method further includes: obtaining a target protein sequence, the target vector corresponding to the target protein sequence, and the actual sequence length of the target protein sequence; based on the target vector, using the protein length prediction model to obtain the predicted sequence length of the target protein sequence; in the case where the predicted sequence length is less than the actual sequence length, optimizing the target protein sequence to obtain an optimized protein sequence, wherein the sequence length of the optimized protein sequence is the predicted sequence length.

[0064] Optionally, after obtaining the protein sequence length prediction model, for any protein sequence, such as the target protein sequence, when the target protein sequence and the corresponding target vector are known, the target vector can be used as an input, and the protein sequence length prediction model is used to predict the predicted sequence length of the target protein sequence. By comparing the predicted sequence length of the target protein sequence with the actual sequence length, it is determined whether the length of the target protein sequence needs to be optimized; if the predicted sequence length is less than the actual sequence length, the length of the target protein sequence needs to be optimized to reduce the length of the target protein sequence. Specifically, in the case where the predicted sequence length is less than the actual sequence length, the gradient descent algorithm is used to optimize the target protein sequence to obtain an optimized protein sequence. If the predicted sequence length is greater than or equal to the actual sequence length, it indicates that the sequence length of the current target protein is already close to the optimal value and no processing is required. Through the above method, the purpose of obtaining high-quality protein sequences can be achieved.

[0065] Through the above steps S102 to step S108, the purpose of training a protein sequence generation model that integrates multi-resolution protein sequence features based on a generator network and a discriminator network to efficiently generate protein sequences can be achieved, so as to realize the technical effect of improving the accuracy and efficiency of protein sequence generation, and further solve the technical problem of poor protein sequence generation effect existing in the protein sequence generation method in the related art.

[0066] In the embodiments of the present invention, the generator network not only generates a single high-resolution (i.e., predetermined resolution) protein image, but also generates protein images with multiple resolutions, including feature maps with resolutions reduced to one-half, one-quarter, one-eighth, and one-sixteenth. At the same time, the discriminator network needs to process protein images with multiple resolutions and accurately classify and discriminate images with different resolutions. In this framework of multi-resolution generation, the initial generator network not only focuses on the generation of protein images at a single resolution, but also can process protein sequence information at different scales. Based on the feedback of the discriminator network, the generator network can first ensure that the generated protein sequence is consistent with the real protein sequence at a low resolution, and then remain consistent at a high resolution, thereby improving the stability of the generator network. The discriminator network also has the ability to discriminate and distinguish multi-resolution images, and can first judge the consistency between the predicted protein sequence and the actual protein sequence from the low resolution, and then judge the consistency at the high resolution, reducing the difficulty of discrimination and improving the robustness and adaptability of the overall model at the same time.

[0067] Based on the above embodiments and optional embodiments, the present invention proposes an optional implementation manner. Figure 4 It is an optional protein sequence processing flowchart according to the embodiments of the present invention, as Figure 4 shown. This method is based on the generator network and the discriminator network to construct a protein sequence generation model that integrates protein sequence features with multiple resolutions.

[0068] Specifically, it includes:

[0069] Step S1, on the UniProt website, all Cas9 sequences (Cas9 amino acid sequences) were screened according to the gene name. Then, based on several classical Iscb proteins (Protein Data Bank PDB numbers: 7UTN, 8CTL, 8CSZ), sequence alignment searches in the protein library were performed using the protein sequence alignment tool Blastp, and protein sequences with sequence identity greater than the set threshold were screened out. In this way, a large amount of sequence data of Cas9 and Iscb-like proteins was obtained. As Figure 2 shown in the schematic diagram of the protein sequence length distribution, it shows the length distribution of the sequence dataset composed of multiple actual protein sequences. It can be seen that the length distribution of the sequences is quite extensive, ranging from 50 to 1750. Excluding extremely few sequences with lengths exceeding 1600, a total of 8392 sequence data were finally obtained as multiple actual protein sequences. Among them, 8041 of the multiple actual protein sequences can be used for the training set, and 351 for the test set.

[0070] Step S2, use the ResNet (residual network) architecture in computer vision to encode the protein sequence, regard each actual protein sequence as a highly abstract feature image, and obtain the initial protein sequence map of each actual protein sequence. First, describe the quantization method of protein sequence data. Define an alphabet containing 21 elements to describe different types of amino acids (i.e., 20 basic amino acids and one letter "X" representing a space), and use integers between 1 and 21 to encode each type of amino acid. For each actual protein sequence S=(a1,a2,…,a n ), it is padded to reach the maximum length of 1600, that is, S=(a1,a2,…,a n ,a n+1 ,…,a 1600 ), where a i is the integer of the amino acid type at the i-th position, representing the residue at the i-th position. Use one-hot encoding for the protein sequence S to obtain a three-dimensional image of (21,1,1600) as the initial protein sequence map of each actual protein sequence, where S i,1,j indicates whether the residue at the j-th position is the i-th amino acid type (1 means yes, 0 means no). To achieve accurate generation and discrimination of protein sequences, adopt a multi-resolution processing and generation strategy for protein sequence images.

[0071] Step S3, taking the resolution reduction by half as an example, for the protein sequence image of (21,1,1600), by averaging the vector representations of adjacent residues, a protein sequence map of (21,1,800) is obtained as the initial protein sequence map. In Figure 4 the processing of training data in the upper right corner, to obtain multi-scale features, not only obtain the feature map with the resolution reduced by half, but also obtain the protein sequence images with the resolutions reduced by one-fourth, one-eighth, and one-sixteenth. Specifically, the vector representations of adjacent four residues, eight residues, and sixteen residues are respectively averaged to achieve comprehensive capture of different-scale features.

[0072] Step S4, the input of the initial generator network is a (128,1) multi-vector, and the multi-vector is sampled from a random distribution with a mean of 0 and a standard deviation of 1. The initial generator network first converts and reshapes the (128,1) vector through a multi-layer perceptron into a three-dimensional vector, where c and w respectively represent the number of channels and the vector length. Then, it is processed through multiple residual basic units, and each residual basic unit includes the following steps: taking the input as an example, use a transposed convolutional layer with a stride of 2 to magnify the vector width to Then, a 1×1 convolutional layer is used to process the generated feature map to capture higher-level features; then, another transposed convolutional layer with a stride of 2 is used to expand the width of the feature map, and finally, a residual connection is made with the feature map obtained in the previous step to get the output The feature map output by the last residual basic unit is passed through a 1×1 convolutional layer to obtain an output of (21, 1, L). To capture multi-scale features more comprehensively, the output of the generator network not only includes the (21, 1, L) of the last layer, but also includes the low-resolution protein sequence maps with dimensions obtained by passing the intermediate feature maps through 1×1 convolutional layers For example, the longest protein sequence length is 1600, and the lengths of the low-resolution protein sequence maps generated by the initial generator network are 100, 200, 400, and 800 respectively

[0073] Step S5, the input of the initial discriminator network is the sum of the multi-resolution predicted protein sequence maps and the actual protein sequence map, and its size is where x = (0, 1, 2, 3, 4) is the resolution scaling factor. To model the association relationships between different amino acids more effectively, the initial discriminator network first processes the protein sequence maps of different resolutions through a representation layer. This representation layer is designed to learn a (1, emb_dim)-dimensional vector for each type of amino acid, where emb_dim represents the embedding representation of the amino acid, so as to capture and learn the amino acid features. After this processing, the initial discriminator network can more fully understand the complex associations between the amino acids in the protein sequence and provide a richer and more meaningful feature representation for the subsequent discrimination process. Subsequently, the characterized three-dimensional image is processed through multiple residual basic units. Each residual basic unit includes the following steps: taking the input of (emb_dim, 1, L) as an example, a convolutional layer with a filter size of 3 and a stride of 2 is used to extract features from it to obtain feature images with several channels. Then, the feature images corresponding to the resolution after passing through the representation layer are concatenated with this feature image to obtain a new feature map representation. The feature images of other resolutions, including are also input into the initial discriminator network in the same way

[0074] ​Step S6, during the machine learning process, the initial generator network and the initial discriminator network are continuously trained and optimized. The corresponding loss function can adopt non-saturating loss and R1 regularization to ensure the stability of training. When the discrimination result output by the trained discriminator network meets the predetermined discrimination condition, the trained target generator network and target discriminator network are output to form a protein sequence generation model; based on this protein sequence generation model, batch generation of protein sequences can be achieved. To further improve stability, spectral normalization is implemented in all layers of the generator network and the discriminator network.

[0075] Figure 5 is an optional model effect diagram according to an embodiment of the present invention, Figure 5 showing the change in consistency between the protein sequences generated during the continuous training of the protein sequence generation model and the training set and the validation set. In Figure 5 , the solid line and the dashed line respectively represent the training set and the validation set. Using a single 4090, about 9 days are required for 500,000 training steps. Every 600 training steps, the model generates 64 protein sequences. Each generated protein sequence is compared with the sequences in the training set and the validation set using the Blastp program, and their sequence consistency is calculated. From Figure 5 , it can be observed that as the training progresses, the consistency between the generated sequences and the training set and the validation set gradually increases and finally stabilizes. In the stable stage, the average consistency between the generated sequences and the training set is about 70%, and the consistency with the validation set is about 60%. This indicates that the model training is stable and no overfitting has occurred.

[0076] Figure 6 is an optional protein sequence distribution comparison diagram according to an embodiment of the present invention. As Figure 6 shown, all ProteinBert-based sequence characterizations are converted into two-dimensional characterizations using the t-distributed stochastic neighbor embedding (t-SNE) algorithm for visual analysis. To ensure that the t-SNE algorithm has the same weight for each sequence category, sequences with a quantity almost equal to that of the training set and the validation set are generated, totaling 8,600. Figure 6 The left figure (a) in Figure 6The right middle figure (b) shows the relationship between the sequences generated by the protein sequence generation model proposed in the embodiment of the present invention and the training set and the validation set. Compared with the left figure (a), the sequences generated by the model are indistinguishable from the training set and the validation set. ProteinBert considers these generated sequences to be homologous sequences.

[0077] Step S7, further determine whether the protein sequence generation model provided in the embodiment of the present invention contains semantic features related to sequence length in the latent space. First, 9000 points were sampled on random noise, and 9000 corresponding protein sequences were generated. Subsequently, the length of each generated sequence was calculated and used as a label. This data set was divided into a test set, accounting for 20% of the total data volume. Then, the training set was used to fit the parameters of the support vector machine model to obtain a protein length prediction model to achieve the function of predicting the sequence length based on the latent vector.

[0078] The simple linear prediction model of support vector machine was selected because the goal of the embodiment of the present invention is to explore whether the latent space contains semantic features, rather than improving the prediction accuracy of the latent space through a complex model. Too complex models may more reflect the performance of the prediction model rather than the semantic features of the latent space. To quantify the prediction error, the R 2 value was used to measure the fitting degree between the predicted value and the true value. Figure 7 is a schematic diagram of the comparison between the predicted length and the true length of a sequence according to an optional embodiment of the present invention. As Figure 7 shown, the R 2 value of the protein length prediction model on the test set reached 0.598, showing an obvious linear correlation between the predicted value and the true value. This proves that the protein sequence latent space defined by the protein sequence generation model not only maps to the protein sequence spaces of Cas9 and Iscb, but also has powerful length semantic features.

[0079] The protein length prediction model trained in the embodiment of the present invention can be used to effectively reduce the length of Cas9 protein and Iscb protein. The currently concerned Cas9 and Iscb protein sequences were selected from the genomic Blast database and the Uniprot database. This process first includes determining the positions of these two target protein sequences in the latent space. 8000 sequences were generated using the protein sequence generation model, and the target sequences were compared with these sequences to find the generated sequence with the highest consistency with the target sequence. It is assumed that the latent vector corresponding to the generated sequence with the highest consistency is the latent vector of the target sequence.

[0080] Step S8: Using the hidden vector of the target protein sequence as the starting point, apply the protein sequence length prediction model based on the hidden vector. This model can effectively calculate the direction of the change in the hidden vector when the sequence length rapidly decreases. Based on these calculation results, update the hidden vector of the target sequence and repeat this process until the obtained protein sequence length is reduced to the set expected range. During this process, in order to ensure that the protein sequence represented by the updated hidden vector is regarded as real and reliable in the protein sequence generation model, the hidden vector positions with relatively low sampling probabilities in the generation model's hidden space are deliberately excluded.

[0081] In addition, Figure 8 is a schematic diagram of the consistency comparison after the length reduction of an optional Cas9 protein sequence according to an embodiment of the present invention. Figure 9 is a schematic diagram of the consistency comparison after the length reduction of an optional IscB protein sequence according to an embodiment of the present invention. As Figure 8 and Figure 9 shown in part (a) of Figure 8 and Figure 9 it can be seen that due to the significant length semantic features of the hidden vector, the reduction of the sequence length can be effectively achieved through the gradient backpropagation of the protein sequence length prediction model based on the hidden vector. At the same time,

[0082] The embodiments of the present invention can at least achieve at least one of the following effects: 1) The embodiments of the present invention propose a protein sequence generation model constructed based on a multi-scale generative adversarial network for learning the generation pattern of protein sequences. Compared with the existing sequence generation models, the embodiments of the present invention can more effectively capture the long-range correlation of sequences with a length exceeding 1024, providing researchers with a more accurate and comprehensive protein sequence generation tool. 2) For the miniaturization application requirements of Cas9 protein, the proposed multi-scale generation model has been experimentally proven to be able to effectively train the generation model, enabling it to generate sequences highly similar to the real Cas9 and Iscb proteins. In addition, the latent space encoded by the protein sequence generation model in the embodiments of the present invention shows significant ability in sequence length prediction. When using the support vector machine (SVM) model for prediction and analysis, the coefficient of determination (R^2) of the model reaches 0.598, indicating the effectiveness and accuracy of the protein sequence generation model proposed in the embodiments of the present invention in predicting the length of protein sequences. 3) Based on the protein sequence generation model for generating Cas9-like and IscB-like proteins, combined with the protein sequence length prediction model, the method of backpropagation of gradients is used to achieve the reduction of sequence length. The experimental results show that even after the significant reduction of sequence length, the obtained protein sequences can still maintain a high degree of consistency with the real Cas9 and IscB protein sequences. It is confirmed that the embodiments of the present invention not only demonstrate an efficient protein sequence design method but also provide a new protein length optimization technology, which is crucial for various biotechnological applications.

[0083] In this embodiment, a protein sequence processing device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the terms "module" and "device" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation in hardware, or a combination of software and hardware, is also possible and contemplated.

[0084] According to an embodiment of the present invention, an apparatus embodiment for implementing the above protein sequence processing method is also provided. Figure 10 It is a schematic structural diagram of a protein sequence processing device according to an embodiment of the present invention. As Figure 10 shown, the above protein sequence processing device includes: a first machine learning module 1010, a second machine learning module 1012, an acquisition module 1014, and a generation module 1016, where:

[0085] The first machine learning module 1010 is configured to perform machine learning on an initial generator network based on multiple vectors to obtain a target generator network, and a set of predicted protein images corresponding to each of the multiple vectors. The set of predicted protein images includes multiple predicted protein sequence maps, and the multiple predicted protein sequence maps correspond to different resolutions;

[0086] The second machine learning module 1012 is connected to the first machine learning module 1010 and is configured to perform machine learning on an initial discriminator network based on the set of predicted protein images corresponding to each of the multiple vectors and a set of actual protein images to obtain a target discriminator network and a discrimination result. The discrimination result is used to indicate the discrimination degree between the predicted protein sequence maps corresponding to each of the multiple vectors and the corresponding actual protein sequence maps. The set of actual protein images includes multiple actual protein sequence maps, and the multiple actual protein sequence maps correspond to different resolutions;

[0087] The obtaining module 1014 is connected to the second machine learning module 1012 and is configured to obtain a protein sequence generation model based on the target generator network and the discriminator network when the discrimination result meets a predetermined discrimination condition;

[0088] The generating module 1016 is connected to the obtaining module 1014 and is configured to generate multiple protein sequences by using the protein sequence generation model.

[0089] In an embodiment of the present invention, by setting a first machine learning module 1010, which is used to perform machine learning on an initial generator network based on multiple vectors to obtain a target generator network, and a set of predicted protein images corresponding to the multiple vectors respectively, wherein the set of predicted protein images includes multiple predicted protein sequence maps, and the multiple predicted protein sequence maps correspond to different resolutions; a second machine learning module 1012, connected to the first machine learning module 1010, which is used to perform machine learning on an initial discriminator network based on the set of predicted protein images corresponding to the multiple vectors respectively and a set of actual protein images to obtain a target discriminator network and a discrimination result, wherein the discrimination result is used to indicate the discrimination degree between the predicted protein sequence maps corresponding to the multiple vectors respectively and the corresponding actual protein sequence maps, the set of actual protein images includes multiple actual protein sequence maps, and the multiple actual protein sequence maps correspond to different resolutions; an acquisition module 1014, connected to the second machine learning module 1012, which is used to obtain a protein sequence generation model based on the target generator network and the discriminator network when the discrimination result meets a predetermined discrimination condition; a generation module 1016, connected to the acquisition module 1014, which is used to generate multiple protein sequences by using the protein sequence generation model, achieving the purpose of training a protein sequence generation model that integrates protein sequence features of multiple resolutions based on a generator network and a discriminator network for efficient generation of protein sequences, thereby realizing the technical effect of improving the accuracy and efficiency of protein sequence generation, and further solving the technical problem of poor protein sequence generation effect existing in the protein sequence generation method in the related art.

[0090] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following manner: the above-mentioned various modules can be located in the same processor; or, the above-mentioned various modules are located in different processors in any combination.

[0091] It should be noted here that the above-mentioned first machine learning module 1010, second machine learning module 1012, acquisition module 1014, and generation module 1016 correspond to steps S102 to step S108 in the embodiment. The examples and application scenarios implemented by the above-mentioned modules and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned embodiment. It should be noted that the above-mentioned modules, as part of the device, can run in a computer terminal.

[0092] It should be noted that the optional or preferred implementation manners of this embodiment can refer to the relevant descriptions in the embodiment, and will not be repeated here.

[0093] The above protein sequence processing device may further include a processor and a memory. The first machine learning module 1010, the second machine learning module 1012, the acquisition module 1014, the generation module 1016, etc. are all stored in the memory as program modules, and the processor executes the above program modules stored in the memory to implement corresponding functions.

[0094] The processor contains a kernel, and the kernel retrieves the corresponding program modules from the memory. One or more kernels can be set. The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one memory chip.

[0095] According to an embodiment of the present application, an embodiment of a non-volatile storage medium is also provided. Optionally, in this embodiment, the non-volatile storage medium includes a stored program, and when the above program runs, it controls the device where the non-volatile storage medium is located to execute any one of the above protein sequence processing methods.

[0096] Optionally, in this embodiment, the non-volatile storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group, and the non-volatile storage medium includes a stored program.

[0097] Optionally, when the program runs, it controls the device where the non-volatile storage medium is located to execute the following functions: performing machine learning on an initial generator network based on multiple vectors to obtain a target generator network, and a set of predicted protein images corresponding to the multiple vectors respectively, where the set of predicted protein images includes multiple predicted protein sequence diagrams, and the multiple predicted protein sequence diagrams correspond to different resolutions; performing machine learning on an initial discriminator network based on the set of predicted protein images corresponding to the multiple vectors respectively, and a set of actual protein images to obtain a target discriminator network and a discrimination result, where the discrimination result is used to indicate the discrimination degree between the predicted protein sequence diagrams corresponding to the multiple vectors respectively and the corresponding actual protein sequence diagrams, the set of actual protein images includes multiple actual protein sequence diagrams, and the multiple actual protein sequence diagrams correspond to different resolutions; in the case where the discrimination result meets a predetermined discrimination condition, obtaining a protein sequence generation model based on the target generator network and the discriminator network; and generating multiple protein sequences using the protein sequence generation model.

[0098] According to an embodiment of the present application, an embodiment of a processor is also provided. Optionally, in this embodiment, the above processor is used to run a program, and when the above program runs, it executes any one of the above protein sequence processing methods.

[0099] According to an embodiment of the present application, an embodiment of a computer program product is further provided. When executed on a data processing device, it is adapted to execute a program for initializing the steps of the protein sequence processing method of any one of the above.

[0100] Optionally, when the above computer program product is executed on a data processing device, it is adapted to execute a program with the following method steps: performing machine learning on an initial generator network based on multiple vectors to obtain a target generator network and a set of predicted protein images respectively corresponding to the multiple vectors, wherein the set of predicted protein images includes multiple predicted protein sequence diagrams, and the multiple predicted protein sequence diagrams correspond to different resolutions; performing machine learning on an initial discriminator network based on the set of predicted protein images respectively corresponding to the multiple vectors and a set of actual protein images to obtain a target discriminator network and a discrimination result, wherein the discrimination result is used to indicate the discrimination degree between the predicted protein sequence diagrams respectively corresponding to the multiple vectors and the corresponding actual protein sequence diagrams, the set of actual protein images includes multiple actual protein sequence diagrams, and the multiple actual protein sequence diagrams correspond to different resolutions; in the case where the discrimination result meets a predetermined discrimination condition, obtaining a protein sequence generation model based on the target generator network and the discriminator network; and generating multiple protein sequences by using the protein sequence generation model.

[0101] An embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, the following steps are implemented: performing machine learning on an initial generator network based on multiple vectors to obtain a target generator network and a set of predicted protein images respectively corresponding to the multiple vectors, wherein the set of predicted protein images includes multiple predicted protein sequence diagrams, and the multiple predicted protein sequence diagrams correspond to different resolutions; performing machine learning on an initial discriminator network based on the set of predicted protein images respectively corresponding to the multiple vectors and a set of actual protein images to obtain a target discriminator network and a discrimination result, wherein the discrimination result is used to indicate the discrimination degree between the predicted protein sequence diagrams respectively corresponding to the multiple vectors and the corresponding actual protein sequence diagrams, the set of actual protein images includes multiple actual protein sequence diagrams, and the multiple actual protein sequence diagrams correspond to different resolutions; in the case where the discrimination result meets a predetermined discrimination condition, obtaining a protein sequence generation model based on the target generator network and the discriminator network; and generating multiple protein sequences by using the protein sequence generation model.

[0102] The order of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments.

[0103] In the above embodiments of the present invention, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0104] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the above-mentioned module division can be a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of modules or modules can be in an electrical or other form.

[0105] The modules described above as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0106] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0107] If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable non-volatile storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned non-volatile storage media include: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical disks, etc., which can store program codes.

[0108] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A protein sequence processing method, characterized in that: include: Based on the multiple vectors, machine learning is performed on the initial generator network to obtain a target generator network and a set of predicted protein images corresponding to the multiple vectors, wherein the set of predicted protein images includes multiple predicted protein sequence graphs corresponding to different resolutions; Based on the predicted protein image sets corresponding to the multiple vectors and multiple actual protein image sets, performing machine learning on an initial discriminator network to obtain a target discriminator network and a discrimination result, wherein the discrimination result is used to indicate the degree of distinction between the predicted protein sequence graphs corresponding to the multiple vectors and the corresponding actual protein sequence graphs, the actual protein image set including multiple actual protein sequence graphs corresponding to different resolutions; When the discrimination result satisfies a predetermined discrimination condition, obtaining a protein sequence generation model based on the target generator network and the discriminator network; A plurality of protein sequences are generated using the protein sequence generation model.

2. The method according to claim 1, characterized in that Before performing machine learning on the initial discriminator network based on the predicted protein image sets corresponding to the multiple vectors and the multiple actual protein image sets to obtain the target discriminator network and the discrimination results, the method further includes: Obtain multiple actual protein sequences; performing encoding processing on the multiple actual protein sequences according to the amino acid types corresponding to the amino acids included in the multiple actual protein sequences, to obtain encoding results corresponding to the multiple actual protein sequences, wherein different types of amino acids correspond to different encoding identifiers; generating initial protein sequence graphs corresponding to the plurality of actual protein sequences based on the encoding results corresponding to the plurality of actual protein sequences; The initial protein sequence graphs corresponding to the multiple actual protein sequences are compressed according to multiple resolutions to obtain the multiple actual protein image sets, wherein the multiple actual protein sequences correspond to the multiple actual protein image sets one-to-one.

3. The method according to claim 2, characterized in that The obtaining of multiple actual protein sequences comprises: Obtaining protein sequences of a first preset protein type from a protein sequence library to obtain a first number of actual protein sequences; Determining, from the protein sequence library, protein sequences whose sequence identity with protein sequences of a second preset type is greater than a preset threshold, to obtain a second number of actual protein sequences; The plurality of actual protein sequences are obtained based on the first number of actual protein sequences and the second number of actual protein sequences.

4. The method according to claim 1, wherein The method of performing machine learning on the initial generator network based on the multiple vectors to obtain a target generator network and a set of predicted protein images corresponding to the multiple vectors, including: Based on the multiple vectors, using a multilayer perceptron in the initial generator network, obtain three-dimensional vectors corresponding to the multiple vectors, wherein the multiple vectors are respectively two-dimensional vectors; Based on the three-dimensional vectors corresponding to the multiple vectors, the multiple residual basic units in the initial generator network are used to obtain the predicted protein image sets corresponding to the multiple vectors, wherein the multiple residual basic units in the initial generator network correspond to different resolutions.

5. The method according to claim 4, characterized in that The method of obtaining the predicted protein image sets corresponding to the multiple vectors respectively based on the three-dimensional vectors corresponding to the multiple vectors by using multiple residual basic units in the initial generator network comprises: Based on the three-dimensional vectors corresponding to the multiple vectors, any predicted protein sequence graph corresponding to any one of the multiple vectors is obtained in the following manner: Based on the three-dimensional vector of the arbitrary one vector, a first transposed convolution layer in a first residual basic unit is used to obtain a first feature map, wherein the first residual basic unit is any one of the plurality of residual basic units in the initial generator network; Based on the first feature map, using a convolutional layer of a predetermined first dimension in the first residual basic unit to obtain a second feature map; Based on the second feature map, using the second transposed convolution layer in the first residual basic unit to obtain a third feature map; Perform a residual connection on the second feature map and the third feature map to obtain the arbitrary predicted protein sequence map.

6. The method according to claim 1, characterized in that The method of performing machine learning on the initial discriminator network based on the predicted protein image sets corresponding to the multiple vectors and the multiple actual protein image sets to obtain a target discriminator network and a discrimination result includes: The predicted protein sequence graphs in the predicted protein image set corresponding to the multiple vectors are respectively used as target predicted protein sequence graphs, and the corresponding actual protein sequence graphs are used as target actual protein sequence graphs. The discrimination between the predicted protein sequence graphs corresponding to the multiple vectors and the corresponding actual protein sequence graphs is obtained in the following manner: Based on the target predicted protein sequence graph and the target actual protein sequence graph, using the representation layer in the initial discriminator network, a fourth feature graph corresponding to the target predicted protein sequence graph and a fifth feature graph corresponding to the target actual protein sequence graph are obtained; Based on the fourth feature map and the fifth feature map, a second residual basic unit is used to obtain the discrimination between the target predicted protein sequence map and the target actual protein sequence map, wherein the second residual basic unit is any one of the plurality of residual basic units in the initial discriminator network; The discrimination result is obtained based on the degree of distinction between the predicted protein sequence graphs and the corresponding actual protein sequence graphs corresponding to the multiple vectors.

7. The method according to claim 6, characterized in that The method of obtaining the discrimination between the target predicted protein sequence graph and the target actual protein sequence graph based on the fourth feature graph and the fifth feature graph and using a second residual basic unit includes: Using the filter in the second residual basic unit to perform feature extraction processing on the fourth feature map and the fifth feature map, respectively, to obtain a sixth feature map and a seventh feature map; The fourth feature map and the sixth feature map are spliced together to obtain an eighth feature map; and the fifth feature map and the seventh feature map are spliced together to obtain a ninth feature map; Based on the degree of distinction between the eighth feature graph and the ninth feature graph, the degree of distinction between the target predicted protein sequence graph and the target actual protein sequence graph is determined.

8. The method according to any one of claims 1 to 7, characterized in that After generating a plurality of protein sequences using the protein sequence generation model, the method further includes: Determining vectors corresponding to the plurality of protein sequences, and sequence lengths corresponding to the plurality of protein sequences; Based on the vectors corresponding to the multiple protein sequences and the sequence lengths corresponding to the multiple protein sequences, machine learning is performed to obtain a protein length prediction model.

9. The method according to claim 8, characterized in that After performing machine learning based on the vectors corresponding to the multiple protein sequences and the sequence lengths corresponding to the multiple protein sequences to obtain a protein length prediction model, the method further includes: Obtaining a target protein sequence, a target vector corresponding to the target protein sequence, and an actual sequence length of the target protein sequence; Based on the target vector, using the protein length prediction model, obtaining a predicted sequence length of the target protein sequence; In a case where the predicted sequence length is less than the actual sequence length, the target protein sequence is optimized to obtain an optimized protein sequence, wherein the sequence length of the optimized protein sequence is the predicted sequence length.

10. A protein sequence processing device, characterized in that: include: a first machine learning module, configured to perform machine learning on an initial generator network based on a plurality of vectors to obtain a target generator network and a set of predicted protein images corresponding to the plurality of vectors, wherein the set of predicted protein images includes a plurality of predicted protein sequence graphs corresponding to different resolutions; a second machine learning module, configured to perform machine learning on an initial discriminator network based on the predicted protein image sets corresponding to the multiple vectors and multiple actual protein image sets, to obtain a target discriminator network and a discrimination result, wherein the discrimination result is used to indicate the degree of discrimination between the predicted protein sequence graphs corresponding to the multiple vectors and the corresponding actual protein sequence graphs, the actual protein image set including multiple actual protein sequence graphs corresponding to different resolutions; an acquisition module, configured to obtain a protein sequence generation model based on the target generator network and the discriminator network when the discrimination result satisfies a predetermined discrimination condition; A generation module is used to generate multiple protein sequences using the protein sequence generation model.

11. An electronic device, characterized in that: The method comprises one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the protein sequence processing method according to any one of claims 1 to 9.