A method and system for amino acid sequence generation and screening based on a diffusion denoising probability model.

By using a diffusion denoising probability model and a deep learning framework, the problem of single design results in amino acid sequence generation and screening methods was solved, realizing the diversity and efficient generation of viral amino acid sequences, and improving the efficiency and adaptability of viral amino acid sequence design.

CN118136107BActive Publication Date: 2025-10-31ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211533061.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-01
Publication Date
2025-10-31
Estimated Expiration
2042-12-01

AI Technical Summary

Technical Problem

Existing methods for generating and screening amino acid sequences are limited by their singular design results and excessive sensitivity to the details of the main chain structure design. This restricts the diversity and variability of the main chain structure design and makes it difficult to generate highly active amino acid sequences, especially in the design of viral amino acid sequences, where there are technical bottlenecks.

Method used

A diffusion denoising probability model is adopted, combined with word vector representation and maximum independent variable point set method. A deep learning generation framework is used to learn the distribution of amino acid sequence information to generate viral amino acid sequences with high specificity and high production fitness. Neural networks are used to perform forward diffusion and backward diffusion processes, and denoising and maximum likelihood reduction are performed to achieve amino acid sequence generation.

Benefits of technology

This improved the diversity and variability of amino acid sequence design, enhanced the efficiency and speed of viral amino acid sequence design, reduced experimental costs, and generated viral amino acid sequences with high specificity and high production adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118136107B_ABST
    Figure CN118136107B_ABST
Patent Text Reader

Abstract

This invention discloses a general method for amino acid sequence generation and screening based on machine learning. This method can be used to generate various functional amino acid sequences, such as viral amino acid sequences. It uses a diffusion-based denoising probability model as a deep learning generation framework. Based on the prior distribution of the data, it progressively denoises the diffused data or random noise through a back-diffusion process. Based on existing amino acid sequence databases, it can achieve de novo design of various functional amino acid sequences. For example, when applied to viral amino acid sequence generation and screening, the method includes the following steps: S1, constructing a dataset for training the model from experimental data; S2, training the viral amino acid sequence generation and screening device using the dataset; S3, using the trained viral amino acid sequence generation and screening device to generate and screen viral amino acid sequences; and S4, validating the viral amino acid sequences in the target viral amino acid sequence library through wet experiments. This invention also provides systems and devices related to the general method for amino acid sequence generation and screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a general method for amino acid sequence generation and screening based on machine learning, which can be used for the generation of various functional amino acid sequences. In particular, it relates to a method and system for generating and screening viral amino acid sequences based on machine learning, including viral vector design and construction, generation of viral amino acid sequences with high specificity and high adaptability to viral amino acid sequence production, and verification of algorithm results. Background Technology

[0002] The structure and function of proteins are determined by their amino acid sequences, which are formed through long-term evolution in nature. In recent years, with the growth of biological data and the rapid development of artificial intelligence (AI) technology, AI has made substantial breakthroughs in the design and development of biotechnology and medicine. For example, in the field of protein crystal structure analysis, obtaining high-precision protein crystal structures previously required expensive and time-consuming methods such as cryo-electron microscopy, nuclear magnetic resonance (NMR), or X-ray diffraction. Precise protein structures help us analyze life processes and disease mechanisms at the molecular level and design more effective drugs based on disease targets. However, the number of known proteins is in the hundreds of millions, and currently, less than 50,000 high-precision protein structures have been resolved. On July 28, 2022, DeepMind announced that its AlphaFold could predict the structures of over 200 million proteins from 1 million species with high confidence, covering almost all known proteins on Earth, greatly advancing human research on protein function. However, when the structure and function of natural proteins cannot meet the needs of industrial or medical applications, protein design is necessary to obtain specific functional proteins. Based on the hypothesis that sequence determines structure and structure determines function, de novo protein design is almost equivalent to de novo design of amino acid sequences. On September 15, 2022, a research team from Professor David Baker's laboratory published a new study in Science showing that artificial intelligence technology can create practical protein molecules faster and more accurately than before.

[0003] However, the methods described above are insufficient to address all the challenges of current protein design. Developing a universal method for generating and screening amino acid sequences to aid in the design of functional protein molecules is a long-sought goal in the field. However, existing methods suffer from limitations such as limited design results and excessive sensitivity to the design details of the main chain structure, restricting the diversity and variability of the designed main chain structure. Currently, the design of protein amino acid sequences and the screening of highly active amino acid sequences based on experimental or computer-aided methods remain at a technological bottleneck and represent one of the most challenging scientific problems.

[0004] The Denoising Diffusion Probabilistic Model (DDPM) was proposed in 2015 and became widely known in 2020. It has mainly demonstrated its powerful capabilities in computer vision and has also been reported to have applications in small molecule design. However, there has been no attempt to apply this theoretical method to amino acid sequence design. This is because the traditional continuous diffusion denoising probabilistic model has been very successful in computer vision and audio processing, but due to the inherent discreteness of amino acid sequence data, the continuous diffusion denoising probabilistic model is difficult to apply to the generation of amino acid sequence data. Summary of the Invention

[0005] To address the aforementioned challenges, the inventors considered the inherent discreteness of sequence data when using the diffusion denoising probability model. In representing sequence data, they cleverly and specifically used word vector representation to enable gradient updates of the continuous latent variables in the diffusion denoising probability model. Furthermore, they used the method of using the maximum value of the independent variable point set to restore the discrete values ​​of the sequence data. Finally, the diffusion denoising probability model can be applied to the generation of amino acid sequence data.

[0006] The method of this invention has demonstrated excellent empirical results in predicting functional amino acid sequences such as viral capsid proteins, membrane-penetrating peptide sequences, antimicrobial peptide sequences, and antibody sequences. Specifically, this invention provides a machine learning-based method for generating and screening amino acid sequences, characterized by the following steps:

[0007] SI, constructing a dataset from experimental data for training models;

[0008] SII, using a dataset to train an amino acid sequence generation and screening device; includes the following steps:

[0009] SII-1 encodes the amino acid sequences in the dataset using features;

[0010] SII-2 trains the specific amino acid sequence generation module;

[0011] In the specific amino acid sequence generation module, a deep learning generation framework based on a diffusion denoising probability model is used to learn how to denoise and demaximize the diffusion process in order to learn an effective distribution of amino acid sequence information. This enables the generation of highly specific amino acid sequences from noise that meets a specific prior distribution during the sampling stage.

[0012] SII-3 Training Amino Acid Sequence Performance Evaluation Module;

[0013] In the amino acid sequence performance evaluation module, the relevant functions of the amino acid sequence are predicted based on the amino acid sequence.

[0014] SIII uses a trained amino acid sequence generation and screening device to generate and screen amino acid sequences. The specific steps are as follows:

[0015] SIII-1: After setting the amino acid sequence length and the number of sequences to be generated, input the parameters into the specific amino acid sequence generation module.

[0016] In S III-2, the specific amino acid sequence generation module, noise is removed from noise that meets a specific prior distribution to generate a preset number and length of highly specific amino acid sequences.

[0017] In S III-3, the performance evaluation module for amino acid sequences, after receiving all generated amino acid sequences, the corresponding amino acid sequence-related functions are evaluated for each amino acid sequence to obtain the corresponding amino acid sequence-related functions.

[0018] S III-4, Scoring Module

[0019] Based on the generated amino acid sequences, different scoring functions are used to score the generated amino acid sequences for different functional amino acid sequence generation tasks;

[0020] S III-5, generating a target amino acid sequence library.

[0021] The amino acid sequences are sorted according to their scores, and the top P amino acid sequences with the highest scores are selected to form the target amino acid sequence library, where P is a positive integer.

[0022] SIV is used to verify the amino acid sequences in the target amino acid sequence library;

[0023] SIV-1 was used to conduct experiments on amino acid sequences in the target amino acid sequence library to obtain experimental data. The functions associated with the experimentally obtained amino acid sequences were compared with the predicted functions associated with the amino acid sequences using evaluation indicators to verify the effectiveness of the amino acid sequence generation and screening device in generating amino acid sequences.

[0024] The method of this invention can be well applied to the design of viral vectors in gene therapy. Specifically, gene therapy technology is based on molecular biology methods to edit and modify the target gene or gene expression product in a patient's body, thereby achieving disease treatment. Adeno-associated virus (AAV) is a small virus with a single-stranded DNA genome. After artificial modification and transformation, it is not pathogenic. Therefore, when AAV enters the human body to participate in gene therapy, most people's immune system will not produce an immune response. Currently, due to the low genotoxicity of AAV in humans, AAV is the most widely used vector tool for gene therapy. The tissue tropism of natural AAV serotypes prevents the wide and effective targeting of target tissues and cells during treatment, which hinders the further development of gene therapy. However, there is still considerable room for improvement in AAV's targeted delivery and tissue enrichment. Experiments have shown that some recombinant AAVs have a certain selectivity in infecting living animal tissues. It is worth noting that the different infectivity characteristics of AAV mainly depend on the capsid protein encoded by its structural gene Cap. Adenovirus (Ad) is one of the earliest viral vectors used clinically for in vivo gene therapy. Ad viruses (Ad) are a type of DNA virus with a genome size of 34-43 kb, encapsulated in a non-enveloped icosahedral viral particle. Ad viruses possess high immunogenicity. Due to these characteristics, Ad vector-based gene therapy is primarily used for cancer treatment and infectious disease vaccination. Retroviruses, also known as retrotraceromycin, are a type of RNA virus whose genetic information is stored on RNA, not DNA. Unlike other RNA viruses, retroviral RNA does not self-replicate. After entering a host cell, the reverse transcriptase in the viral nucleus transcribes the RNA into cDNA, which is then used to synthesize double-stranded DNA. Integrase integrates the double-stranded DNA into the host cell's chromosomal DNA, thereby introducing non-viral genes into the cell. These genes can then be transferred to daughter cells through mitosis in vitro. Retroviruses specifically infect dividing cells, such as embryonic stem cells, neural stem cells, hematopoietic stem cells, and blood cells. In practical applications, gamma-retroviral vectors are often used in gene therapy due to their wide transfection range and high transfection rate. Lentivirals are a type of retrovirus, named for their long incubation period and slow development of clinical symptoms. Lentiviral vectors, on the other hand, are created by modifying and recombining lentiviruses, thereby removing their biological hazards while utilizing their high infectivity to express target genes. However, lentiviral vectors are not suitable for in vivo studies because their titers are insufficient for in vivo applications and they exhibit strong immunogenicity. Herpes simplex virus (HSV) is an enveloped virus with a double-stranded DNA genome exceeding 150 kb. This viral genome encodes approximately 90 genes; half of these genes are non-essential and can be removed / replaced in recombinant vectors, thus providing a high capacity of exogenous DNA.Eight human HSV serotypes have been identified, each exhibiting different tropisms. Currently, three main types of HSV vectors are used for gene therapy: amplified HSV, replication-defective HSV, and replication-capable HSV.

[0025] Based on the above, the method of this invention can efficiently generate high-quality viral amino acid sequences. This invention uses only the sequence information of viral amino acids as algorithm training data and employs a deep learning generation framework based on a diffusion denoising probability model to remove noise from the prior distribution to generate data, achieving the generation of viral amino acid sequences from scratch. This framework uses a neural network to learn the random noise superimposed at each time step in the forward diffusion process and performs denoising and demaximization in the backward diffusion process to learn an effective distribution of amino acid sequence information. In the sampling stage, it generates highly specific and production-fit viral amino acid sequences from noise that satisfies a specific prior distribution. This preserves the diversity of viral amino acid sequence design while also considering design efficiency, improving the experimental speed of screening viral amino acid sequences while reducing experimental costs. More specifically, this invention also proposes a machine learning-based method and system for generating and screening viral amino acid sequences, and a method for predicting the production fitness of viruses based on viral amino acid sequences.

[0026] This invention employs strategies such as word embedding, tag encoding, vector representation, and positional encoding to characterize viral amino acid sequences. By extracting features from the viral amino acid sequences and extracting semantic and syntactic information, it predicts the production fitness of the virus and screens viral amino acid sequences. Specifically, it includes the following steps:

[0027] S1, Construct a dataset from experimental data to train the model;

[0028] S2, Training the Virus Amino Acid Sequence Generation and Screening Device Using a Dataset; including the following steps:

[0029] S21 encodes the amino acid sequences in the dataset using features;

[0030] S22 trains the specific amino acid sequence generation module;

[0031] The specific amino acid sequence generation module uses a deep learning generation framework based on a diffusion denoising probability model. It uses a neural network to learn the random noise superimposed at each time step in the forward diffusion process, and performs denoising and demaximization in the reverse diffusion process to learn the effective distribution of amino acid sequence information. In the sampling stage, it generates viral amino acid sequences with a preset number and length from noise that meets a specific prior distribution.

[0032] S23 Training Virus Amino Acid Sequence Production Fitness Prediction Module;

[0033] In the viral amino acid sequence production fitness prediction module, the production fitness of the viral amino acid sequence is predicted based on the viral amino acid sequence (it should be noted that the higher the production fitness, the stronger the ability of the amino acid sequence to generate virus).

[0034] S3 uses a trained viral amino acid sequence generation and screening device to generate and screen viral amino acid sequences. The specific steps are as follows:

[0035] S31, After setting the viral amino acid sequence length and generation quantity, input the parameters into the specific amino acid sequence generation module;

[0036] In S32, the specific amino acid sequence generation module generates highly specific viral amino acid sequences of a preset number and length by denoising from noise that satisfies a specific prior distribution.

[0037] S33, in the viral amino acid sequence production fitness prediction module, after receiving all generated viral amino acid sequences, the viral amino acid sequence production fitness is predicted for each viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness.

[0038] S34, Virus scoring module

[0039] The viral amino acid sequence is scored based on the predicted viral amino acid sequence production fitness.

[0040] S35, generate the target virus amino acid sequence library.

[0041] The viral amino acid sequences are sorted according to their scores, and the top P amino acid sequences with the highest scores are selected to form the target viral amino acid sequence library, where P is a positive integer.

[0042] S4, verify the viral amino acid sequences in the target viral amino acid sequence library;

[0043] S41, Experiments are conducted on viral amino acid sequences in the target viral amino acid sequence library to obtain experimental data. The production fitness of the experimentally obtained viral amino acid sequences is compared with the predicted production fitness of viral amino acid sequences using evaluation indicators to verify the effectiveness of the viral amino acid sequence generation and screening device in generating viral amino acid sequences.

[0044] In a preferred embodiment of the present invention, step S1, which constructs a dataset for training the model from experimental data, specifically includes the following steps;

[0045] S11, Steps for constructing a viral plasmid library

[0046] H amino acid sequences of length L are randomly generated. Each amino acid sequence is linked with a specific barcode. All amino acid sequences linked with barcodes are pooled together to construct an amino acid sequence pool. The amino acid sequence library in the amino acid sequence pool is used to replace some sites of the target plasmid to obtain a viral plasmid. Different viral plasmids constitute a viral plasmid library.

[0047] S12, Viral amino acid sequence production fitness data collection steps

[0048] The frequency of viral plasmids was calculated by high-throughput sequencing of plasmid DNA extracted from viral plasmids.

[0049] Simultaneously, after transfecting the viral plasmid into the cells, the virus was purified, and then the viral DNA was extracted and subjected to high-throughput calculations to obtain the viral frequency.

[0050] Finally, the fitness of a single viral amino acid sequence is calculated using viral plasmid frequency and viral frequency.

[0051] S13 Constructing the Dataset

[0052] The dataset uses amino acid sequences that can become viruses as samples, and the corresponding amino acid sequence produces fitness as the dataset label.

[0053] In a preferred embodiment of the present invention, S22 trains the specific amino acid sequence generation module, specifically including the following steps;

[0054] Viral amino acid sequences are selected from the training set according to a set batch size. After tagging the viral amino acid sequences (i.e., numbering amino acids 'A', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'K', 'L', 'M', 'N', 'P', 'Q', 'R', 'S', 'T', 'V', 'W', 'Y' according to their natural amino acid count from 0 to 19), the tagged viral amino acid sequences are used as input to the model. The tagged viral amino acid sequences are defined as... Where N is the number of amino acids in a single amino acid sequence, a iThis is the encoded value of the i-th amino acid tag. The purpose of the diffusion model is to learn the spatial distribution p(x) of the amino acid sequence. The internal workflow of the model is as follows: First, the diffusion time step is set to t (t is a positive integer from 0 to 100000). The diffusion time step is used to limit the number of steps of predicted noise during model training. A diffusion probability model defines two Markov chains for the diffusion process. The forward diffusion process gradually adds noise into x until it becomes all noise. The reverse diffusion process learns how to reverse the forward diffusion process and gradually eliminate some noise from the noise to restore the true data x.

[0055] Forward diffusion can be defined by the following formula:

[0056]

[0057]

[0058] α t =1-β t α t x is the weight value of the added noise, which increases with the time step. z is the noise that follows a Gaussian distribution. t Let I represent the amino acid sequence at time step t, and let I represent the identity matrix, which is a matrix with all elements equal to 1.

[0059] The reverse denoising process can be defined by the following formula:

[0060]

[0061]

[0062] μ θ The parameterized neural network learns the mean and variance of the inverse distribution. The parameters are randomly set, therefore what needs to be learned during training is μ. θ .

[0063] We can use variational inference to obtain the variational lower bound (VLB) function to optimize the negative log-likelihood as the maximization objective:

[0064]

[0065] make This is the objective function we want to optimize.

[0066] During model training, the noise introduced during forward diffusion is used as training labels. The model is then used to predict the noise that needs to be removed during backward diffusion, ensuring that the distributions of the introduced noise and the noise to be removed are as similar as possible. Based on this principle, all data is divided into multiple batches. Each batch of original viral amino acid sequence features is input into the model, and random noise is added to the original viral amino acid sequence features sequentially over time steps—essentially "noising" the original data. The data with added random noise and its corresponding time step are then input into the neural network to predict the noise that needs to be removed during backward diffusion at that time step. The internal training steps of the neural network are as follows:

[0067] The data with added random noise and its corresponding time step are input into the neural network. First, the data with added random noise is input into the linear layer to extract feature information. After the feature information is extracted, the corresponding time step value is added, and then the feature is delinearized by the ReLU activation function to obtain the activated feature. This process is repeated N times (N is a positive integer from 0 to 1000).

[0068] The predicted noise is then output after feature aggregation through a linear layer.

[0069] The mean squared error is calculated using the predicted noise and the actual added noise as the loss function for backpropagation of the neural network.

[0070] The above amino acid generation steps are completed by iteratively selecting viral amino acid sequences from the entire training set according to the batch size. This process is repeated M times until the loss function is stable, where M is a positive integer. The model training parameters are then saved.

[0071] In a preferred embodiment of the present invention, the S23 training virus amino acid sequence production fitness prediction module specifically includes the following steps;

[0072] The viral amino acid sequence features are simultaneously input into multiple convolutional blocks in parallel. The features extracted from these convolutional blocks are concatenated in the hidden layer. A residual module is used to connect the residual network to retain the original features with a certain probability while updating the features using convolutional layers. The previously extracted features are then input into the residual module, and layer normalization (LN) is used to aggregate the features to obtain aggregated features. The aggregated features are then processed using the sigmoid activation function to obtain sequence information weights. Simultaneously, the aggregated features are delinearized using the ReLU activation function to obtain activation features. The sequence information weights are multiplied by the activation features to obtain weighted activation information. The sequence information weights obtained earlier are subtracted from 1 and multiplied by the previous viral amino acid sequence features to obtain the original weight information. The weighted activation information and the original weight information are added together to form the predicted features. The predicted features are then input into the temporary fallback method and then into the linear layer. Finally, the leaky ReLU activation function is used as the output for predicting the fitness of the viral amino acid sequence.

[0073] After Q iterations, where Q is a positive integer, the training of the viral amino acid sequence production fitness prediction module ends.

[0074] In a preferred embodiment of the present invention, the execution steps of the S32 specific amino acid sequence generation module specifically include the following steps;

[0075] After adding noise to the sequence data and generating random noise through the forward diffusion process, the random noise (i.e., the data that needs to be denoised) and the current time step are input into the neural network to predict the noise that should be removed. The internal training steps of the neural network are as follows:

[0076] The data to be denoised and its corresponding time step are input into the neural network. First, the data to be denoised is input into a linear layer to extract feature information. After the feature information is extracted, the corresponding time step value is added, and then the feature is delinearized by the ReLU activation function to obtain the activated feature. This process is repeated N times (N is a positive integer from 0 to 1000). Then, the feature is aggregated by a linear layer and the predicted noise to be removed is output.

[0077] The model predicts the noise that needs to be removed based on the current time step, starts with random noise, and removes noise step by step. It generates a preset number and length of highly specific viral amino acid sequences by removing noise from noise that meets a specific prior distribution.

[0078] Repeat the above steps. Once the number of viral amino acid sequences reaches the preset number, the specific amino acid sequence generation module stops generating viral amino acid sequences.

[0079] This invention also provides a viral amino acid sequence generation and screening device, including a parameter setting module, a specific amino acid sequence generation module, a viral amino acid sequence production fitness prediction module, a virus scoring module, and a target viral amino acid sequence library, the specific connection structure of which is as follows:

[0080] The parameter setting module is used to set the length and number of viral amino acid sequences generated, and the parameters are input into the specific amino acid sequence generation module.

[0081] In the specific amino acid sequence generation module, highly specific viral amino acid sequences of a preset number and length are generated by denoising noise from noise that meets a specific prior distribution; the generated viral amino acid sequences are then input into the viral amino acid sequence production fitness prediction module.

[0082] The viral amino acid sequence production fitness prediction module receives the viral amino acid sequence and performs viral amino acid sequence production fitness prediction on the viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness.

[0083] The virus scoring module scores the viral amino acid sequence based on the predicted viral amino acid sequence production fitness.

[0084] The target virus amino acid sequence library is sorted according to the score of the viral amino acid sequence, and the top P amino acid sequences with the highest score values ​​are selected and saved, where P is a positive integer;

[0085] In a preferred embodiment of the present invention, the specific amino acid sequence generation module operates as follows:

[0086] After adding noise to the sequence data and generating random noise through the forward diffusion process, the random noise, i.e. the data that needs to be denoised, and the current time step are input into the neural network. The neural network predicts the noise added at the current time step in the forward diffusion process and removes it from the data. The model removes the random noise in the sequence data step by step according to the noise predicted at the current time step. Finally, by denoising from the noise that meets a specific prior distribution, a highly specific viral amino acid sequence of a preset number and length is generated.

[0087] Repeat the above steps. Once the number of viral amino acid sequences reaches the preset number, the specific amino acid sequence generation module stops generating viral amino acid sequences.

[0088] This invention also provides a system for implementing the machine learning-based viral amino acid sequence generation and screening method described in the claims, comprising four main modules: a cloud computing and supercomputing platform, a viral vector design and development laboratory, a viral amino acid sequence generation and screening device, and an algorithm result verification laboratory; wherein:

[0089] The cloud computing and supercomputing platform accepts operation instructions from users or administrators through the I / O interface and assigns them corresponding permissions. It is responsible for collecting and managing online information data related to virus sequence design, transmitting local experimental data into the storage unit, and allocating the available computing resources to the computing tasks submitted by users according to priority through the computing unit to execute the corresponding tasks.

[0090] The viral vector design and development laboratory is used to obtain viral amino acid sequences experimentally, as well as the production fitness of each viral amino acid sequence;

[0091] A viral amino acid sequence generation and screening device is used to generate a target viral amino acid sequence library with high specificity and high viral amino acid sequence production fitness, including: a parameter setting module, a specific amino acid sequence generation module, a viral amino acid sequence production fitness prediction module, a virus scoring module, and a target viral amino acid sequence library.

[0092] In the algorithm result verification laboratory, the production fitness of the corresponding viral amino acid sequences stored in the target viral amino acid sequence library was obtained through experiments, and compared with the production fitness predicted by the production fitness prediction module to verify the effectiveness of the viral amino acid sequence generation and screening device.

[0093] Compared with the prior art, the present invention has the following contributions:

[0094] 1. In the process of generating amino acid sequences using machine learning methods, a diffusion-based denoising probability model is used to learn from existing amino acid sequence information. During forward diffusion, random noise is added to the existing training data step-by-step according to the diffusion time step. During backward diffusion, the trained model is used to deduce the added noise step-by-step from the noisy sequence data obtained during the multi-step noise addition process in forward diffusion, removing random noise from the prior distribution to generate data. The amino acid sequence samples generated by this method of a specific length make the characteristics of the generated sequence samples more closely resemble the real sequences in terms of spatial distribution (see reference). Figure 9 and Figure 10 );

[0095] 2. Viral amino acid sequences are designed from two aspects: high specificity and high production adaptability. This ensures sequence diversity while enhancing the feasibility of viral synthesis, enabling the model to generate more meaningful viral amino acid sequence data and improving the efficiency of viral amino acid sequence design and screening.

[0096] 3. Apply forward diffusion and back diffusion methods to enable the model to learn more abstract and effective viral amino acid information and features, thereby learning viral amino acid sequence information more comprehensively and enhancing the model's ability to extract and represent sequence features.

[0097] 4. Traditional continuous diffusion denoising probabilistic models have been very successful in computer vision and audio processing. However, due to the inherent discreteness of amino acid sequence data, continuous diffusion denoising probabilistic models are difficult to apply to amino acid sequence data generation. Therefore, when using the diffusion denoising probabilistic model, we considered the inherent discreteness of sequence data. When representing sequence data, we used word vector representation to enable gradient updates of the continuous latent variables in the diffusion denoising probabilistic model. Furthermore, we used the method of maximizing the independent variable point set to restore the discrete values ​​of the sequence data, making the diffusion denoising probabilistic model applicable to amino acid sequence data generation.

[0098] 5. It should be emphasized that the specific embodiments of this invention are merely examples of viral amino acid sequence design to illustrate the general amino acid sequence design method of this invention. The method of this invention can be used for the design of all functional amino acid sequences, including but not limited to the prediction and design of functional amino acid sequences such as membrane-penetrating peptide sequence design, antimicrobial peptide sequence design, and antibody sequence design. Attached Figure Description

[0099] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for initial preferred embodiment purposes only and are not intended to limit the embodiments of the invention. Furthermore, the same reference numerals denote the same parts throughout all the drawings. In the drawings:

[0100] Figure 1 An overview diagram of the machine learning-based viral amino acid sequence generation and screening method and system provided in the embodiments of the present invention is shown;

[0101] Figure 2 This invention provides a flowchart for constructing a dataset for training a model from experimental data, according to an embodiment of the invention.

[0102] Figure 3 This illustrates the steps of training a viral amino acid sequence generation and screening device using a dataset, as provided in an embodiment of the present invention.

[0103] Figure 4 The present invention illustrates the steps of generating and screening viral amino acid sequences using the viral amino acid sequence generation and screening apparatus provided in this embodiment.

[0104] Figure 5The following steps for verifying viral amino acid sequences in a target viral amino acid sequence library, as provided in an embodiment of the present invention, are illustrated:

[0105] Figure 6 This embodiment of the invention illustrates the probability density distribution of viral amino acid sequences generating fitness when constructing a dataset for training a model using experimental data.

[0106] Figure 7 This invention illustrates a Pearson correlation heatmap showing the frequency of viral amino acid sequences in multiple plasmid replication experiments and multiple virus replication experiments when constructing a dataset for training a model using experimental data.

[0107] Figure 8 This embodiment of the invention illustrates a Pearson correlation analysis of the viral amino acid sequence production fitness predicted by the viral amino acid sequence production fitness prediction module and the actual viral amino acid sequence production fitness.

[0108] Figure 9 This illustrates the spatial distribution of sequences generated using a diffusion model and real viral amino acid sequence data after tsne dimensionality reduction in an embodiment of the present invention.

[0109] Figure 10 This invention illustrates the spatial distribution of viral amino acid sequences generated using VAE (Variable Differential Autoencoder), viral amino acid sequences generated by a diffusion model, and real viral amino acid sequence data after dimensionality reduction using tsne in an embodiment of the invention.

[0110] Figure 11 This illustration shows the spatial distribution of viral amino acid sequences generated using a diffusion model through different diffusion time steps (100 steps were used here) and the actual viral amino acid sequence data after dimensionality reduction using tsne in an embodiment of the present invention. Detailed Implementation

[0111] To better understand the technical solution of the present invention, the specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The same reference numerals in the drawings indicate elements with the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0112] This invention proposes a method for generating and screening viral amino acid sequences based on machine learning, such as... Figure 1 As shown, a viral amino acid sequence generation method based on a diffusion denoising probability model generates and experimentally verifies viral amino acid sequences based on viral production fitness information, aiming to find viral amino acid sequences with high transfection efficiency and high virus-forming ability. The specific steps are as follows:

[0113] S1, constructing a dataset from experimental data for training the model, such as... Figure 2 As shown, it includes the following steps:

[0114] S11, Steps for constructing a viral plasmid library

[0115] M amino acid sequences (M ranging from 10,000 to 200,000) of length N (N ranging from 1 to 100) are randomly generated. Each amino acid sequence is linked with a specific barcode. All amino acid sequences linked with barcodes are pooled together to construct an amino acid sequence pool. The amino acid sequences in the pool are used to replace certain sites on the target plasmid to obtain viral plasmids. Different amino acid sequences or substitutions of different sites on the target plasmid will yield different viral plasmids. These different viral plasmids together constitute a viral plasmid library.

[0116] Each amino acid sequence is linked to a specific barcode for subsequent high-throughput sequencing.

[0117] The viral plasmids in the current viral plasmid library are only potential viral plasmids. Further testing, including viral vector packaging, viral purity detection, and in vitro biological activity detection, is needed to determine whether the obtained viral plasmids can ultimately become viruses.

[0118] S12, Viral amino acid sequence production fitness data collection steps

[0119] The viral plasmid frequency was obtained by extracting plasmid DNA from the viral plasmid and performing high-throughput sequencing. Specifically, the frequency of a single viral amino acid sequence in the plasmid was calculated by extracting plasmid DNA from the replaced plasmid and performing high-throughput sequencing.

[0120] Simultaneously, after transfecting the viral plasmid into cells, the virus was purified, and the viral DNA was extracted. The viral frequency was then calculated using high-throughput sequencing. The specific process is as follows: the viral plasmid was transfected into cells, and the virus was produced (i.e., the virus spread) in the cells. After purification, the virus purity was tested, and the viral DNA was extracted. The viral frequency was then calculated using high-throughput sequencing. The viral frequency refers to the number of times the same viral amino acid sequence appears in the virus after high-throughput sequencing.

[0121] Finally, the production fitness of a single viral amino acid sequence is calculated using viral plasmid frequency and viral frequency; production fitness is a quantitative representation of the ability of a viral amino acid sequence to generate virus. A higher production fitness indicates a stronger ability of the amino acid sequence to generate virus.

[0122] S13 Constructing the Dataset

[0123] The dataset uses viral amino acid sequences as samples, and the production fitness of the corresponding viral amino acid sequences as the dataset labels. In this embodiment, viral amino acid sequences containing premature stop codons and those with sequencing errors discovered during high-throughput sequencing are mainly used. The samples in the dataset are divided into training set: validation set: test set in a ratio of 7:1:2.

[0124] To avoid imbalanced training data and ensure the accuracy and research significance of the model training process, it is necessary to first check the normality of the overall distribution of the viral amino acid sequence production fitness data labels. For imbalanced data, strategies such as normalization, downsampling, and gradient pruning should be used to balance the data to ensure that the model is unbiased during the learning process. Figure 6 This describes the probability density distribution of fitness generated by viral amino acid sequences when constructing a dataset for training the model from experimental data.

[0125] S2 uses a dataset to train the virus amino acid sequence generation and screening device;

[0126] The specific amino acid sequence generation module and the viral amino acid sequence production fitness prediction module in the viral amino acid sequence generation and screening device were trained separately, such as... Figure 3 As shown, it includes the following steps:

[0127] S21 performs feature encoding on the amino acid sequences in the dataset.

[0128] The viral amino acid sequences in the dataset are input into the data preprocessing module to encode their features, resulting in viral amino acid sequence features. These features include the characteristics of each amino acid in the viral amino acid sequence. The data preprocessing module can employ existing strategies such as word embedding, tag encoding, vector representation, or positional encoding to characterize the feature information of the viral amino acid sequences.

[0129] S22 trains the specific amino acid sequence generation module;

[0130] There are currently 20 known amino acids. The amino acid sequence is composed of combinations of amino acids, but only some of these combinations can generate a virus.

[0131] The specific amino acid sequence generation module uses a deep learning generation framework based on a diffusion denoising probability model. It uses a neural network to learn how to denoise and demaximize the diffusion process in order to learn an effective distribution of amino acid sequence information. This enables the generation of highly specific viral amino acid sequences from noise that meets a specific prior distribution during the sampling stage.

[0132] In this embodiment, the specific amino acid sequence generation module adopts a deep learning generation framework based on a diffusion denoising probability model. The training steps for the specific amino acid sequence generation module are as follows: select viral amino acid sequences from the training set according to a set batch size; after tagging the viral amino acid sequences (i.e., numbering amino acids 'A', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'K', 'L', 'M', 'N', 'P', 'Q', 'R', 'S', 'T', 'V', 'W', 'Y' according to the number of natural amino acids, from 0 to 19), use the tagged viral amino acid sequences as input to the model. The tagged viral amino acid sequences are defined as... Where N is the number of amino acids in a single amino acid sequence, a i This is the encoded value of the i-th amino acid tag. The purpose of the diffusion model is to learn the spatial distribution p(x) of the amino acid sequence. The internal workflow of the model is as follows: First, the diffusion time step is set to t (t is a positive integer from 0 to 100000). The diffusion time step is used to limit the number of steps of predicted noise during model training. A diffusion probability model defines two Markov chains for the diffusion process. The forward diffusion process gradually adds noise into x until it becomes all noise. The reverse diffusion process learns how to reverse the forward diffusion process and gradually eliminate some noise from the noise to restore the true data x.

[0133] Forward diffusion can be defined by the following formula:

[0134]

[0135]

[0136] α t =1-β t α t x is the weight value of the added noise, which increases with the time step. z is the noise that follows a Gaussian distribution. t Let I represent the amino acid sequence at time step t, and let I represent the identity matrix, which is a matrix with all elements equal to 1.

[0137] The reverse denoising process can be defined by the following formula:

[0138]

[0139]

[0140] μ θ The parameterized neural network learns the mean and variance of the inverse distribution. The parameters are randomly set, therefore what needs to be learned during training is μ. θ .

[0141] We can use variational inference to obtain the variational lower bound (VLB) to optimize the negative log-likelihood as the maximization objective:

[0142]

[0143] make This is the objective function we want to optimize.

[0144] During model training, the noise introduced during forward diffusion is used as training labels. The model is then used to predict the noise that needs to be removed during backward diffusion, ensuring that the distributions of the introduced noise and the noise to be removed are as similar as possible. Based on this principle, all data is divided into multiple batches. Each batch of original viral amino acid sequence features is input into the model, and random noise is added to the original viral amino acid sequence features sequentially over time steps—essentially "noising" the original data. The data with added random noise and its corresponding time step are then input into the neural network to predict the noise that needs to be removed during backward diffusion at that time step. The internal training steps of the neural network are as follows:

[0145] The data with added random noise and its corresponding time step are input into the neural network. First, the data with added random noise is input into a linear layer to extract feature information. After the feature information is extracted, the corresponding time step value is added, and then the feature is delinearized by the ReLU activation function to obtain the activated feature. This process is repeated N times (N is a positive integer from 0 to 1000), and then the feature is aggregated by a linear layer to output the predicted noise.

[0146] The mean squared error is calculated using the predicted noise and the actual noise added, and used as the loss function for backpropagation of the neural network. Backpropagation is then performed, iteratively selecting viral amino acid sequences from the entire training set according to the batch size to complete the above amino acid generation steps. This process is repeated M times until the loss function is stable, and the model training parameters are saved. Therefore, when training the specific amino acid sequence generation module using a dataset containing N viral amino acid sequences, it requires (N / batch size)*M iterations before the specific amino acid sequence generation module training ends.

[0147] Where 1≤N≤50000000, N is a positive integer, and 1≤M≤50000, M is a positive integer.

[0148] S23 Training Virus Amino Acid Sequence Production Fitness Prediction Module

[0149] The viral amino acid sequence production fitness prediction module can predict the viral amino acid sequence production fitness based on the viral amino acid sequence. The higher the production fitness, the stronger the ability of the amino acid sequence to generate virus.

[0150] The training steps of the viral amino acid sequence production fitness prediction module are as follows: Feature information of a viral amino acid sequence is simultaneously input into multiple convolutional blocks in parallel. The features extracted from multiple convolutional blocks are concatenated in the hidden layer. A residual module is used to connect the network with a certain probability to retain the original features while updating the features using convolutional layers. The previously extracted features are input into the residual module, and then layer normalization (LN) is used to aggregate the features to obtain aggregated features. The aggregated features are then processed using the sigmoid activation function to obtain sequence information weights. Simultaneously, the aggregated features are delinearized using the ReLU activation function to obtain activation features. The sequence information weights are multiplied by the activation features to obtain weighted activation information. The sequence information weights obtained earlier are subtracted from 1 and multiplied by the previous viral amino acid sequence features to obtain the original weight information. The weighted activation information and the original weight information are added together to form the prediction features, allowing the original feature information to be input with a certain probability while updating the features. The prediction features are then input into a temporary fallback method and then into a linear layer. Finally, the LeakyReLU activation function is used as the output for predicting the viral amino acid sequence production fitness. After Q (1≤Q≤50000, where Q is a positive integer) rounds of iteration, the training of the viral amino acid sequence production fitness prediction module ends.

[0151] S3 uses a trained viral amino acid sequence generation and screening device to generate and screen viral amino acid sequences, such as... Figure 4 As shown, the specific steps are as follows:

[0152] S31, After setting the viral amino acid sequence length and generation quantity, input the parameters into the specific amino acid sequence generation module;

[0153] S32, the execution steps of the specific amino acid sequence generation module are as follows:

[0154] After adding noise to the sequence data and generating random noise through the forward diffusion process, the random noise (i.e., the data that needs to be denoised) and the current time step are input into the neural network to predict the noise that should be removed. The internal training steps of the neural network are as follows:

[0155] The data to be denoised and its corresponding time step are input into the neural network. First, the data to be denoised is input into a linear layer to extract feature information. After the feature information is extracted, the corresponding time step value is added, and then the feature is delinearized by the ReLU activation function to obtain the activated feature. This process is repeated N times (N is a positive integer from 0 to 1000). Then, the feature is aggregated by a linear layer and the predicted noise to be removed is output.

[0156] The model predicts the noise to be removed based on the current time step, removes the noise from the original random noise step by step, and generates a preset number and length of highly specific viral amino acid sequences by denoising the noise that satisfies a specific prior distribution.

[0157] Repeat the above steps. Once the number of viral amino acid sequences reaches the preset number, the specific amino acid sequence generation module stops generating viral amino acid sequences.

[0158] All generated viral amino acid sequences are sent to the viral amino acid sequence production fitness prediction module.

[0159] Besides the virus-specific sequence generation module described in this embodiment, which uses a diffusion denoising probability model to generate viral amino acid sequences, reliable sequence generation algorithms such as Variable Differential Autoencoders (VAEs) and Generative Adversarial Models (GANs) can also be used to generate virus-specific sequences. However, the VAE architecture consists of an encoder and a decoder. The encoder compresses the original data into low-dimensional feature vectors through a neural network, and the decoder restores the compressed feature vectors through a neural network to generate data that conforms to the original distribution. Although the generated data has high diversity, the generated results are usually rather ambiguous, the quality is difficult to guarantee, and there are problems such as the posterior distribution being assumed to be a decomposable Gaussian distribution, which is based on strong assumptions of the encoder.

[0160] S33, After receiving all the generated viral amino acid sequences, the viral amino acid sequence production fitness prediction module uses the trained viral amino acid sequence production fitness prediction module to predict the viral amino acid sequence production fitness for each viral amino acid sequence, and obtains the corresponding viral amino acid sequence production fitness.

[0161] S34, Virus scoring module

[0162] The viral amino acid sequences were scored based on the predicted production fitness of the viral amino acid sequences, and then sorted in descending order according to production fitness. The top N high-scoring sequences were selected for synthesis experiments using biosynthesis methods and verified.

[0163] S35, generate the target virus amino acid sequence library.

[0164] The viral amino acid sequences are sorted according to their scores, and the top P amino acid sequences with the highest scores are selected to form the target viral amino acid sequence library, where P is a positive integer.

[0165] S4, verify the viral amino acid sequences in the target viral amino acid sequence library;

[0166] S41, Experiments were conducted on the viral amino acid sequences in the target viral amino acid sequence library to obtain experimental data, such as... Figure 5 As shown:

[0167] If the number of viral amino acid sequences in the target viral amino acid sequence library is greater than 100, each viral amino acid sequence in the library is barcoded to construct an amino acid sequence pool. If the number of viral amino acid sequences is less than or equal to 100, a viral vector is constructed and packaged separately for each amino acid sequence. After replacing appropriate sites on the target plasmid with viral amino acid sequences, a viral plasmid is obtained. The viral plasmid frequency is calculated by high-throughput sequencing of plasmid DNA extracted from the viral plasmid. Simultaneously, the viral plasmid is transfected into cells for viral production in cells, followed by viral purification. After viral purification, viral purity is tested. Viral DNA is extracted, and viral frequency is calculated by high-throughput sequencing. In vitro and in vivo biological activity tests are then performed. The viral amino acid sequence production fitness is calculated based on the viral plasmid frequency and viral frequency. This fitness is compared with the predicted viral amino acid sequence production fitness using evaluation indicators to verify the effectiveness of the viral amino acid sequence generation and screening device. Finally, a viral vector with high specificity and high viral amino acid sequence production fitness, validated by biological experiments, is obtained.

[0168] S42, Determine the evaluation indicators.

[0169] In selecting evaluation metrics to assess the predictive performance of the model, this study involves a regression prediction task; therefore, the root mean square error, Pearson correlation coefficient, Spearman correlation coefficient, and coefficient of determination R0 were chosen. 2 To evaluate the performance of viral nucleocapsid production fitness data prediction; the root mean square error describes the distance between the predicted and actual values; the Pearson correlation coefficient and the Spearman correlation coefficient describe the correlation between the predicted and actual values, where the Pearson correlation coefficient describes the linear correlation between the two values, and the Spearman correlation coefficient is the rank form of the Pearson correlation coefficient, which describes the correlation between the two variables (e.g., when one variable increases, the other variable also increases), and it is related to the monotonicity of the function; the coefficient of determination R... 2 It is a dimensionless score that describes the effectiveness of the model, comparing the predictions to random guesses based on the average of the true values;

[0170] Figure 7-11 This is an example showing some actual measured metrics;

[0171] Figure 7 This is a Pearson correlation heatmap showing the frequency of viral amino acid sequences in multiple plasmid replication experiments and multiple virus replication experiments when constructing a dataset from experimental data for training the model.

[0172] Figure 8 The Pearson correlation analysis of the viral amino acid sequence production fitness prediction module and the actual viral amino acid sequence production fitness is performed.

[0173] Figure 9 This illustrates the spatial distribution of sequences generated using a diffusion model and real viral amino acid sequence data after tsne dimensionality reduction in an embodiment of the present invention.

[0174] Figure 10 This invention illustrates the spatial distribution of viral amino acid sequences generated using VAE (Variable Differential Autoencoder), viral amino acid sequences generated by a diffusion model, and real viral amino acid sequence data after dimensionality reduction using tsne in an embodiment of the invention.

[0175] Figure 11 This illustration shows the spatial distribution of viral amino acid sequences generated using a diffusion model through different diffusion time steps (100 steps were used here) and the actual viral amino acid sequence data after dimensionality reduction using tsne in an embodiment of the present invention.

[0176] This invention also discloses a viral amino acid sequence generation and screening device, comprising a parameter setting module, a specific amino acid sequence generation module, a viral amino acid sequence production fitness prediction module, a virus scoring module, and a target viral amino acid sequence library, the specific connection structure of which is as follows:

[0177] The parameter setting module is used to set the length and number of viral amino acid sequences generated, and the parameters are input into the specific amino acid sequence generation module.

[0178] The specific amino acid sequence generation module generates highly specific viral amino acid sequences of a preset number and length by denoising noise from noise that satisfies a specific prior distribution; the generated viral amino acid sequences are then input into the viral amino acid sequence production fitness prediction module; the working process of the specific amino acid sequence generation module is as follows:

[0179] After adding noise to the sequence data and generating random noise through the forward diffusion process, the random noise, i.e. the data that needs to be denoised, and the current time step are input into the neural network. The neural network predicts the noise added at the current time step in the forward diffusion process and removes it from the data. The model removes the random noise in the sequence data step by step according to the noise predicted at the current time step. Finally, by denoising from the noise that meets a specific prior distribution, a highly specific viral amino acid sequence of a preset number and length is generated.

[0180] Repeat the above steps. Once the number of viral amino acid sequences reaches the preset number, the specific amino acid sequence generation module stops generating viral amino acid sequences.

[0181] The viral amino acid sequence production fitness prediction module receives the viral amino acid sequence and performs viral amino acid sequence production fitness prediction on the viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness.

[0182] The virus scoring module scores the viral amino acid sequence based on the predicted viral amino acid sequence production fitness.

[0183] The target virus amino acid sequence library is sorted according to the score of the viral amino acid sequence, and the top P amino acid sequences with the highest score values ​​are selected and saved, where P is a positive integer;

[0184] This invention also discloses a system for analyzing viral amino acid sequences based on machine learning, such as... Figure 1 As shown, it comprises four main modules: a cloud computing and supercomputing platform, a virus vector design and development laboratory, a virus amino acid sequence generation and screening device, and an algorithm result verification laboratory; among which:

[0185] The cloud computing and supercomputing platform accepts operation instructions from users or administrators through the I / O interface and assigns them corresponding permissions. It is responsible for collecting and managing online information data related to virus sequence design, transmitting local experimental data into the storage unit, and allocating the available computing resources to the computing tasks submitted by users according to priority through the computing unit to execute the corresponding tasks.

[0186] The viral vector design and development laboratory is used to obtain viral amino acid sequences experimentally, as well as the production fitness of each viral amino acid sequence;

[0187] A viral amino acid sequence generation and screening device for generating a library of target viral amino acid sequences that are highly specific and have high adaptability to viral amino acid sequence production.

[0188] In the algorithm result verification laboratory, the production fitness of the corresponding viral amino acid sequences stored in the target viral amino acid sequence library was obtained through experiments, and compared with the production fitness predicted by the production fitness prediction module to verify the effectiveness of the viral amino acid sequence generation and screening device.

[0189] This invention has also achieved good prediction accuracy in predicting membrane-penetrating peptide sequences, antimicrobial peptide sequences, and antibody sequences. Based on the principles of prediction and screening methods, this invention can be applied to the prediction of any kind of functional amino acid sequences.

[0190] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating and screening amino acid sequences based on machine learning, characterized in that, The specific steps are as follows: SI constructs a dataset from experimental data to train the model; SII uses a dataset to train the amino acid sequence generation and screening device; this includes the following steps: SII-1 encodes the amino acid sequences in the dataset using features; SII-2 trains the specific amino acid sequence generation module; In the specific amino acid sequence generation module, a deep learning generation framework based on a diffusion denoising probability model is used. A neural network learns the noise added at each time step during the forward diffusion process. During the reverse diffusion process, the noisy sequence data is denoised and demaximized to learn an effective distribution of amino acid sequence information. This allows for the generation of highly specific amino acid sequences from noise that satisfies a specific prior distribution during the sampling phase. The process involves selecting viral amino acid sequences from the training set according to a set batch size; encoding the viral amino acid sequences to obtain viral amino acid sequence features; and using these features as input to the diffusion model. The internal operation of the model is as follows: A diffusion time step of t is set, where t is a positive integer from 0 to 100,000. The forward diffusion process gradually adds noise to the data until the data distribution approximately reaches the prior distribution. The reverse diffusion process starts from the prior distribution and iteratively transforms it into the desired distribution. Model training relies on the forward diffusion process to simulate noisy data. Data is fed into the model in batches, with the batch size set as needed. Each time, a random noise is added to the original viral amino acid sequence features of a batch according to different time steps. The data with added random noise and its corresponding time step are then input into the neural network to predict the noise. The diffusion time step is used to limit the number of steps for predicting noise during model training. A diffusion probability model defines two Markov chains for the diffusion process. The internal training steps of a neural network are as follows: First, the data with added random noise and its corresponding time step are input into the neural network. In the neural network, the data with added random noise is first input into a linear layer to extract feature information. After the feature information is extracted, the corresponding time step value is added, and then the ReLU activation function is used to delinearize the features to obtain activated features. This process is repeated P times, where P is a positive integer from 0 to 1000. Subsequently, the predicted noise is output after feature aggregation through a linear layer. Subsequently, the predicted noise and the actual added noise were used to calculate the mean squared error as the loss function for backpropagation of the neural network. Based on the batch size, the entire viral amino acid sequence in the training set is selected cyclically to complete all the steps of the above amino acid generation. The above loop is repeated M times until the loss function converges smoothly. M is a positive integer from 0 to 100000. The model training parameters are then saved. SII-3 Training Amino Acid Sequence Performance Evaluation Module; In the amino acid sequence performance evaluation module, the relevant functions of short peptides or proteins composed of amino acid sequences are predicted based on the amino acid sequence. SIII uses a trained amino acid sequence generation and screening device to generate and screen amino acid sequences. The specific steps are as follows: After setting the amino acid sequence length and number of sequences to be generated in S III-1, input the parameters into the specific amino acid sequence generation module; In the S III-2 specific amino acid sequence generation module, noise is removed from noise that meets a specific prior distribution to initially generate a preset number and length of highly specific amino acid sequences. In the S III-3 amino acid sequence performance evaluation module, after receiving all generated amino acid sequences, the performance evaluation function in the performance evaluation module performs a preliminary evaluation of each amino acid sequence generated by the model based on the corresponding functional requirements, and obtains the preliminary score of the corresponding amino acid sequence. S III-4 scoring module, Based on the generated amino acid sequences, different scoring functions are used to further score the generated amino acid sequences for different functional amino acid sequence generation tasks. S III-5 generates a target amino acid sequence library. The amino acid sequences are sorted by score, and the top P amino acid sequences with the highest scores are selected to form a target amino acid sequence library. Then, amino acid sequences with high specificity and relevant functions are screened out based on the scores; where P is a positive integer. SIV validates the amino acid sequences in the target amino acid sequence library; SIV-1 conducted wet experiments based on amino acid sequences from the target amino acid sequence library to obtain experimental data. The functions associated with the obtained amino acid sequences were compared with the predicted functions associated with the amino acid sequences using evaluation indicators to verify the effectiveness of the amino acid sequence generation and screening device in generating amino acid sequences.

2. A method for generating and screening viral amino acid sequences based on machine learning, characterized in that, The specific steps are as follows: S1, Construct a dataset from experimental data to train the model; S2, Training the Virus Amino Acid Sequence Generation and Screening Device Using a Dataset; including the following steps: S21, Encode the viral amino acid sequences in the dataset using features; S22, Training the specific amino acid sequence generation module; In the specific amino acid sequence generation module, a deep learning generation framework based on a diffusion denoising probability model is used. A neural network learns the noise added at each time step during the forward diffusion process. During the reverse diffusion process, the noisy sequence data is denoised and de-maximized to learn an effective distribution of amino acid sequence information. This allows for the generation of a preset number and length of viral amino acid sequences from noise satisfying a specific prior distribution during the sampling phase. Specifically, viral amino acid sequences are selected from the training set according to a set batch size. After feature encoding of the viral amino acid sequences, viral amino acid sequence features are obtained and used as input to the diffusion model. The internal operation of the model is as follows: A diffusion time step of t is set, where t is a positive integer from 0 to 100,000. The forward diffusion process gradually adds noise to the data until the data distribution approximately reaches the prior distribution. The reverse diffusion process starts from the prior distribution and iteratively transforms it into the desired distribution. Model training relies on the forward diffusion process to simulate noisy data. Data is fed into the model in batches, with the batch size set as needed. Each time, a random noise is added to the original viral amino acid sequence features of a batch according to different time steps. The data with added random noise and its corresponding time step are then input into the neural network to predict the noise. The diffusion time step is used to limit the number of steps for predicting noise during model training. A diffusion probability model defines two Markov chains for the diffusion process. The internal training steps of a neural network are as follows: First, the data with added random noise and its corresponding time step are input into the neural network. In the neural network, the data with added random noise is first input into a linear layer to extract feature information. After the feature information is extracted, the corresponding time step value is added, and then the ReLU activation function is used to delinearize the features to obtain activated features. This process is repeated P times, where P is a positive integer from 0 to 1000. Subsequently, the predicted noise is output after feature aggregation through a linear layer. Subsequently, the predicted noise and the actual added noise were used to calculate the mean squared error as the loss function for backpropagation of the neural network. Based on the batch size, the entire viral amino acid sequence in the training set is selected cyclically to complete all the steps of the above amino acid generation. The above loop is repeated M times until the loss function converges smoothly. M is a positive integer from 0 to 100000. The model training parameters are then saved. S23, Training Virus Amino Acid Sequence Production Fitness Prediction Module; In the viral amino acid sequence production fitness prediction module, the production fitness of the viral amino acid sequence is predicted based on the viral amino acid sequence; the higher the production fitness, the stronger the ability of the amino acid sequence to generate virus. S3. Using the trained viral amino acid sequence generation and screening device, viral amino acid sequence generation and screening are performed. The specific steps are as follows: S31, After setting the viral amino acid sequence length and generation quantity, input the parameters into the specific amino acid sequence generation module; In S32, the specific amino acid sequence generation module generates highly specific viral amino acid sequences of a predetermined number and length by denoising from noise that satisfies a specific prior distribution. S33, in the viral amino acid sequence production fitness prediction module, after receiving all generated viral amino acid sequences, the viral amino acid sequence production fitness is predicted for each viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness. S34, Virus scoring module The viral amino acid sequence is scored based on the predicted viral amino acid sequence production fitness. S35, generate the target virus amino acid sequence library. The viral amino acid sequences are sorted by score, and the top P amino acid sequences with the highest scores are selected to form a target viral amino acid sequence library. Amino acid sequences with high specificity and high production adaptability are then selected based on the scores; where P is a positive integer. S4, verify the viral amino acid sequences in the target viral amino acid sequence library; S41. Wet experiments are conducted based on viral amino acid sequences in the target viral amino acid sequence library to obtain experimental data. The production fitness of the viral amino acid sequences obtained in the experiment is compared with the predicted production fitness of the viral amino acid sequences using evaluation indicators to verify the effectiveness of the viral amino acid sequence generation and screening device in generating viral amino acid sequences.

3. The method for generating and screening viral amino acid sequences based on machine learning according to claim 2, characterized in that, S1 constructs a dataset for training the model from the experimental data, specifically including the following steps; S11, Steps for constructing a viral plasmid library H amino acid sequences of length L are randomly generated. Each amino acid sequence is linked with a specific barcode. All amino acid sequences linked with barcodes are pooled together to construct an amino acid sequence pool. The amino acid sequence library in the amino acid sequence pool is used to replace some sites of the target plasmid to obtain a viral plasmid. Different viral plasmids constitute a viral plasmid library. S12, Viral amino acid sequence production fitness data collection steps The frequency of viral plasmids was calculated by high-throughput sequencing of plasmid DNA extracted from viral plasmids. Simultaneously, after transfecting the viral plasmid into the cells, the virus was purified, and then the viral DNA was extracted and subjected to high-throughput calculations to obtain the viral frequency. Finally, the fitness of a single viral amino acid sequence is calculated using viral plasmid frequency and viral frequency. S13, Constructing the dataset The dataset uses amino acid sequences that can become viruses as samples, and the corresponding amino acid sequence produces fitness as the dataset label.

4. The method for generating and screening viral amino acid sequences based on machine learning according to claim 2, characterized in that, The S23 training virus amino acid sequence production fitness prediction module specifically includes the following steps; The viral amino acid sequence features are simultaneously input into multiple convolutional blocks in parallel. The features extracted from these blocks are concatenated in the hidden layer. A residual module is used to maintain the original features with a certain probability while updating the features using convolutional layers. The features extracted from the convolutional blocks are then input into the residual module, and layer normalization (LN) is used to converge the features to obtain converged features. The converged features are then processed using the sigmoid activation function to obtain sequence information weights. Simultaneously, the converged features are delinearized using the relu activation function to obtain activation features. The sequence information weights are multiplied by the activation features to obtain weighted activation information. The sequence information weights are subtracted from the previously obtained sequence information weights and multiplied by the previous viral amino acid sequence features to obtain the original weight information. The weighted activation information and the original weight information are added together to form the prediction features. The prediction features are then input into the temporary fallback method and then into the linear layer. Finally, the leaky relu activation function is used as the output for predicting the fitness of the viral amino acid sequence. The prediction features are then input into the temporary fallback method and then into the linear layer, where the sigmoid activation function is used as the output for predicting whether the viral sequence is real. After Q rounds of iteration, Q is a positive integer from 0 to 100,000, and the training of the viral amino acid sequence production fitness prediction module ends.

5. The method for generating and screening viral amino acid sequences based on machine learning according to claim 2, characterized in that, The execution steps of the S32 specific amino acid sequence generation module specifically include the following steps; After adding noise to the sequence data and generating random noise through a forward diffusion process, the random noise (i.e., the data that needs to be denoised) and the current time step are input into a neural network to predict the noise that should be removed. The internal training steps of the neural network are as follows: First, the data that needs to be denoised is input into a linear layer to extract feature information. After extracting the feature information, the corresponding time step value is added, and then the features are delinearized using the ReLU activation function to obtain activated features. This process is repeated Q times, and then the features are aggregated through a linear layer to output the predicted noise that needs to be removed, where Q is a positive integer from 0 to 1000. The model predicts the noise that needs to be removed based on the current time step and removes the noise from the original random noise step by step. By denoising from the noise that satisfies a specific prior distribution, a preset number and length of highly specific viral amino acid sequences are generated. Repeat the above steps. Once the number of viral amino acid sequences reaches the preset number, the specific amino acid sequence generation module stops generating viral amino acid sequences.

6. A viral amino acid sequence generation and screening device, comprising a parameter setting module, a specific amino acid sequence generation module, a viral amino acid sequence production fitness prediction module, a virus scoring module, and a target viral amino acid sequence library, the specific connection structure of which is as follows: The parameter setting module is used to set the length and number of viral amino acid sequences generated, and the parameters are input into the specific amino acid sequence generation module. The specific amino acid sequence generation module generates highly specific viral amino acid sequences of a preset number and length by denoising noise from noise that satisfies a specific prior distribution; the generated viral amino acid sequences are then input into the viral amino acid sequence production fitness prediction module. The viral amino acid sequence production fitness prediction module receives the viral amino acid sequence and performs viral amino acid sequence production fitness prediction on the viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness. The virus scoring module scores the viral amino acid sequence based on the predicted viral amino acid sequence production fitness. The target viral amino acid sequence library is sorted according to the score of the viral amino acid sequence, and the top P amino acid sequences with the highest score values ​​are selected for storage. Viral amino acid sequences with high specificity and high production adaptability are screened according to the score, where P is a positive integer from 0 to 100,000. in, The specific amino acid sequence generation module works as follows: After adding noise to the sequence data and generating random noise through the forward diffusion process, the random noise (i.e., the data that needs to be denoised) and the current time step are input into the neural network. The model predicts the noise added at the current time step in the forward diffusion process and removes it from the data. Based on the noise predicted at the current time step, the model removes random noise from the sequence data step by step. Finally, by denoising from noise that meets a specific prior distribution, a highly specific viral amino acid sequence of a predetermined number and length is generated. Repeat the above steps. Once the number of viral amino acid sequences reaches the preset number, the specific amino acid sequence generation module stops generating viral amino acid sequences. The specific amino acid sequence generation module is trained using the following method: Viral amino acid sequences are selected from the training set according to a set batch size; after feature encoding of the viral amino acid sequences, viral amino acid sequence features are obtained, and these features are used as input to the diffusion model. The internal operation of the model is as follows: A diffusion time step of t is set, where t is a positive integer from 0 to 100,000. The forward diffusion process gradually adds noise to the data until the data distribution approximately reaches the prior distribution. The reverse diffusion process starts from the prior distribution and iteratively transforms it into the desired distribution. Model training relies on the forward diffusion process to simulate noisy data. Data is fed into the model in batches, with the batch size set as needed. Each time, a random noise is added to the original viral amino acid sequence features of a batch according to different time steps. The data with added random noise and its corresponding time step are then input into the neural network to predict the noise. The diffusion time step is used to limit the number of steps for predicting noise during model training. A diffusion probability model defines two Markov chains for the diffusion process. The internal training steps of a neural network are as follows: First, the data with added random noise and its corresponding time step are input into the neural network. In the neural network, the data with added random noise is first input into a linear layer to extract feature information. After the feature information is extracted, the corresponding time step value is added, and then the ReLU activation function is used to delinearize the features to obtain activated features. This process is repeated P times, where P is a positive integer from 0 to 1000. Subsequently, the predicted noise is output after feature aggregation through a linear layer. Subsequently, the predicted noise and the actual added noise were used to calculate the mean squared error as the loss function for backpropagation of the neural network. Based on the batch size, the entire viral amino acid sequence in the training set is selected cyclically to complete all the steps of the above amino acid generation. The above loop is repeated M times until the loss function converges smoothly. M is a positive integer from 0 to 100000. The model training parameters are then saved.

7. A machine learning-based viral amino acid sequence generation and screening system, comprising the system for implementing the machine learning-based viral amino acid sequence generation and screening method according to any one of claims 2-5, including four main modules: a cloud computing and supercomputing platform, a viral vector design and development laboratory, a viral amino acid sequence generation and screening device, and an algorithm result verification laboratory; wherein: The cloud computing and supercomputing platform accepts operation instructions from users or administrators through the I / O interface and assigns them corresponding permissions. It is responsible for collecting and managing online information data related to virus sequence design, transmitting local experimental data into the storage unit, and allocating the available computing resources to the computing tasks submitted by users according to priority through the computing unit to execute the corresponding tasks. The viral vector design and development laboratory is used to obtain viral amino acid sequences experimentally, as well as the production fitness of each viral amino acid sequence; A viral amino acid sequence generation and screening device is used to generate a target viral amino acid sequence library with high specificity and high viral amino acid sequence production fitness, including: a parameter setting module, a specific amino acid sequence generation module, a viral amino acid sequence production fitness prediction module, a virus scoring module, and a target viral amino acid sequence library. In the algorithm result verification laboratory, the production fitness of the corresponding viral amino acid sequences stored in the target viral amino acid sequence library was obtained through experiments, and compared with the production fitness predicted by the production fitness prediction module to verify the effectiveness of the viral amino acid sequence generation and screening device.

8. A machine learning-based device for generating and screening viral amino acid sequences, characterized in that, include: Memory, used to store executable instructions; The processor, when executing executable instructions stored in the memory, implements the method for generating and screening viral amino acid sequences according to any one of claims 2 to 5.

9. A computer-readable storage medium, characterized in that, The device stores executable instructions for causing a processor to execute the executable instructions to implement the method for generating and screening viral amino acid sequences according to any one of claims 2 to 5.

Citation Information

Patent Citations

  • Machine learning implementation for multi-analyte assay of biological samples

    CN112292697A

  • Protein interaction site prediction method based on deep learning

    CN113643756A