A Machine Learning-Based Method and System for Viral Amino Acid Sequence Generation and Screening
By using machine learning methods to generate and screen viral amino acid sequences, the problem of low efficiency in viral amino acid sequence screening in existing technologies has been solved. This enables the efficient generation and screening of highly adaptable viral amino acid sequences, thereby improving the effectiveness of gene therapy.
Patent Information
- Application Number
- CN202211214161.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-09-30
AI Technical Summary
Existing technologies struggle to efficiently screen viral amino acid sequences with high viral adaptability and high targeting performance, hindering the further development of gene therapy.
A machine learning-based approach was adopted to extract and generate features from viral amino acid sequences using long short-term memory neural networks and convolutional neural networks. The production fitness of viral amino acid sequences was predicted by combining positional coding and residual connection networks, thereby generating and screening viral amino acid sequences with high specificity and high production fitness.
It improved the efficiency and accuracy of viral amino acid sequence generation, reduced experimental costs, enhanced the targeting and adaptability of viral vectors, and promoted the progress of gene therapy.
Smart Images

Figure CN117854585B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene therapy, and more particularly to a method and system for generating and screening viral amino acid sequences based on machine learning, including viral vector design and construction, generation of viral amino acid sequences with high specificity and high adaptability to viral amino acid sequence production, and verification of algorithm results. Background Technology
[0002] Gene therapy is a technology that uses molecular biology methods to edit and modify target genes or gene expression products in a patient's body, thereby treating the disease.
[0003] Adeno-associated virus (AAV) is a small virus with a single-stranded DNA genome. After artificial modification and alteration, it becomes non-pathogenic. Therefore, when AAV enters the human body for gene therapy, the immune system in most cases does not mount an immune response. Currently, due to its low genotoxicity in humans, AAV is the most widely used vector tool for gene therapy. The tissue tropism of natural AAV serotypes prevents broad and effective targeting of specific tissues and cells during treatment, hindering further development of gene therapy. However, there is still significant room for improvement in targeted delivery and tissue enrichment of AAV. Experiments have demonstrated that some recombinant AAVs exhibit selectivity in in vivo animal tissue infection. Notably, the different infection characteristics of AAV primarily depend on the capsid protein encoded by its structural gene Cap.
[0004] Adenovirus (Ad) is one of the earliest viral vectors used clinically for in vivo gene therapy. Ad is a type of DNA virus with a genome size of 34-43 kb, encapsulated in a non-enveloped icosahedral viral particle. Ad possesses high immunogenicity. Due to these characteristics, Ad-based gene therapy is primarily used for cancer treatment and infectious disease vaccination.
[0005] Retroviruses, also known as retroterrorites, are a type of RNA virus whose genetic information is stored on RNA, not DNA. Unlike other RNA viruses, retroviral RNA does not self-replicate. After entering a host cell, the reverse transcriptase in the viral nucleus transcribes the RNA into cDNA, which is then used to synthesize double-stranded DNA. Integrase integrates this double-stranded DNA into the host cell's chromosomal DNA, thus introducing non-viral genes into the cell. These genes can then be transferred to daughter cells through mitosis in vitro. Retroviruses specifically infect dividing cells, such as embryonic stem cells, neural stem cells, hematopoietic stem cells, and blood cells. In practical applications, gamma-retroviral vectors are often used in gene therapy due to their broad transfection range and high transfection rate.
[0006] Lentivirals are a type of retrovirus, named for their long incubation period and slow development of clinical symptoms. Lentiviral vectors are created by modifying and recombining lentiviruses to remove their biological hazards while utilizing their high infectivity to express target genes. However, lentiviral vectors are not suitable for in vivo studies because their titers are insufficient for in vivo applications and they exhibit strong immunogenicity.
[0007] Herpes simplex virus (HSV) is an enveloped virus with a double-stranded DNA genome exceeding 150 kb. The viral genome encodes approximately 90 genes; half of these genes are non-essential and can be removed / replaced in recombinant vectors, thus providing a high capacity of exogenous DNA. Eight human HSV serotypes have been identified, each exhibiting distinct tropisms. Currently, three main types of HSV vectors are used for gene therapy: amplified HSV, replication-defective HSV, and replication-capable HSV.
[0008] Designing viral protein amino acid sequences based on experimental or computer-aided methods and screening for viral amino acid sequences with high viral adaptability and high targeting performance is a key to breaking through the current bottleneck of gene therapy technology and is also one of the most challenging scientific problems at present. Summary of the Invention
[0009] To overcome the shortcomings of traditional experimentally-dependent viral amino acid design methods, this invention proposes a machine learning-based method and system for generating and screening viral amino acid sequences, incorporating computer-aided design. This method efficiently extracts high-quality viral amino acid sequences. It uses only the sequence information of viral amino acids as training data, treating the amino acid sequence as a time series and extracting its syntactic and semantic information. Based on the logical structure characteristics of the time series, information is extracted from a specifically distributed semantic space, generating highly specific and production-adaptive viral amino acid sequences step-by-step at each time step. This approach preserves the diversity of viral amino acid sequence design while maintaining design efficiency, improving the experimental speed of screening viral amino acid sequences while reducing experimental costs.
[0010] This invention proposes a method and system for generating and screening viral amino acid sequences based on machine learning, and a method for predicting the production fitness of viruses based on viral amino acid sequences. It applies strategies such as word embedding, graph representation, and positional encoding to characterize viral amino acid sequences, extracts semantic and syntactic information from the sequences to predict the production fitness of viruses, and screens viral amino acid sequences. Specifically, it includes the following steps: S1, constructing a dataset for training the model from experimental data;
[0011] S2, Training the Virus Amino Acid Sequence Generation and Screening Device Using a Dataset; including the following steps:
[0012] S21 performs feature encoding on the amino acid sequences in the dataset.
[0013] S22 trains the specific amino acid sequence generation module;
[0014] The specific amino acid sequence generation module generates viral amino acid sequences by learning the logical structure and syntactic and semantic features between existing amino acid sequences that can generate viruses.
[0015] S23 Training Virus Amino Acid Sequence Production Fitness Prediction Module
[0016] The viral amino acid sequence production fitness prediction module predicts the viral amino acid sequence production fitness based on the viral amino acid sequence. The higher the production fitness, the stronger the ability of the amino acid sequence to generate virus.
[0017] S3 uses a trained viral amino acid sequence generation and screening device to generate and screen viral amino acid sequences. The specific steps are as follows:
[0018] S31, After setting the viral amino acid sequence length and generation quantity, input the parameters into the specific amino acid sequence generation module;
[0019] S32, the specific amino acid sequence generation module generates a viral amino acid sequence of a preset number and length based on a randomly generated amino acid;
[0020] S33, After receiving all the generated viral amino acid sequences, the viral amino acid sequence production fitness prediction module performs viral amino acid sequence production fitness prediction on each viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness and the judgment of whether the viral amino acid sequence is real.
[0021] S34, Virus scoring module
[0022] The viral amino acid sequence is scored based on the predicted viral amino acid sequence production fitness and the judgment of whether the viral amino acid sequence is true.
[0023] S35, generate the target virus amino acid sequence library.
[0024] The viral amino acid sequences are sorted according to their scores, and the top P amino acid sequences with the highest scores are selected to form the target viral amino acid sequence library, where P is a positive integer.
[0025] S4, verify the viral amino acid sequences in the target viral amino acid sequence library;
[0026] S41, Experiments were conducted on viral amino acid sequences in the target viral amino acid sequence library to obtain experimental data. The production fitness of the experimentally obtained viral amino acid sequences was compared with the predicted production fitness of the viral amino acid sequences using evaluation indicators to verify the effectiveness of the viral amino acid sequence generation and screening device in generating viral amino acid sequences.
[0027] In a preferred embodiment of the present invention, step S1, which constructs a dataset for training the model from experimental data, specifically includes the following steps;
[0028] S11, Virus plasmid library construction steps
[0029] H amino acid sequences of length L are randomly generated. Each amino acid sequence is linked with a specific barcode. All amino acid sequences linked with barcodes are pooled together to construct an amino acid sequence pool. The amino acid sequence library in the amino acid sequence pool is used to replace some sites of the target plasmid to obtain a viral plasmid. Different viral plasmids constitute a viral plasmid library.
[0030] S12, Viral amino acid sequence production fitness data collection steps
[0031] The frequency of viral plasmids was calculated by high-throughput sequencing of plasmid DNA extracted from viral plasmids.
[0032] Simultaneously, after transfecting the viral plasmid into the cells, the virus was purified, and the viral DNA was extracted and the viral frequency was obtained by high-throughput calculation.
[0033] Finally, the fitness of a single viral amino acid sequence is calculated using viral plasmid frequency and viral frequency.
[0034] S13 Constructing the Dataset
[0035] The dataset uses amino acid sequences that can become viruses as samples, and the corresponding amino acid sequence produces fitness as the dataset label.
[0036] In a preferred embodiment of the present invention, S22 trains the specific amino acid sequence generation module, specifically including the following steps;
[0037] Select viral amino acid sequences from the training set according to the set batch size; after text encoding the viral amino acid sequences, prepend a number 0 to the beginning of each viral amino acid sequence and truncate the last amino acid;
[0038] All features of the processed viral amino acid sequence are used as input to the Long Short-Term Memory (LSTM) neural network. The original viral amino acid sequence, after text encoding and without further processing, is used as the target to be generated by the LTM neural network. The internal workflow of the LTM neural network is as follows: an input viral amino acid sequence is split into multiple individual amino acids. The previous amino acid is used as input to the LTM neural network to generate the amino acid at the next position. The next generated amino acid is used as input to the LTM neural network to generate the amino acid at the next position, and so on, until an amino acid sequence with the same length as the target amino acid sequence is generated.
[0039] The generated amino acid sequence is input into a linear layer for amino acid sequence feature extraction. The extracted features are then input into the softmax activation function to obtain the final generated amino acid sequence.
[0040] The selected viral amino acid sequence features are compared with the generated amino acid sequence features, the loss function is calculated, and backpropagation is performed.
[0041] The above amino acid generation steps are completed by iteratively selecting viral amino acid sequences from the entire training set according to the batch size. This process is repeated M times until the loss function is stable, where M is a positive integer. The model training parameters are then saved.
[0042] In a preferred embodiment of the present invention, the S23 training virus amino acid sequence production fitness prediction module specifically includes the following steps;
[0043] The viral amino acid sequence features are simultaneously input into multiple convolutional blocks in parallel. The features extracted from these blocks are concatenated in the hidden layer. A residual module is used to connect the network with a certain probability to preserve the original features while updating the features using convolutional layers. The previously extracted features are then input into the residual module, and layer normalization (LN) is used to aggregate the features to obtain aggregated features. The aggregated features are then processed using the sigmoid activation function to obtain sequence information weights. Simultaneously, the aggregated features are delinearized using the ReLU activation function to obtain activation features. The sequence information weights are multiplied by the activation features to obtain weighted activation information. The sequence information weights are subtracted from the previously obtained sequence information weights and multiplied by the previous viral amino acid sequence features to obtain the original weight information. The weighted activation information and the original weight information are added together to form the prediction features. The prediction features are then input into the temporary fallback method and then into the linear layer. Finally, the leaky ReLU activation function is used as the output for predicting the fitness of the viral amino acid sequence. The prediction features are then input into the temporary fallback method and then into the linear layer, where the sigmoid activation function is used as the output for predicting whether the viral sequence is real.
[0044] After Q iterations, where Q is a positive integer, the training of the viral amino acid sequence production fitness prediction module ends.
[0045] In a preferred embodiment of the present invention, the execution steps of the S32 specific amino acid sequence generation module specifically include the following steps;
[0046] An amino acid is randomly generated, and its features are characterized. The amino acid features are then input into a trained long short-term memory neural network.
[0047] The Long Short-Term Memory (LSTM) neural network generates viral amino acid sequences based on a set amino acid sequence length. Specifically, after inputting amino acid features into the trained LSM, the LSM generates the next amino acid based on the current amino acid features. The generated amino acid is then input into the LSM again to generate the next amino acid, and so on, until the set number of amino acids is generated. These generated amino acids are then concatenated along the sequence length dimension. The above steps are repeated until the number of viral amino acid sequences reaches the preset number, at which point the specific amino acid sequence generation module stops generating viral amino acid sequences.
[0048] This invention also provides a viral amino acid sequence generation and screening device, including a parameter setting module, a specific amino acid sequence generation module, a viral amino acid sequence production fitness prediction module, a virus scoring module, and a target viral amino acid sequence library, the specific connection structure of which is as follows:
[0049] The parameter setting module is used to set the length and number of viral amino acid sequences generated, and the parameters are input into the specific amino acid sequence generation module.
[0050] The specific amino acid sequence generation module generates viral amino acid sequences by learning the logical structure and syntactic and semantic features between existing viral amino acid sequences; the generated viral amino acid sequences are then input into the viral amino acid sequence production fitness prediction module.
[0051] The viral amino acid sequence production fitness prediction module receives the viral amino acid sequence and performs viral amino acid sequence production fitness prediction on the viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness and a judgment on whether the viral amino acid sequence is real.
[0052] The virus scoring module scores the viral amino acid sequence based on the predicted viral amino acid sequence production fitness and the judgment of whether the viral amino acid sequence is authentic.
[0053] The target virus amino acid sequence library is sorted according to the score of the viral amino acid sequence, and the top P amino acid sequences with the highest score values are selected and saved, where P is a positive integer;
[0054] In a preferred embodiment of the present invention, the specific amino acid sequence generation module operates as follows:
[0055] An amino acid is randomly generated, its features are characterized, and then these features are input into a long short-term memory neural network.
[0056] The Long Short-Term Memory Neural Network generates viral amino acid sequences based on a set amino acid sequence length. Specifically, after inputting the amino acid features into the trained Long Short-Term Memory Neural Network, the Neural Network generates the next amino acid based on the current amino acid features. The generated amino acid is then input into the Neural Network again to generate the next amino acid at the next position, and so on, until the set number of amino acids is generated. Finally, these generated amino acids are spliced together along the sequence length dimension.
[0057] Repeat the above steps. Once the number of viral amino acid sequences reaches the preset number, the specific amino acid sequence generation module stops generating viral amino acid sequences.
[0058] This invention also provides a system for implementing the machine learning-based viral amino acid sequence generation and screening method described in the claims, comprising four main modules: a cloud computing and supercomputing platform, a viral vector design and development laboratory, a viral amino acid sequence generation and screening device, and an algorithm result verification laboratory; wherein:
[0059] The cloud computing and supercomputing platform accepts operation instructions from users or administrators through the I / O interface and assigns them corresponding permissions. It is responsible for collecting and managing online information data related to virus sequence design, transmitting local experimental data into the storage unit, and allocating the available computing resources to the computing tasks submitted by users according to priority through the computing unit to execute the corresponding tasks.
[0060] The viral vector design and development laboratory is used to obtain viral amino acid sequences experimentally, as well as the fitness of each viral amino acid sequence.
[0061] A viral amino acid sequence generation and screening device is used to generate a target viral amino acid sequence library with high specificity and high viral amino acid sequence production fitness, including: a parameter setting module, a specific amino acid sequence generation module, a viral amino acid sequence production fitness prediction module, a virus scoring module, and a target viral amino acid sequence library.
[0062] In the algorithm result verification laboratory, the generation fitness of the corresponding viral amino acid sequences stored in the target viral amino acid sequence library was obtained through experiments. The fitness was then compared with the production fitness predicted by the production fitness prediction module to verify the effectiveness of the viral amino acid sequence generation and screening device.
[0063] Compared with the prior art, the present invention has the following contributions:
[0064] 1. In the process of generating amino acid sequences, the logical structure and syntactic and semantic features between sequences are extracted based on the method of "Natural Language Processing" (also referred to as NLP in this invention). Amino acid sequence samples of a specific length are generated step by step according to the time step, so that the generated sequence samples are closer to the real sequence in terms of feature distribution and information, and learn the logical order information unique to time series. Here, natural language processing is a metaphor. Specifically, it means that, for example, "I like to eat apples" in Chinese is a sentence that can express a complete meaning, which includes six language elements (Chinese characters, or words, etc.) such as I, like, like, eat, apple, and fruit. If these six language elements are randomly arranged and combined, hundreds of combinations can be obtained. However, as we all know, in the context of Chinese, there are not many ways to combine the language elements of I, like, like, eat, apple, and fruit into combinations with actual meaning, such as "I like to eat apples", "I like to eat apples", "I like to eat apples", etc. If we consider the viral capsid amino acid sequence as sentences and viral amino acids as language elements, we can assume that the existing known amino acid sequences capable of forming capsids—these "sentences"—exist according to certain specific rules (analogous to the grammar of natural language). If this "grammar" is mastered during sequence generation, the computational load can undoubtedly be greatly reduced. This invention demonstrates that first learning this "grammar" through artificial intelligence, and then using this "grammar" to guide amino acid sequence generation, is highly effective in improving computational speed and the adaptability of viral amino acid sequence production (see [reference]). Figure 9 and Figure 10 );
[0065] 2. Viral amino acid sequences are designed from two aspects: virus specificity and high specificity with high adaptability to viral amino acid sequence production. This ensures sequence diversity while enhancing the feasibility of virus construction, enabling the model to generate more meaningful sequence data and improve the efficiency of viral amino acid sequence screening.
[0066] 3. Apply the lexical representation method to represent the viral amino acid sequence information, so as to learn the semantic and syntactic information within the viral amino acid sequence more comprehensively and enhance the model's ability to extract and represent sequence features.
[0067] 4. In terms of model design, we first used a more advanced generative adversarial network (GAN) with highly diverse generated sequences as the basic idea of the model. At the same time, we considered the logical order of sequence information, which enabled the model to learn the semantic and syntactic information of the amino acid sequence. Compared with the traditional noise generation method, it generated the viral amino acid sequence with high specificity and high viral amino acid sequence production fitness in a more logical way. Then, we performed viral amino acid sequence production fitness prediction on the generated viral amino acid sequence with high specificity and high viral amino acid sequence production fitness, and selected the viral amino acid sequence with high score for experimental verification. Attached Figure Description
[0068] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for initial preferred embodiment purposes only and are not intended to limit the embodiments of the invention. Furthermore, the same reference numerals denote the same parts throughout all the drawings. In the drawings:
[0069] Figure 1 An overview diagram of the machine learning-based viral amino acid sequence generation and screening method and system provided in the embodiments of the present invention is shown;
[0070] Figure 2 This illustrates a flowchart of constructing a dataset for training a model from experimental data, as provided in an embodiment of the present invention.
[0071] Figure 3 This illustrates the steps for training a viral amino acid sequence generation and screening device using a dataset, as provided in an embodiment of the present invention.
[0072] Figure 4 The present invention illustrates the steps of generating and screening viral amino acid sequences using the viral amino acid sequence generation and screening apparatus provided in this embodiment.
[0073] Figure 5 The following steps for verifying viral amino acid sequences in a target viral amino acid sequence library, as provided in an embodiment of the present invention, are illustrated:
[0074] Figure 6 This embodiment of the invention illustrates the distribution of viral amino acid sequences that generate fitness probability density when constructing a dataset for training a model using experimental data.
[0075] Figure 7 This invention illustrates a Pearson correlation heatmap showing the frequency of viral amino acid sequences in multiple plasmid replication experiments and multiple virus replication experiments when constructing a dataset for training a model using experimental data.
[0076] Figure 8 This embodiment of the invention illustrates a Pearson correlation analysis of the viral amino acid sequence production fitness predicted by the viral amino acid sequence production fitness prediction module and the actual viral amino acid sequence production fitness.
[0077] Figure 9 This illustration shows the spatial distribution of viral amino acid sequences generated according to time steps based on time series logical structure information features, sequences generated from spatial noise using traditional methods, and real viral amino acid sequence data after tsne dimensionality reduction. In the figure, NLP represents the generation of viral amino acid sequences according to time steps using natural language processing, and CV represents the generation of sequences from spatial noise using image processing.
[0078] Figure 10 This illustration shows the spatial distribution of viral amino acid sequences generated using VAE (Variable Differential Autoencoder) in an embodiment of the present invention, as well as the data of viral amino acid sequences generated according to time steps and real viral amino acid sequences after dimensionality reduction by tsne. In the figure, NLP represents the generation of viral amino acid sequences according to time steps using natural language processing. Detailed Implementation
[0079] To better understand the technical solution of the present invention, the specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The same reference numerals in the drawings indicate elements with the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0080] This invention proposes a method for generating and screening viral amino acid sequences based on machine learning, such as... Figure 1 As shown, a viral amino acid sequence generation method based on generative adversarial networks (GANs) generates and experimentally verifies viral amino acid sequences based on viral production fitness information, aiming to find viral amino acid sequences with high transfection efficiency and high viral formation ability. The specific steps are as follows:
[0081] S1, constructing a dataset from experimental data for training the model, such as... Figure 2 As shown, it includes the following steps:
[0082] S11, Virus plasmid library construction steps
[0083] M amino acid sequences (M ranging from 10,000 to 200,000) of length N (N ranging from 1 to 100) are randomly generated. Each amino acid sequence is linked with a specific barcode. All amino acid sequences linked with barcodes are pooled together to construct an amino acid sequence pool. The amino acid sequences in the pool are used to replace certain sites on the target plasmid to obtain viral plasmids. Different amino acid sequences or substitutions of different sites on the target plasmid will yield different viral plasmids. These different viral plasmids together constitute a viral plasmid library.
[0084] Each amino acid sequence is linked to a specific barcode for subsequent high-throughput sequencing.
[0085] The viral plasmids in the current viral plasmid library are only potential viral plasmids. Further testing, including viral vector packaging, viral purity detection, and in vitro biological activity detection, is needed to determine whether the obtained viral plasmids can ultimately become viruses.
[0086] S12, Viral amino acid sequence production fitness data collection steps
[0087] The viral plasmid frequency was obtained by extracting plasmid DNA from the viral plasmid and performing high-throughput sequencing. Specifically, the frequency of a single viral amino acid sequence in the plasmid was calculated by extracting plasmid DNA from the replaced plasmid and performing high-throughput sequencing.
[0088] Simultaneously, after transfecting the viral plasmid into cells, the virus was purified, and the viral DNA was extracted. The viral frequency was then calculated using high-throughput sequencing. The specific process is as follows: the viral plasmid was transfected into cells, and the virus was produced (i.e., the virus spread) in the cells. After purification, the virus purity was tested, and the viral DNA was extracted. The viral frequency was then calculated using high-throughput sequencing. The viral frequency refers to the number of times the same viral amino acid sequence appears in the virus after high-throughput sequencing.
[0089] Finally, the production fitness of a single viral amino acid sequence is calculated using viral plasmid frequency and viral frequency; production fitness is a quantitative representation of the ability of a viral amino acid sequence to generate virus. A higher production fitness indicates a stronger ability of the amino acid sequence to generate virus.
[0090] S13 Constructing the Dataset
[0091] The dataset uses viral amino acid sequences as samples, and the production fitness of the corresponding viral amino acid sequences as the dataset labels. In this embodiment, viral amino acid sequences containing premature stop codons and those with sequencing errors discovered during high-throughput sequencing are mainly used. The samples in the dataset are divided into training set: validation set: test set in a ratio of 7:1:2.
[0092] To avoid imbalanced training data and ensure the accuracy and research significance of the model training process, it is necessary to first check the normality of the overall distribution of the viral amino acid sequence production fitness data labels. For imbalanced data, strategies such as normalization, downsampling, and gradient pruning should be used to balance the data to ensure that the model is unbiased during the learning process. Figure 6 This describes the probability density distribution of fitness generated by viral amino acid sequences when constructing a dataset for training the model from experimental data.
[0093] S2 uses a dataset to train the virus amino acid sequence generation and screening device;
[0094] The specific amino acid sequence generation module and the viral amino acid sequence production fitness prediction module in the viral amino acid sequence generation and screening device were trained separately, such as... Figure 3 As shown, it includes the following steps:
[0095] S21 performs feature encoding on the amino acid sequences in the dataset.
[0096] The viral amino acid sequences in the dataset are input into the data preprocessing module to encode their features, resulting in viral amino acid sequence features. These features include the characteristics of each amino acid in the viral amino acid sequence. The data preprocessing module can employ existing strategies such as word embedding, graph representation, or positional encoding to characterize the feature information of the viral amino acid sequences.
[0097] S22 trains the specific amino acid sequence generation module;
[0098] There are currently 20 known amino acids. The amino acid sequence is composed of combinations of amino acids, but only some of these combinations can generate a virus.
[0099] The specific amino acid sequence generation module generates viral amino acid sequences by learning the logical structure and syntactic-semantic features of existing viral amino acid sequences. Specificity refers to conforming to the logical structure and syntactic-semantic features of viral amino acid sequences.
[0100] In this embodiment, the specific amino acid sequence generation module employs a long short-term memory (LSTM) neural network. The training steps for the specific amino acid sequence generation module are as follows: Viral amino acid sequences are selected from the training set according to a set batch size, where the batch size refers to the number of viral amino acid sequences that can be input into the specific amino acid sequence generation module at one time; after text encoding of the viral amino acid sequences, a zero is appended to the beginning of each viral amino acid sequence, and the last amino acid is truncated. All features of the processed viral amino acid sequences are used as input to the LTM neural network. The original viral amino acid sequences, after text encoding and without further processing, are used as the targets to be generated by the LTM neural network. The internal workflow of the LTM neural network is that an input viral amino acid sequence is split into multiple single... The process involves generating a specific amino acid sequence from a dataset containing N viral amino acid sequences. The previous amino acid is used as input to a long short-term memory (LSM) neural network to generate the next amino acid sequence. This process continues until an amino acid sequence matching the target sequence length is generated. The generated amino acid sequence is then input into a linear layer for feature extraction. The extracted features are then passed to a softmax activation function to obtain the final amino acid sequence. The selected viral amino acid sequence features are compared with the generated amino acid sequence features, a loss function is calculated, and backpropagation is performed. This process is repeated M times until the loss function stabilizes, and the model training parameters are saved. Therefore, when training the specific amino acid sequence generation module using a dataset containing N viral amino acid sequences, it requires (N / batch size)*M iterations before training concludes.
[0101] Where 1≤N≤50000000, N is a positive integer, and 1≤M≤50000, M is a positive integer.
[0102] S23 Training Virus Amino Acid Sequence Production Fitness Prediction Module
[0103] The viral amino acid sequence production fitness prediction module can predict the viral amino acid sequence production fitness based on the viral amino acid sequence. The higher the production fitness, the stronger the ability of the amino acid sequence to generate virus.
[0104] The training steps for the viral amino acid sequence production fitness prediction module are as follows: Feature information of a viral amino acid sequence is simultaneously input into multiple convolutional blocks in parallel. Features extracted from these convolutional blocks are concatenated in the hidden layer. A residual module is used to maintain the original features with a certain probability while updating the features using convolutional layers. The previously extracted features are then input into the residual module, followed by layer normalization (LN). Feature aggregation is performed using Normalization to obtain aggregated features. The aggregated features are then weighted using the Sigmoid activation function. Simultaneously, the aggregated features are delinearized using the ReLU activation function to obtain activation features. The sequence information weights are multiplied by the activation features to obtain weighted activation information. The sequence information weights are subtracted from the previously obtained weights and multiplied by the previous viral amino acid sequence features to obtain the original weight information. The weighted activation information and the original weight information are added together to form the prediction features. This allows the original feature information to be passed in with a certain probability while updating the features. The prediction features are then input into a temporary fallback method and then into a linear layer. Finally, the LeakyReLU activation function is used as the output for predicting the fitness of the viral amino acid sequence production. The prediction features are then input into a temporary fallback method and then into a linear layer, where the Sigmoid activation function is used as the output for predicting whether the viral sequence is real. After Q iterations (1 ≤ Q ≤ 50000, where Q is a positive integer), the training of the viral amino acid sequence production fitness prediction module is completed.
[0105] S3 uses a trained viral amino acid sequence generation and screening device to generate and screen viral amino acid sequences, such as... Figure 4 As shown, the specific steps are as follows:
[0106] S31, After setting the viral amino acid sequence length and generation quantity, input the parameters into the specific amino acid sequence generation module;
[0107] S32, the execution steps of the specific amino acid sequence generation module are as follows: randomly generate an amino acid, characterize the randomly generated amino acid, input the amino acid features into a trained long short-term memory neural network, and the long short-term memory neural network generates a viral amino acid sequence according to the set amino acid sequence length. Specifically, after inputting the amino acid features into the trained long short-term memory neural network, the long short-term memory neural network will generate the next amino acid according to the current amino acid features. The generated amino acid will then be input into the long short-term memory neural network to generate the next amino acid at the next position, and so on until the set number of amino acids is generated. Then, these generated amino acids are spliced together in the dimension of sequence length. The above steps are repeated. When the number of viral amino acid sequences reaches the preset number of generated sequences, the specific amino acid sequence generation module stops generating viral amino acid sequences.
[0108] All generated viral amino acid sequences are sent to the viral amino acid sequence production fitness prediction module.
[0109] Besides the virus-specific sequence generation module described in this embodiment, which uses a long short-term memory neural network with a generative adversarial network architecture to generate viral amino acid sequences, reliable sequence generation algorithms such as variable differential autoencoders (VAEs) and diffusion models can also be used to generate virus-specific sequences. However, the VAE architecture consists of an encoder and a decoder. The encoder compresses the original data into low-dimensional feature vectors through a neural network, while the decoder restores the compressed feature vectors through a neural network and generates data that conforms to the original distribution. Although the generated data has high diversity, the generated results are usually rather ambiguous. Without the adversarial process of a discriminator, the quality is difficult to guarantee, and there are problems such as the posterior distribution being assumed to be a decomposable Gaussian distribution, which is based on strong assumptions of the encoder. The diffusion model includes a diffusion process and a reverse diffusion process. Under the condition of a Markov chain, the diffusion process gradually tends to a Gaussian distribution by continuously adding Gaussian noise to the original data, eventually obtaining globally diffused data. The reverse diffusion process restores the globally diffused data through Gaussian noise and generates the original data. The generative model based on this architecture can generate relatively robust sequence information, but due to current technological limitations, the generalization ability of the generated data still needs to be improved. Therefore, the virus-specific sequence generation module in this embodiment is the most suitable choice.
[0110] S33, after receiving all generated viral amino acid sequences, the viral amino acid sequence production fitness prediction module uses the trained module to predict the production fitness of each sequence, obtaining the corresponding production fitness and a judgment on whether the sequence is genuine.
[0111] S34, Virus scoring module
[0112] The viral amino acid sequences are scored based on their predicted production fitness and the authenticity of the sequences. The authenticity assessment is primarily used during model training to constrain the logical relationships learned by the generator among real viral amino acid sequences. The score for sequence authenticity is limited to 0-1; a score of 0.5 or higher indicates a real viral amino acid sequence, while a score less than 0.5 indicates a fake sequence. In the scoring module, both viral production fitness and sequence authenticity are considered. Only sequences assessed as real by the model are selected, and then sorted in descending order of production fitness. The top N high-scoring sequences are then used for biosynthesis experiments to verify their authenticity.
[0113] S35, generate the target virus amino acid sequence library.
[0114] The viral amino acid sequences are sorted according to their scores, and the top P amino acid sequences with the highest scores are selected to form the target viral amino acid sequence library, where P is a positive integer.
[0115] S4, verify the viral amino acid sequences in the target viral amino acid sequence library;
[0116] S41, Experiments were conducted on the viral amino acid sequences in the target viral amino acid sequence library to obtain experimental data, such as... Figure 5 As shown:
[0117] If the number of viral amino acid sequences in the target viral amino acid sequence library is greater than 100, each viral amino acid sequence in the library is barcoded to construct an amino acid sequence pool. If the number of viral amino acid sequences is less than or equal to 100, a viral vector is constructed and packaged separately for each amino acid sequence. After replacing appropriate sites on the target plasmid with viral amino acid sequences, a viral plasmid is obtained. The viral plasmid frequency is calculated by high-throughput sequencing of plasmid DNA extracted from the viral plasmid. Simultaneously, the viral plasmid is transfected into cells for viral production in cells, followed by viral purification. After viral purification, viral purity is tested. Viral DNA is extracted, and viral frequency is calculated by high-throughput sequencing. In vitro and in vivo biological activity tests are then performed. The viral amino acid sequence production fitness is calculated based on the viral plasmid frequency and viral frequency. This fitness is compared with the predicted viral amino acid sequence production fitness using evaluation indicators to verify the effectiveness of the viral amino acid sequence generation and screening device. Finally, a viral vector with high specificity and high viral amino acid sequence production fitness, validated by biological experiments, is obtained.
[0118] S42, Determine the evaluation indicators.
[0119] In selecting evaluation metrics to assess the predictive performance of the model, this study involves a regression prediction task; therefore, the root mean square error, Pearson correlation coefficient, Spearman correlation coefficient, and coefficient of determination R0 were chosen. 2 To evaluate the performance of viral nucleocapsid production fitness data prediction; the root mean square error describes the distance between the predicted and actual values; the Pearson correlation coefficient and the Spearman correlation coefficient describe the correlation between the predicted and actual values, where the Pearson correlation coefficient describes the linear correlation between the two values, and the Spearman correlation coefficient is the rank form of the Pearson correlation coefficient, which describes the correlation between the two variables (e.g., when one variable increases, the other variable also increases), and it is related to the monotonicity of the function; the coefficient of determination R... 2It is a dimensionless score that describes the effectiveness of the model, comparing the predictions to random guesses based on the average of the true values;
[0120] Figure 7-10 This is an example showing some actual measured metrics;
[0121] Figure 7 This is a Pearson correlation heatmap showing the frequency of viral amino acid sequences in multiple plasmid replication experiments and multiple virus replication experiments when constructing a dataset from experimental data for training the model.
[0122] Figure 8 The Pearson correlation analysis of the viral amino acid sequence production fitness prediction module and the actual viral amino acid sequence production fitness is performed.
[0123] Figure 9 It compares the spatial distribution of viral amino acid sequences generated according to time steps based on the logical structure information features of time series with that generated from spatial noise using traditional methods and the spatial distribution of real viral amino acid sequence data after tsne dimensionality reduction.
[0124] Figure 10 This is the spatial distribution of viral amino acid sequences generated using VAE, viral amino acid sequences generated according to time steps, and real viral amino acid sequence data after dimensionality reduction by tsne.
[0125] This invention also discloses a viral amino acid sequence generation and screening device, comprising a parameter setting module, a specific amino acid sequence generation module, a viral amino acid sequence production fitness prediction module, a virus scoring module, and a target viral amino acid sequence library, the specific connection structure of which is as follows:
[0126] The parameter setting module is used to set the length and number of viral amino acid sequences generated, and the parameters are input into the specific amino acid sequence generation module.
[0127] The specific amino acid sequence generation module generates viral amino acid sequences by learning the logical structure and syntactic and semantic features between existing viral amino acid sequences. The generated viral amino acid sequences are then input into the viral amino acid sequence production fitness prediction module. The working process of the specific amino acid sequence generation module is as follows: a random amino acid is generated, its features are characterized, and these features are input into a trained long short-term memory (LSTM) neural network. The LSM generates viral amino acid sequences according to a set amino acid sequence length. Specifically, after inputting the amino acid features into the trained LSM, the LSM generates the next amino acid based on the current amino acid features. This generated amino acid is then input into the LSM again to generate the next amino acid at the next position, and so on, until a set number of amino acids of a predetermined length are generated. These generated amino acids are then concatenated along the sequence length dimension. This process is repeated until the number of viral amino acid sequences reaches the preset generation limit, at which point the specific amino acid sequence generation module stops generating viral amino acid sequences.
[0128] The viral amino acid sequence production fitness prediction module receives the viral amino acid sequence and performs viral amino acid sequence production fitness prediction on the viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness and a judgment on whether the viral amino acid sequence is real.
[0129] The virus scoring module scores the viral amino acid sequence based on the predicted viral amino acid sequence production fitness and the judgment of whether the viral amino acid sequence is authentic.
[0130] The target virus amino acid sequence library is sorted according to the score of the viral amino acid sequence, and the top P amino acid sequences with the highest score values are selected and saved, where P is a positive integer;
[0131] This invention also discloses a system for analyzing viral amino acid sequences based on machine learning, such as... Figure 1 As shown, it comprises four main modules: a cloud computing and supercomputing platform, a virus vector design and development laboratory, a virus amino acid sequence generation and screening device, and an algorithm result verification laboratory; among which:
[0132] The cloud computing and supercomputing platform accepts operation instructions from users or administrators through the I / O interface and assigns them corresponding permissions. It is responsible for collecting and managing online information data related to virus sequence design, transmitting local experimental data into the storage unit, and allocating the available computing resources to the computing tasks submitted by users according to priority through the computing unit to execute the corresponding tasks.
[0133] The viral vector design and development laboratory is used to obtain viral amino acid sequences experimentally, as well as the fitness of each viral amino acid sequence.
[0134] A viral amino acid sequence generation and screening device for generating a library of target viral amino acid sequences that are highly specific and have high adaptability for viral amino acid sequence production.
[0135] In the algorithm result verification laboratory, the generation fitness of the corresponding viral amino acid sequences stored in the target viral amino acid sequence library was obtained through experiments. The fitness was then compared with the production fitness predicted by the production fitness prediction module to verify the effectiveness of the viral amino acid sequence generation and screening device.
Claims
1. A method for generating and screening viral amino acid sequences based on machine learning, characterized in that, The specific steps are as follows: S1, Construct a dataset from experimental data to train the model; S2, Training the Virus Amino Acid Sequence Generation and Screening Device Using a Dataset; including the following steps: S21 encodes the amino acid sequences in the dataset using features; S22 trains the specific amino acid sequence generation module; The specific amino acid sequence generation module generates viral amino acid sequences by learning the logical structure and syntactic and semantic features between existing amino acid sequences that can generate viruses. The specific steps are as follows: select viral amino acid sequences from the training set according to the set batch number; after text encoding the viral amino acid sequences, concatenate a number 0 in front of each viral amino acid sequence and truncate the last amino acid. All features of the processed viral amino acid sequence are used as input to the Long Short-Term Memory (LSTM) neural network. The original viral amino acid sequence, after text encoding and without further processing, is used as the target to be generated by the LTM neural network. The internal workflow of the LTM neural network is as follows: an input viral amino acid sequence is split into multiple individual amino acids. The previous amino acid is used as input to the LTM neural network to generate the amino acid at the next position. The next generated amino acid is used as input to the LTM neural network to generate the amino acid at the next position, and so on, until an amino acid sequence with the same length as the target amino acid sequence is generated. The generated amino acid sequence is input into a linear layer for amino acid sequence feature extraction. The extracted features are then input into the softmax activation function to obtain the final generated amino acid sequence. The selected viral amino acid sequence features are compared with the generated amino acid sequence features, the loss function is calculated, and backpropagation is performed. The above amino acid generation steps are completed by iteratively selecting viral amino acid sequences from the entire training set according to the batch size. The above cycle is repeated M times until the loss function is stable, where M is a positive integer. The model training parameters are then saved. S23 Training Virus Amino Acid Sequence Production Fitness Prediction Module; The viral amino acid sequence production fitness prediction module predicts the viral amino acid sequence production fitness based on the viral amino acid sequence. The higher the production fitness, the stronger the ability of the amino acid sequence to generate virus. S3 uses a trained viral amino acid sequence generation and screening device to generate and screen viral amino acid sequences. The specific steps are as follows: S31, After setting the viral amino acid sequence length and generation quantity, input the parameters into the specific amino acid sequence generation module; S32, the specific amino acid sequence generation module generates a viral amino acid sequence of a preset number and length based on a randomly generated amino acid; S33, After receiving all the generated viral amino acid sequences, the viral amino acid sequence production fitness prediction module performs viral amino acid sequence production fitness prediction on each viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness and the judgment of whether the viral amino acid sequence is real. S34, Virus scoring module The viral amino acid sequence is scored based on the predicted viral amino acid sequence production fitness and the judgment of whether the viral amino acid sequence is true. S35, generate the target virus amino acid sequence library. The viral amino acid sequences are sorted according to their scores, and the top P amino acid sequences with the highest scores are selected to form the target viral amino acid sequence library, where P is a positive integer. S4, verify the viral amino acid sequences in the target viral amino acid sequence library; S41, Experiments were conducted on viral amino acid sequences in the target viral amino acid sequence library to obtain experimental data. The production fitness of the experimentally obtained viral amino acid sequences was compared with the predicted production fitness of the viral amino acid sequences using evaluation indicators to verify the effectiveness of the viral amino acid sequence generation and screening device in generating viral amino acid sequences.
2. The method for generating and screening viral amino acid sequences based on machine learning according to claim 1, characterized in that, S1 constructs a dataset for training the model from the experimental data, specifically including the following steps; S11, Steps for constructing a viral plasmid library H amino acid sequences of length L are randomly generated. Each amino acid sequence is linked with a specific barcode. All amino acid sequences linked with barcodes are pooled together to construct an amino acid sequence pool. The amino acid sequence library in the amino acid sequence pool is used to replace some sites of the target plasmid to obtain a viral plasmid. Different viral plasmids constitute a viral plasmid library. S12, Viral amino acid sequence production fitness data collection steps The frequency of viral plasmids was calculated by high-throughput sequencing of plasmid DNA extracted from viral plasmids. Simultaneously, after transfecting the viral plasmid into the cells, the virus was purified, and the viral DNA was extracted and the viral frequency was obtained by high-throughput calculation. Finally, the fitness of a single viral amino acid sequence is calculated using viral plasmid frequency and viral frequency. S13 constructs the dataset. The dataset uses amino acid sequences that can become viruses as samples, and the corresponding amino acid sequences generate fitness as labels for the dataset.
3. The method for generating and screening viral amino acid sequences based on machine learning according to claim 1, characterized in that, The S23 training virus amino acid sequence production fitness prediction module specifically includes the following steps; The viral amino acid sequence features are simultaneously input into multiple convolutional blocks in parallel. The features extracted from these blocks are concatenated in the hidden layer. A residual module is used to connect the network with a certain probability to preserve the original features while updating the features using convolutional layers. The previously extracted features are then input into the residual module, and layer normalization (LN) is used to converge the features to obtain converged features. The converged features are then processed using the sigmoid activation function to obtain sequence information weights. Simultaneously, the converged features are delinearized using the ReLU activation function to obtain activation features. The sequence information weights are multiplied by the activation features to obtain weighted activation information. The sequence information weights are subtracted from the previously obtained sequence information weights and multiplied by the previous viral amino acid sequence features to obtain the original weight information. The weighted activation information and the original weight information are added together to form the prediction features. The prediction features are then input into the temporary fallback method and then into the linear layer. Finally, the leaky ReLU activation function is used as the output for predicting the fitness of the viral amino acid sequence. The prediction features are then input into the temporary fallback method and then into the linear layer, where the sigmoid activation function is used as the output for predicting whether the viral sequence is real. After Q iterations, where Q is a positive integer, the training of the viral amino acid sequence production fitness prediction module ends.
4. The method for generating and screening viral amino acid sequences based on machine learning according to claim 1, characterized in that, The execution steps of the S32 specific amino acid sequence generation module specifically include the following steps; An amino acid is randomly generated, and its features are characterized. The amino acid features are then input into a trained long short-term memory neural network. The Long Short-Term Memory (LSTM) neural network generates viral amino acid sequences based on a set amino acid sequence length. Specifically, after inputting amino acid features into the trained LSM, the LSM generates the next amino acid based on the current amino acid features. The generated amino acid is then input into the LSM again to generate the next amino acid, and so on, until the set number of amino acids is generated. These generated amino acids are then concatenated along the sequence length dimension. The above steps are repeated until the number of viral amino acid sequences reaches the preset number, at which point the specific amino acid sequence generation module stops generating viral amino acid sequences.
5. A viral amino acid sequence generation and screening device, comprising a parameter setting module, a specific amino acid sequence generation module, a viral amino acid sequence production fitness prediction module, a virus scoring module, and a target viral amino acid sequence library, the specific connection structure of which is as follows: The parameter setting module is used to set the length and number of viral amino acid sequences generated, and the parameters are input into the specific amino acid sequence generation module. The specific amino acid sequence generation module generates viral amino acid sequences by learning the logical structure and syntactic and semantic features between existing viral amino acid sequences, and inputs the generated viral amino acid sequences into the viral amino acid sequence production fitness prediction module. The viral amino acid sequence production fitness prediction module receives the viral amino acid sequence and performs viral amino acid sequence production fitness prediction on the viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness and a judgment on whether the viral amino acid sequence is real. The virus scoring module scores the viral amino acid sequence based on the predicted viral amino acid sequence production fitness and the judgment of whether the viral amino acid sequence is authentic. The target virus amino acid sequence library is sorted according to the score of the viral amino acid sequence, and the top P amino acid sequences with the highest score values are selected and saved, where P is a positive integer; in, The specific amino acid sequence generation module is pre-trained through the following steps: selecting viral amino acid sequences from the training set according to a set batch size; after text encoding the viral amino acid sequences, prepending a 0 to the beginning of each viral amino acid sequence and truncating the last amino acid; using all features of the processed viral amino acid sequences as input to the Long Short-Term Memory (LSTM) neural network; and using the unprocessed viral amino acid sequences after text encoding as the target to be generated by the LTM neural network. The internal workflow of the LTM neural network is as follows: an input viral amino acid sequence is split into multiple individual amino acids, and the previous amino acid is used as input to the LTM neural network to generate the next sequence. The first amino acid is used as input to the Long Short-Term Memory (LSTM) neural network to generate the next amino acid at the next position, and so on, until an amino acid sequence with the same length as the target amino acid sequence is generated. The generated amino acid sequence is then input into a linear layer for amino acid sequence feature extraction. The extracted features are then input into the softmax activation function to obtain the final generated amino acid sequence. The selected viral amino acid sequence features are compared with the generated amino acid sequence features, the loss function is calculated, and backpropagation is performed. The entire viral amino acid sequence in the training set is selected cyclically according to the batch size to complete the above amino acid generation steps. The above loop is repeated M times until the loss function is stable, where M is a positive integer. The model training parameters are then saved.
6. The viral amino acid sequence generation and screening device according to claim 5, characterized in that, The specific amino acid sequence generation module works as follows: An amino acid is randomly generated, its features are characterized, and then these features are input into a long short-term memory neural network. The Long Short-Term Memory Neural Network generates viral amino acid sequences based on a set amino acid sequence length. Specifically, after inputting the amino acid features into the trained Long Short-Term Memory Neural Network, the Neural Network generates the next amino acid based on the current amino acid features. The generated amino acid is then input into the Neural Network again to generate the next amino acid at the next position, and so on, until the set number of amino acids is generated. Finally, these generated amino acids are spliced together along the sequence length dimension. Repeat the above steps. Once the number of viral amino acid sequences reaches the preset number, the specific amino acid sequence generation module stops generating viral amino acid sequences.
7. A system for generating and screening viral amino acid sequences based on machine learning, comprising four main modules: a cloud computing and supercomputing platform, a viral vector design and development laboratory, a viral amino acid sequence generation and screening device, and an algorithm result verification laboratory, for implementing the viral amino acid sequence generation and screening method described in any one of claims 1-4; wherein: The cloud computing and supercomputing platform accepts operation instructions from users or administrators through the I / O interface and assigns them corresponding permissions. It is responsible for collecting and managing online information data related to virus sequence design, transmitting local experimental data into the storage unit, and allocating the available computing resources to the computing tasks submitted by users according to priority through the computing unit to execute the corresponding tasks. The viral vector design and development laboratory is used to obtain viral amino acid sequences experimentally, as well as the fitness of each viral amino acid sequence. A viral amino acid sequence generation and screening device is used to generate a target viral amino acid sequence library with high specificity and high viral amino acid sequence production fitness, including: a parameter setting module, a specific amino acid sequence generation module, a viral amino acid sequence production fitness prediction module, a virus scoring module, and a target viral amino acid sequence library. In the algorithm result verification laboratory, the generation fitness of the corresponding viral amino acid sequences stored in the target viral amino acid sequence library was obtained through experiments. The fitness was then compared with the production fitness predicted by the production fitness prediction module to verify the effectiveness of the viral amino acid sequence generation and screening device.
Citation Information
Patent Citations
Protein domain detection method based on cost-sensitive LSTM network
CN106295242A
Automatic analysis method and system for virus sequencing sequence
CN112863599A