A method, system and storage medium for designing AAV2 capsid protein variants

By using pre-trained and fine-tuned Progen2 model and fine-tuned AntiBERTy model system, high-quality AAV2 capsid protein variant sequences are generated and screened out, which solves the shortcomings of the AAV2 capsid protein variant design method in the prior art and achieves more efficient gene therapy effects.

CN119380820BActive Publication Date: 2025-06-17WEST CHINA HOSPITAL SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411411219.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-10
Publication Date
2025-06-17
Estimated Expiration
2044-10-10

AI Technical Summary

Technical Problem

The existing AAV2 capsid protein variant design methods have shortcomings in generating high-quality and diverse protein sequences, which cannot effectively overcome the host immune response, resulting in poor therapeutic effects of gene therapy.

Method used

A system based on pre-trained fine-tuned Progen2 model and fine-tuned AntiBERTy model was used to generate and screen high-quality AAV2 capsid protein variant sequences through feature splicing and functional scoring.

Benefits of technology

It improves the homology and diversity of the AAV2 capsid protein sequence, improves the system's prediction and classification performance, effectively overcomes the host immune response, and improves the effect of gene therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380820B_ABST
    Figure CN119380820B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of gene therapy, and particularly relates to a method, system, and storage medium for designing AAV2 capsid protein variants. The system of the present invention includes a preprocessing module configured to preprocess the sequence data of the collected AAV2 capsid protein; an input module configured to input AAV2 capsid protein sequences of different sequence lengths; a prediction module configured to classify the AAV2 capsid protein sequences through a fine-tuned AntiBERTy model and obtain AAV2 capsid protein variant sequences according to the functional scores of the sequences; and an output module configured to output the sequences of the prediction module. The present invention can generate high-quality protein sequences of different lengths and with functionality, significantly improve the homology and diversity of the generated sequences, accurately predict the functional characteristics of the sequences, and thus effectively screen AAV2 capsid protein variants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of gene therapy, and particularly relates to a method, system and storage medium for designing AAV2 capsid protein variants. Background Art

[0002] As a revolutionary treatment method, gene therapy has shown great potential in the treatment of various diseases such as malignant tumors and genetic diseases in recent years. The key step of gene therapy is to design and construct gene vectors for safely and effectively delivering target genes into host cells to ensure the correct expression of genes in cells and play a therapeutic role. Adeno-associated virus (AAV) is the simplest known non-enveloped single-stranded DNA virus, with a genome length of about 4.7 Kb and belonging to the parvovirus family. Due to the non-pathogenicity and wide host range of AAV, and the fact that it does not cause strong immune responses in the human body, it has now become one of the most important gene vectors in the field of gene therapy. The AAV2 capsid is a component of the first gene therapy approved by the US Food and Drug Administration for human use.

[0003] The key requirement of gene vectors is to be able to resist the human immune defense mechanism. However, natural AAV has some drawbacks. Many human sera have the ability to neutralize the virus, which can prevent the virus from further infecting host cells, resulting in the failure of the vector to successfully deliver drugs or target genes to the designated location and causing treatment failure, posing challenges to gene therapy. To overcome this problem, it is necessary to design new and diverse AAV variants, which can alleviate the problem of natural immunity to a certain extent and is of great significance for gene therapy. Currently, most variant designs focus on the amino acid sites 561-588 in the AAV2 capsid protein VP3. The specific natural wild-type capsid protein sequence is: DEEEIRTTNPVATEQYGSVSTNLQRGNR (SEQ ID NO.1). This region is near the three-fold symmetry axis of the AAV2 VP1 protein, including a buried region, a surface region and an interface region, and overlaps with known heparin and antibody binding sites. However, traditional protein design methods often rely on artificial experience to perform directed evolution and a large number of experimental screenings on the wild-type (WT) protein sequence, including strategies such as random mutagenesis, gene shuffling, and targeted recombination of protein fragments. Directed evolution is a powerful method. When the mechanism is poorly understood, repeated applications of random mutagenesis and artificial selection are usually the default experimental strategies. However, although experimental-based techniques can detect a large number of biological sequences, engineered protein libraries rarely exceed the sequence diversity of natural protein families and most mutations result in non-functional proteins. Since the generated sequences are still very similar to natural isolates, these variants have limited ability to overcome antibody neutralization and do not have the ability of immune evasion.

[0004] With the continuous development of artificial intelligence technology, data-driven machine learning methods for protein engineering design have begun to emerge. The application of data-driven machine learning in protein engineering design not only improves the design efficiency but also opens up new possibilities for biomedical research. In the article "Bryant, D.H., Bashir, A., Sinai, S. et al. Deep diversification of an AAV capsid protein by machine learning. Nat Biotechnol 39, 691–696 (2021).", a large amount of AAV2 capsid protein data is used to train classification models through supervised machine learning methods to predict whether the capsid protein sequence variants are active, including logistic regression models (LR), convolutional neural networks (CNN), and recurrent neural networks (LSTM). By sampling different sequences and using these machine learning algorithms to classify and discriminate the functionality of sequences, the diversification generation of AAV2 capsid proteins is guided. Compared with traditional methods, the advantage is that feasible AAV2 variant sequences can be quickly generated through programs, but there is still room for improvement in the classification prediction accuracy of the model. In the article "Sinai S, Jain N, Church G M, et al. Generative AAV capsid diversification by latent interpolation [J]. bioRxiv, 2021: 2021.04.16.440236.", the authors use an unsupervised method, using the evolutionary data and mutation data of AAV capsid proteins, and use a generative model of variational autoencoder (VAE) to directly generate new AAV2 variant sequences. Although this method can generate many AAV2 variant sequences, the generated sequences have a fixed length, and the generated sequences have insufficient homology and there are still a large number of non-functional proteins.

[0005] It can be seen that the existing data-driven methods still have problems in terms of both the quality of generated sequences and the sequence classification and discrimination ability, and a high-quality and effective solution has not been given. Therefore, there is an urgent need for a new method that can efficiently design diverse AAV2 capsid protein variants. Summary of the Invention

[0006] In view of the problems of the prior art, the present invention provides a system for designing AAV2 capsid protein variants.

[0007] A system for designing AAV2 capsid protein variants includes:

[0008] A preprocessing module, configured to: preprocess the collected sequence data of AAV2 capsid proteins;

[0009] An input module, configured to: input AAV2 capsid protein sequences of different sequence lengths;

[0010] A prediction module, configured to: classify AAV2 capsid protein sequences of different sequence lengths through a fine-tuned AntiBERTy model, and obtain AAV2 capsid protein variant sequences according to the functional scores of the sequences;

[0011] An output module, configured to: output the sequences of the prediction module.

[0012] Preferably, the fine-tuned AntiBERTy model includes an AntiBERTy network layer and an MLP network layer.

[0013] Preferably, the AntiBERTy network layer includes a characterization sequence hidden layer and an average pooling layer; the characterization sequence hidden layer includes a first-layer feature extraction layer, a second-layer feature extraction layer, and a last-layer feature extraction layer, and the average pooling layer splices the features obtained by the first-layer feature extraction layer, the second-layer feature extraction layer, and the last-layer feature extraction layer.

[0014] Preferably, the number of the MLP network layers is two. The first-layer MLP network layer is used to project the vector after splicing the protein sequence features to be evaluated and the mutation features to a low dimension, and the second-layer MLP network layer is used to classify and evaluate the score of the projected vector.

[0015] Preferably, the mutation feature is the difference feature obtained by subtracting the protein sequence feature to be evaluated from the wild-type AAV2 capsid protein sequence feature; the wild-type AAV2 capsid protein sequence is as shown in SEQ ID NO.1.

[0016] Preferably, the operations of the preprocessing include sequence cleaning, redundancy removal, and alignment.

[0017] Preferably, the AAV2 capsid protein sequences of different sequence lengths are generated by a fine-tuned Progen2 model.

[0018] Preferably, the fine-tuned Progen2 model adjusts the model parameters through a loss function to generate AAV2 capsid protein sequences of different sequence lengths; the calculation formula of the loss function is as follows:

[0019]

[0020] In the above formula, L represents the calculated loss, y tk represents the value of the actual token, represents the token output value predicted by the model.

[0021] The present invention also provides a method for designing AAV2 capsid protein variants using the above system, including the following steps:

[0022] Step 1, collect sequence data of AAV2 capsid protein and perform preprocessing;

[0023] Step 2, input AAV2 capsid protein sequences with different sequence lengths;

[0024] Step 3, classify AAV2 capsid protein sequences with different sequence lengths through a fine-tuned AntiBERTy model, and obtain AAV2 capsid protein variant sequences according to the functional scores of the sequences.

[0025] A computer-readable storage medium stores thereon: a computer program for implementing the above system and a program for implementing the method.

[0026] The AAV2 capsid protein variant design system provided by the present invention mainly includes a pre-trained and fine-tuned Progen2 model and a fine-tuned AntiBERTy model; the fine-tuned Progen2 model generates AAV2 capsid protein variant sequences with different sequence lengths by adjusting model parameters, and then inputs them into the fine-tuned AntiBERTy model. The functional scores of the sequences are obtained through feature splicing processing, so as to screen out high-quality AAV2 capsid protein sequences. Therefore, the beneficial effects of the present invention are as follows: high-quality AAV2 capsid protein sequences can be obtained, the homology and diversity of the generated sequences are improved, and the system of the present invention has good prediction performance and classification performance.

[0027] Obviously, based on the above content of the present invention, according to the common general knowledge and conventional means in the art, without departing from the above basic technical idea of the present invention, various other forms of modifications, substitutions or changes can be made.

[0028] The above content of the present invention will be further described in detail below through specific embodiments in the form of examples. However, this should not be construed as limiting the scope of the above subject matter of the present invention to the following examples. All technologies implemented based on the above content of the present invention belong to the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 The site Shannon entropy of AAV2 capsid protein sequences generated by different models.

[0030] Figure 2 The cumulative density function distribution of the sequences generated by the AAV2 capsid protein variant design system after functional scoring. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] It should be specifically noted that the algorithms for data collection, transmission, storage, and processing steps that are not specifically described in the embodiments, as well as the hardware structures and circuit connections that are not specifically described, can all be implemented through the content disclosed in the prior art.

[0032] Example 1 AAV2 Capsid Protein Variant Design System and Method

[0033] The AAV2 capsid protein variant design system of this embodiment includes:

[0034] A preprocessing module, configured to: perform preprocessing on the collected sequence data of AAV2 capsid protein;

[0035] An input module, configured to: input AAV2 capsid protein sequences of different sequence lengths;

[0036] A prediction module, configured to: classify the AAV2 capsid protein sequences through a fine-tuned AntiBERTy model, and obtain AAV2 capsid protein variant sequences according to the functional scores of the sequences;

[0037] An output module, configured to: output the sequences of the prediction module.

[0038] This embodiment also provides a method for designing AAV2 capsid protein variants using the above system. The steps of the method include:

[0039] Step 1, collect the sequence data of AAV2 capsid protein and perform preprocessing;

[0040] Step 2, input AAV2 capsid protein sequences of different sequence lengths;

[0041] Step 3, classify the AAV2 capsid protein sequences of different sequence lengths through a fine-tuned AntiBERTy model, and obtain AAV2 capsid protein variant sequences according to the functional scores of the sequences.

[0042] In the above preprocessing module, the preprocessing operations include sequence cleaning, redundancy removal, alignment, etc.

[0043] In the input module, the AAV2 capsid protein sequences of different sequence lengths are generated by a fine-tuned Progen2 model. The network structure and training method of the fine-tuned Progen2 model are as follows:

[0044] (1) The variant data of AAV capsid proteins in the literature was adopted (Bryant, D. H., Bashir, A., Sinai, S. et al. Deep diversification of an AAV capsid protein by machine learning. Nat Biotechnol 39, 691-696 (2021)). By screening out functional protein sequences, 70% was used as the training set and 30% as the test set. The sequence data was processed using the tokenizer of the pre-trained Progen2 to obtain the corresponding tokens;

[0045] (2) When fine-tuning and training the Progen2 model, the AdamW method in pytorch was used as the optimizer, the learning rate was set to 0.00001, and the number of training epochs was 5. In the training stage, starting from the first position of the input sequence, the next position token was continuously predicted, and the loss was calculated using the actual token and all predicted token values. The loss function used the cross-entropy loss function, and the calculation formula is as follows:

[0046]

[0047] In the above formula, L represents the calculated loss, y tk represents the value of the actual token, represents the token output value predicted by the model;

[0048] (3) After completing the model training, the sampling parameters of the model were set. The key parameters included: max-length is the longest sequence length, num-samples is the number of sequence samplings, and t is the sampling temperature (the higher the temperature, the better the diversity of sequence generation); specifically, the parameter settings were: max-length = 45, t = 1.3, num-samples = 10000;

[0049] (4) The generated tokens were translated into protein sequences using the corresponding tokenizer dictionary to obtain AAV2 capsid protein sequences with different sequence lengths.

[0050] In the prediction module, the constructed fine-tuned AntiBERTy model was used to classify and evaluate the AAV2 capsid protein sequences with different sequence lengths generated by the fine-tuned Progen2 model, and high-quality AAV2 capsid protein variant sequences were screened out.

[0051] Among them, the network structure and training method of the fine-tuned AntiBERTy model are as follows:

[0052] (1) The variant data of AAV capsid proteins in the literature was adopted (Bryant, D. H., Bashir, A., Sinai, S. et al. Deep diversification of an AAV capsid protein by machine learning. Nat Biotechnol 39, 691 - 696 (2021)). 70% of the data was used as the training set, 15% as the validation set, and 15% as the test set;

[0053] (2) Feature extraction of protein sequences was performed using a fine - tuned AntiBERTy model, which includes an AntiBERTy network layer and an MLP network layer; the AntiBERTy network layer includes a representation sequence hidden layer and an average pooling layer. The representation sequence hidden layer includes a first - layer feature extraction layer, a second - layer feature extraction layer, and a last - layer feature extraction layer. The average pooling layer concatenates the features obtained from the first - layer feature extraction layer, the second - layer feature extraction layer, and the last - layer feature extraction layer.

[0054] (3) The number of layers of the MLP network layer is two, and the number of neurons is 1024 and 512 respectively. The first - layer MLP network layer is used to project the vector after concatenating the features of the protein sequence to be evaluated and the mutation features to a low dimension. The second - layer MLP network layer is used to classify and evaluate the score of the projected vector, and obtain a binary classification result of predicting whether the given sequence is functional. The mutation features are the difference features obtained by subtracting the features of the wild - type AAV2 capsid protein sequence from the features of the protein sequence to be evaluated. The wild - type AAV2 capsid protein sequence described herein is the sequence of positions 561 - 588 of wild - type AAV2 capsid protein VP3: DEEEIRTTNPVATEQYGSVSTNLQRGNR (SEQ ID NO.1), and the protein sequence to be evaluated is the capsid protein variant sequence designed or generated by mutation (in this example: DEEEIRTTNPVATEQYGSVSTNLQRGEF (SEQ ID NO.2));

[0055] The loss function adopted by the model is an improved weighted cross - entropy loss function L weights , and its formula is:

[0056]

[0057] where L cross is the cross - entropy loss function, here y represents the true label, The label predicted by the model is denoted as, N represents the number of computational samples in a batch. The learning rate during training is 0.0001, the number of training epochs is 5. The training set is used for training, and the validation set is used for validation. The model parameters with the best performance on the validation set are retained. Finally, the test set is used for testing to obtain the fine-tuned AntiBERTy model.

[0058] Comparative Example 1: The untuned Progen2 model

[0059] It is the same as the construction method of the fine-tuned Progen2 model in Example 1, except that the loss function is not adopted, the model is not trained, and the pre-trained model parameters are not updated.

[0060] Comparative Example 2: Variational Autoencoder model (VAE)

[0061] It is trained using the same data as the Progen2 model involved in fine-tuning in Example 1 to directly generate AAV2 variant sequences of different lengths.

[0062] Comparative Example 3: Convolutional Neural Network model (CNN)

[0063] It is trained using the same data as the fine-tuned AntiBERTy model in Example 1 and adopts the same loss function. The difference lies in the model method. The convolutional neural network (CNN) is used to extract features from the variant protein sequences and output the functional scores of the sequences.

[0064] Comparative Example 4: Recurrent Neural Network model (LSTM)

[0065] It is trained using the same data as the fine-tuned AntiBERTy model in Example 1 and adopts the same loss function. The difference lies in the model method. The recurrent neural network (LSTM) is used to extract features from the variant protein sequences and output the functional scores of the sequences.

[0066] Comparative Example 5: Logistic Regression model (LR)

[0067] It is trained using the same data as the fine-tuned AntiBERTy model in Example 1 and adopts the same loss function. The difference lies in the model method. The logistic regression model (LR) is used to extract features from the variant protein sequences and output the functional scores of the sequences.

[0068] Comparative Example 6: [CLS] classification model

[0069] The method for constructing the fine-tuned AntiBERTy model in Example 1 is the same, except that the special classification head [CLS] feature vector of the last hidden layer in the AntiBERTy model is directly used, all pre-trained parameters are frozen without fine-tuning parameter updates, and only a single-layer MLP non-linear classifier is used for training and classification evaluation.

[0070] The technical solution of the present invention will be further described through experiments below.

[0071] Experimental Example 1: Sequence features generated by the fine-tuned Progen2 model

[0072] In this experimental example, the AAV2 group is the AAV2 capsid protein data set used for fine-tuning the Progen2 model in Example 1, the VAE group is the sequence generated by Comparative Example 2, the Progen2-SFT-AAV group is the sequence generated by the fine-tuned Progen2 model in Example 1, and the Progen2-pretrain group is the sequence generated by Comparative Example 1.

[0073] I. Experimental method

[0074] 1. Shannon entropy

[0075] By comparing the representative statistics of amino acid mutations in the generated sequences, the retention of the evolutionary characteristics of the generated sequences in the natural data set is characterized. The calculation formula of Shannon entropy H(i) is as follows:

[0076] H(i) = -∑ a p(a i ) log(p(a i ))

[0077] where p(a i ) is the probability of amino acid a appearing at position i, and a represents all possible amino acid types.

[0078] II. Experimental results

[0079] The Progen2-SFT-AAV group and the AAV2 group were compared for homology by calculating the Shannon entropy, and the homology of the sequence generation results was significant.

[0080] Such as Figure 1As shown, the Shannon entropy values at each site of the generation results in the Progen2-pretrain group are almost the same, which is significantly inconsistent with the capsid protein sequence distribution in the AAV2 group; compared with the AAV2 group, the VAE group has some similarities in the curve trend, but the Shannon entropy value is significantly higher than the data entropy distribution of the AAV2 group; the Progen2-SFT-AAV group is basically consistent with the AAV2 group in the curve trend, indicating that the distributions of their Shannon entropy are very similar. The experimental results show that the AAV2 capsid protein sequence generated by the fine-tuned Progen2 model in the present invention has good homology.

[0081] Experimental Example 2: Prediction Performance of the Fine-Tuned AntiBERTy Model

[0082] The fine-tuned AntiBERTy model used in this experimental example was constructed according to Example 1, and the CNN, LSTM, LR, and [CLS] classification models were constructed according to Comparative Examples 3, 4, 5, and 6.

[0083] I. Experimental Method

[0084] A total of 153,682 complete data of AAV capsid proteins from the literature "Bryant, D.H., Bashir, A., Sinai, S. et al. Deep diversification of an AAV capsid protein by machine learning. Nat Biotechnol 39, 691-696 (2021)." were used, including functional sequences and non-functional sequences. 70% of the data was used as the training set, 15% as the validation set, and 15% as the test set. The training input was different sequences to be estimated, and the prediction result was the functional score of the sequence. A sequence with a score greater than 0.5 was regarded as an available sequence, and a sequence with a score less than 0.5 was regarded as an unavailable sequence. The evaluation criterion used was classification accuracy (Accuracy), which can be used to represent the proportion of samples correctly classified by the model in the total samples. The formula for calculating accuracy is as follows:

[0085]

[0086] Among them, TP (True Positives): True positive examples, that is, the number of actual positive classes that are predicted as positive classes. TN (True Negatives): True negative examples, that is, the number of actual negative classes that are predicted as negative classes. FP (False Positives): False positive examples, that is, the number of actual negative classes that are predicted as positive classes. FN (False Negatives): False negative examples, that is, the number of actual positive classes that are predicted as negative classes; the higher the accuracy, the better the prediction result.

[0087] In addition, the area under the curve (AUC) of the receiver operating characteristic (ROC) curve is used for evaluation. The AUC is the area under the ROC curve, which can measure the ability of the model to distinguish positive and negative samples. The ROC curve is plotted by calculating the true positive rate (TPR) and false positive rate (FPR) at different thresholds.

[0088]

[0089] The relevant definitions of TP, FN, FP, and TN are consistent with those of accuracy. The higher the AUC value, the better the model prediction effect.

[0090] II. Experimental Results

[0091] The results are shown in Table 1. The accuracy of the algorithm of the fine-tuned AntiBERTy model of the present invention is 0.93, and the AUC value is 0.98, which is better than other algorithms participating in the comparison. When the fine-tuning parameter update is not performed, the [CLS] classification model that only uses the pre-trained parameters of the AntiBERTy model for feature extraction has a prediction performance similar to that of general logistic regression classification, which is significantly lower than the method proposed in the present invention. The experimental results show that the fine-tuned AntiBERTy model in the present invention has good prediction performance and classification performance.

[0092] Table 1 Accuracy and AUC values of different model algorithms

[0093]

[0094] Experimental Example 3: Sequence screening is completed using the AAV2 capsid protein variant design system

[0095] The AAV2 capsid protein variant design system used in this experimental example is constructed according to Example 1.

[0096] I. Experimental Method

[0097] The protein sequences generated by the fine-tuned Progen2 model are input into the fine-tuned AntiBERTy model. Using the feature processing method of the prediction module, the generated sequences are processed for features, and finally the prediction scores are output. By calculating the cumulative density function distribution of the prediction scores and visualizing them, the further evaluation results of the generated sequences are shown, realizing the effective screening of AAV2 capsid protein variants.

[0098] II. Experimental Results

[0099] As Figure 2As shown, after the AAV2 capsid protein sequences generated by the fine-tuned Progen2 model are functionally evaluated using the fine-tuned AntiBERTy model, most of the generated sequences obtain high evaluation scores, indicating that the quality of the generated sequences is good, but there are still some sequences with low scores. Finally, according to the descending order of the functional scores of the generated sequences, the variant sequences with high scores can be effectively screened out.

[0100] In summary, the AAV2 capsid protein variant design system of the present invention generates AAV2 capsid protein sequences with good homology, and this system is superior to existing machine learning algorithms, having better prediction performance and classification performance, thereby realizing the effective screening of AAV2 capsid protein variants.

Claims

1. A system for designing AAV2 capsid protein variants, characterized in that: include: The preprocessing module is configured to: preprocess the collected sequence data of AAV2 capsid protein; The input module is configured to: input AAV2 capsid protein sequences of different sequence lengths; the AAV2 capsid protein sequences of different sequence lengths are generated by the fine-tuned Progen2 model; the fine-tuned Progen2 model adjusts model parameters through a loss function to generate AAV2 capsid protein sequences of different sequence lengths; the loss function calculation formula is as follows: In the above formula, L represents the calculated loss, represents the actual token value, and represents the token output value predicted by the model; The prediction module is configured to: classify AAV2 capsid protein sequences of different sequence lengths through a fine-tuned AntiBERTy model, and obtain AAV2 capsid protein variant sequences according to the functional scores of the sequences; the fine-tuned AntiBERTy model includes an AntiBERTy network layer and an MLP network layer; the AntiBERTy network layer includes a sequence representation hidden layer and an average pooling layer; the sequence representation hidden layer includes a first feature extraction layer, a second feature extraction layer and a last feature extraction layer, and the average pooling layer concatenates the features obtained by the first feature extraction layer, the second feature extraction layer and the last feature extraction layer; The output module is configured to: output the sequence of the prediction module.

2. The AAV2 capsid protein variant design system according to claim 1, characterized in that: The number of the MLP network layers is two. The first MLP network layer is used to project the vector obtained by concatenating the protein sequence features to be evaluated and the mutation features to a low dimension, and the second MLP network layer is used to classify and evaluate the vector projection.

3. The AAV2 capsid protein variant design system according to claim 2, characterized in that: The mutation feature is a difference feature obtained by subtracting the protein sequence feature to be evaluated from the wild-type AAV2 capsid protein sequence feature; the wild-type AAV2 capsid protein sequence is shown in SEQ ID NO.

1.

4. The AAV2 capsid protein variant design system according to claim 1, characterized in that: The pre-processing operations include sequence cleaning, redundancy removal, and alignment.

5. A method for designing AAV2 capsid protein variants using the AAV2 capsid protein variant design system according to any one of claims 1 to 4, characterized in that: The steps include: Step 1, collecting sequence data of AAV2 capsid protein and performing preprocessing; Step 2, input AAV2 capsid protein sequences of different sequence lengths; Step 3: Classify AAV2 capsid protein sequences of different sequence lengths using the fine-tuned AntiBERTy model, and obtain AAV2 capsid protein variant sequences based on the functional scores of the sequences.

6. A computer-readable storage medium, characterized in that: Stored thereon is: a computer program for implementing the method described in claim 5.

Citation Information

Patent Citations

  • Capsid-modified RAAV3 vector composition and application thereof in human liver cancer gene therapy

    CN111763690A

  • Optimization method, system and equipment of AAV capsid protein and storage medium

    CN116312795A