Method and system for generating and screening amino acid sequence
Patent Information
- Application Number
- CN202480011352.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-01
- Filing Date
- 2024-03-27
- Publication Date
- 2025-09-16
AI Technical Summary
Existing amino acid sequence design methods have single design results and are too sensitive to the design details of the main chain structure, which limits the diversity and variability of the design. The screening cycle is long and the cost is high, making it difficult to quickly and efficiently generate amino acid sequences with specific functions. .
Using a system based on a combination of machine learning and rational design, through steps such as pre-training models, conditional generation, rational design and filter modules, amino acid sequences with specific functions, such as viral amino acid sequences, are generated and screened to improve design efficiency and diversity.
It achieves the rapid and efficient design of amino acid sequences with specific functions, reduces experimental costs and time, and improves the production fitness and organ enrichment capabilities of viral amino acid sequences.
Smart Images

Figure CN120660140A_ABST
Abstract
Description
[Corrected 10.04.2024 according to Rule 26] Method and system for generating and screening amino acid sequences
[0001] This disclosure claims priority from the following patent applications: Chinese invention patent application CN2023103070185, filed on March 27, 2023; and Chinese invention patent application CN2023109594061, filed on August 1, 2023. The entire contents of these priority applications are incorporated herein by reference. Technical Field
[0002] The present disclosure relates to a general method for generating and screening amino acid sequences, which can design the generation of various functional amino acid sequences based on a rational approach, and in particular to a method and system for generating and screening viral amino acid sequences, which includes viral vector design and construction, generation of viral amino acid sequences with high specificity and high viral amino acid sequence production fitness, and verification of algorithm results. Background Art
[0003] The structure and function of proteins are determined by their amino acid sequences. In natural proteins, the amino acid sequences are formed through long periods of natural evolution. In recent years, with the growth of biological data and the rapid development of artificial intelligence (AI), AI has achieved substantial breakthroughs in the design and development of biotechnology and medicine. For example, in the field of protein crystal structure analysis, obtaining high-precision protein crystal structures previously required expensive and time-consuming methods such as cryo-electron microscopy, nuclear magnetic resonance (NMR), or X-ray diffraction. Accurate protein structures help us understand life processes and disease mechanisms at the molecular level and design more effective drugs based on disease targets. However, the number of known proteins reaches into the hundreds of millions, and currently fewer than 50,000 high-precision protein structures have been resolved. On July 28, 2022, DeepMind announced that its AlphaFold algorithm has predicted the structures of over 200 million proteins from one million species with high confidence, covering virtually every known protein on Earth and significantly advancing research into protein function. However, when the natural protein structure and function do not meet the requirements of industrial or medical applications, protein design is necessary to obtain specific functional proteins. Based on the hypothesis that sequence determines structure and structure determines function, the de novo design of proteins is almost equivalent to the de novo design of amino acid sequences. On September 15, 2022, a research team from Professor David Baker's laboratory published a latest study in Science, showing that artificial intelligence technology can create practical protein molecules faster and more accurately than before.
[0004] However, the methods described above are not sufficient to address all the current challenges of protein design. Providing a universal method for generating and screening amino acid sequences to facilitate the design of functional protein molecules is a goal that the industry is striving for. However, many existing methods have shortcomings such as single design results and excessive sensitivity to the design details of the main chain structure, which limits the diversity and variability of the designed main chain structure. The current situation is that the design of protein amino acid sequences based on experimental or computer-assisted methods and the screening of amino acid sequences with high activity are still at a bottleneck in terms of technology and are one of the most challenging scientific problems currently.
[0005] Summary of the Invention
[0006] Computer-assisted amino acid sequence generation systems can generate amino acid sequences with specific functions (such as organ targeting) and specific morphological characteristics based on rational design under the guidance of prior knowledge, assisting in the design of antibodies or viral vectors, and are the key to breaking through current gene therapy or antibody protein drug therapy.
[0007] Traditional amino acid sequence design suffers from drawbacks such as long screening cycles, high screening costs, and low sequence diversity. To improve the effectiveness, functional diversity, and synthetic stability of amino acid sequence design, compared to traditional amino acid sequence design, the present disclosure utilizes a pre-training model to learn prior knowledge of amino acid sequences before designing amino acid sequences. This method uses rational design to perform targeted modification on amino acid sequences generated based on machine learning, and then selects amino acid sequences with specific functions based on the desired amino acid sequence functions.
[0008] This disclosure proposes a system for generating, evolving, and screening amino acid sequences based on a combination of machine learning and rational design. This system can quickly and efficiently design amino acid sequences with specific functions and can be used as a general-purpose architecture. By inputting training data and the desired amino acid function, the system can quickly and efficiently design and generate amino acid sequences with the user's desired function. This disclosure comprises five modules: a pre-training module that trains the model on a large amount of unlabeled amino acid sequence data; a sequence generation model that is trained on amino acid sequence data labeled with a specific function; a rational design module that optimizes the model-generated sequences based on the specific function; a filter module that ranks the optimized amino acid sequences based on the specific function to filter out invalid sequences and narrow the scope; and finally, a module that scores and ranks amino acid sequences based on their binding to the target. This disclosure uses the generation of viral amino acid sequences with specific functions as an example. The following section focuses on using this disclosure to generate viral amino acid sequences with specific functions. When performing this task, the algorithm uses the sequence information and physicochemical properties of the viral amino acids as training data to generate amino acid sequences with high organ targeting, high specificity, and high viral production adaptability. It can preserve the diversity of amino acid sequence design while taking into account design efficiency, thereby increasing the experimental speed of screening viral amino acid sequences and reducing experimental costs.
[0009] This disclosure proposes a system for generating, evolving, and screening amino acid sequences based on a combination of machine learning and rational design. This system can quickly and efficiently design amino acid sequences with specific functions and can be used as a universal architecture. By inputting training data and the different functions of the required amino acids, this system can quickly and efficiently design and generate amino acid sequences with the functions required by the user. When performing the task of generating viral amino acid sequences with specific functions, word embedding, graph representation, and position encoding strategies are applied to characterize the amino acid sequences. By extracting the features of the amino acid sequences and extracting the "biological semantics" and "biological grammar" information of the sequences, the production fitness of the viruses is predicted, and the viral amino acid sequences are screened. Specifically, the following steps are included:
[0010] S1, constructs a dataset for pre-training models from a large amount of unlabeled amino acid sequence data (tens of millions of amino acid sequences already in the database) to learn the grammatical and semantic information of amino acid sequences;
[0011] S11, using a large language model to perform preliminary training on a pre-training dataset constructed from a large amount of unlabeled amino acid sequence data, to obtain a pre-trained large model with semantic extraction capabilities;
[0012] S2, based on the pre-trained model, uses the dataset of the downstream specific design task (viral amino acid sequence) to preliminarily "conditionally generate" amino acid sequences with high synthesizability, including the following steps:
[0013] S21 performs feature encoding on the amino acid sequences in the dataset;
[0014] S22 trains the specific amino acid sequence generation module;
[0015] S3, based on the "conditionally generated" amino acid sequence, optimizes the amino acid sequence through the rational design module. The specific steps are as follows:
[0016] S31, based on existing positive sample data with specific functions, uses heuristic algorithms such as annealing algorithm, genetic algorithm, swarm intelligence optimization algorithm, etc. to further optimize and transform the amino acid sequence output by the generation model, so that the sequence space evolves towards the positive sample direction, and the optimized sequence is used as the input of the downstream filter.
[0017] The S4 training filter module performs functional prediction on the optimized sequences (viral production fitness and organ enrichment of viral amino acid sequences) based on the required amino acid sequence functions, filters out amino acid sequences with poor functions, and inputs the remaining sequences into the next module. The specific steps are as follows:
[0018] S4a, viral amino acid sequence production fitness prediction module
[0019] The viral amino acid sequence production fitness prediction module predicts the viral amino acid sequence production fitness based on the viral amino acid sequence. The higher the production fitness, the stronger the ability of the amino acid sequence to generate viruses.
[0020] S4b, viral amino acid sequence organ enrichment prediction module
[0021] The viral amino acid sequence organ enrichment prediction module predicts the organ enrichment of viral amino acid sequences based on the viral amino acid sequences. The higher the organ enrichment, the stronger the amino acid sequence enrichment in the organ.
[0022] S5, rating and ranking module
[0023] The amino acid sequences are scored and ranked according to the degree of binding between the viral amino acid sequence and the target. The viral amino acid sequences with high scores and specific functions are selected and input into the experimental verification module. If the function of the final preferred sequence does not meet the user's requirements, a batch of sequences can be regenerated using the generative model and the above steps can be repeated until a sequence that satisfies the user is obtained;
[0024] S6, experimental verification module
[0025] The viral amino acid sequences with specific functions output by the scoring and ranking module are experimentally synthesized and their functionality is verified.
[0026] Compared with the prior art, the present disclosure has the following contributions:
[0027] In the process of generating amino acid sequences, the logical structure and grammatical and semantic features between sequences are extracted based on the method of "natural language processing" (also referred to as NLP in this disclosure), and amino acid sequence samples of a specific length are generated step by step according to time, so that the generated sequence samples are closer to the real sequence in terms of feature distribution and information, and the logical order information unique to the time series is learned; the natural language representation here is a metaphor, and the specific meaning is: for example, in Chinese, "I like to eat apples" is a sentence that can express the complete meaning, which includes the six language elements (Chinese characters, or words, etc.) of I, like, happy, eat, apple, and fruit. If the six language elements are randomly arranged and combined, hundreds of combination methods can be obtained. However, it is well known that in the Chinese environment, there are not many combinations that combine the language elements of I, like, happy, eat, apple, and fruit into combinations with actual meaning, such as "I like to eat apples", "I like that I eat apples", "I like to eat apples", etc. If the viral capsid amino acid sequence is compared to a sentence, and the viral amino acids are compared to language elements, it can be assumed that the existing known amino acid sequences that can form capsids have certain specific patterns (the grammar of natural language). In the sequence generation process, if this "grammar" is mastered, the amount of calculation can undoubtedly be greatly reduced. The present disclosure proves that first learning the above "grammar" through artificial intelligence and then using the "grammar" to guide the generation of amino acid sequences is very effective in improving the calculation speed and increasing the adaptability of viral amino acid sequence production.
[0028] When generating amino acid sequences, multiple required functions of amino acid sequences are taken into consideration. For example, when generating viral amino acid sequences, comprehensive consideration is given to designing and optimizing viral amino acid sequences with high specificity, high viral amino acid sequence production adaptability, and high organ enrichment capacity. This ensures sequence diversity while enhancing the feasibility of constructing target functional viral amino acid sequences, enabling the model to generate more meaningful viral amino acid sequence data with target functions, thereby improving the efficiency of viral amino acid sequence generation and screening;
[0029] Applying word-unit representation methods and using the physicochemical properties of amino acids to characterize viral amino acid sequence information, this allows for a more comprehensive understanding of the semantic and grammatical information within viral amino acid sequences, enhancing the model's ability to extract and represent sequence features.
[0030] The implementation of rational design (Rational Fitness) is another important feature of the present disclosure. After generating the viral amino acid sequences we need based on the above method, which have high specificity, high viral amino acid sequence production fitness, and high organ enrichment, the most important feature of the present disclosure is that rational design is used to perform targeted modification on the amino acid sequences generated by machine learning. Each amino acid site is modified through an heuristic algorithm. Based on the existing positive sample data with the required function, the amino acid sequence output by the generation model is further optimized and modified, so that the sequence space evolves in the direction of the positive sample, the sequence is optimized, and the predicted sequence is quickly iterated and evolved into an excellent amino acid sequence that is closer to the known reality. Compared with a single random generation step based on machine autonomous learning, the hit rate of prediction screening can be greatly improved and the overall machine time occupancy can be reduced. In addition, by adjusting the positive sample library of the required functional sequence, the functional tendency of the amino acid sequence can be artificially controlled, so that the entire method and system have the ability to adjust the biological function preference, which has great advantages over existing methods. After the combination, the viral amino acid sequences generated with high specificity, high viral amino acid sequence production fitness and high organ enrichment are predicted for viral amino acid sequence production fitness and high organ enrichment, and the binding degree between the generated viral amino acid sequences and the target is predicted, and viral amino acid sequences with high binding degree are selected for experimental verification.
[0031] The embodiments of the present disclosure use the prediction and screening of viral capsids, or so-called viral amino acid sequences, as an example. However, those skilled in the art will recognize, based on the principles of the present disclosure, that the concepts of the present disclosure are clearly not only suitable for the prediction of viral capsid proteins, but can also be used to predict other amino acid sequences, such as penetrating peptide sequences, antimicrobial peptide sequences, antibody sequences, and other functional amino acid sequences. Specifically, another aspect of the present disclosure can also provide a general method for generating and screening amino acid sequences based on machine learning, with the following specific steps:
[0032] SI pre-training step: constructing a dataset for pre-training the model based on the unlabeled amino acid sequence database to learn the grammatical and semantic information of the amino acid sequence;
[0033] The SII conditional generation step uses the amino acid sequence dataset to initially generate amino acid sequences based on the pre-trained model.
[0034] The SIII rational design step optimizes the amino acid sequence based on the amino acid sequence generated in the SII step through the rational design module. Based on the existing positive sample data with the required function, the heuristic algorithm is used to further optimize the amino acid sequence output by the generative model, so that the sequence space evolves in the direction of the positive sample and the sequence is optimized. The steps include:
[0035] SIII-1 obtained amino acid sequence as population G t , t is 0 or a positive integer. In the initial generation of population, t is 0, i.e. G0.
[0036] SIII-2 establishes population G t An evaluation function that performs fitness evaluation, where fitness represents the similarity to a target amino acid sequence with the desired properties,
[0037] SIII-3 to population G t The S32 method is used to evaluate the fitness.
[0038] SIII-4 uses a heuristic algorithm to analyze the population G t Each amino acid site is changed to construct a new generation of population G t+1 , G t to G t+1 Defined as one round of evolution,
[0039] SIII-5 cycles through S33 and S34, presetting the maximum number of evolutionary rounds and the required individual fitness value. After reaching the preset value, the cycle stops, and all individuals in the population are ranked from high to low in fitness, and the optimized target amino acid sequence is output.
[0040] In the SIV filtering step, the function prediction of the amino acid sequences in the target amino acid sequence library is performed based on the desired amino acid sequence function, and the amino acid sequences with poor functions are filtered out.
[0041] On the other hand, the present disclosure, based on the above-mentioned amino acid sequence generation and screening method, discovered an AAV capsid modification strategy with excellent effect, as well as an amino acid sequence for changing the hypervariable region of the AAV capsid, and verified the modification results of AAV, and obtained an AAV vector capsid with more prominent blood-brain barrier penetration effect. Specifically, the present disclosure provides an adeno-associated virus AAV vector capsid, which comprises a continuous amino acid sequence with a length of 7-mer to 18-mer and a C-terminal sequence that satisfies the GYSS.
[0042] In a preferred embodiment, the contiguous amino acid sequence has a length of 7-mer to 8-mer.
[0043] In a preferred embodiment, the contiguous amino acid sequence is part of an AAV vector capsid protein.
[0044] In a preferred embodiment, the contiguous amino acid sequence is inserted into the AAV vector capsid at a position corresponding to amino acids 588-589 of the sequence provided in SEQ ID NO: 1 or 2.
[0045] Another aspect of the present disclosure provides an oligomeric peptide for capsid protein modification of an AAV vector, which has a length of 7-mer to 18-mer and a C-terminal sequence that satisfies GYSS.
[0046] In a preferred embodiment, the oligomeric peptide used for capsid protein modification of an AAV vector has a length of 7-mer to 8-mer.
[0047] In a preferred embodiment, the oligomeric peptide for capsid protein modification of an AAV vector is inserted into the AAV vector capsid at a position between amino acids 588-589 corresponding to the sequence provided in SEQ ID NO: 1 or 2.
[0048] Another aspect of the present disclosure provides a method for modifying an AAV vector capsid, characterized in that a continuous amino acid sequence having a length of 7-mer to 18-mer and a C-terminal sequence satisfying GYSS is inserted at a position between amino acids 588-589 corresponding to the sequence provided in SEQ ID NO: 1 or 2 in the AAV vector capsid.
[0049] The present disclosure provides another aspect of a method for modifying an AAV vector capsid, characterized in that a continuous amino acid sequence from the sequence in the following table is inserted into a position between amino acids 588-589 corresponding to the sequence provided in SEQ ID NO: 1 or 2 in the AAV vector capsid: MMRGYSS (SEQ ID NO: 51) or TGFGYSS (SEQ ID NO: 8).
[0050] The present disclosure provides an AAV vector capsid modification strategy with a more prominent blood-brain barrier penetration effect. The AAV vector used for modification can include any known AAV serotype, including, for example, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10 and AAV11.
[0051] In the present disclosure, the AAV serotype used in the examples is AAV9, but based on common principles, it can include or be derived from any natural or recombinant AAV serotype. For example, the AAV vector can utilize or be based on the AAV serotypes described in WO 2017 / 201258A1, the contents of which are incorporated herein by reference in their entirety, such as but not limited to AAV1, AAV2, AAV2G9, AAV3, AAV3a, AAV3b, AAV3-3, AAV4, AAV4-4, AAV5, AAV6, AAV6.1, AAV6.2, AAV7, AAV7.2, AAV8, AAV9, AAV9.11, AAV9.13, AAV9.16, AAV9.24, AAV9.45, AAV9.9.47, AAV9.61, AAV9.6 8. AAV9.84, AAV9.9, AAV10, AAV11, AAV12, AAV16.3, AAV24.1, AAV27.3, AAV42.12, AAV42-1b, AAV42-2, AAV42-3a, AAV42-3b, AAV42-4, AAV42-5a, AAV42-5b, AAV42-6b, AAV42-8, AAV42-10, AAV42-11, AAV42-12, AAV42-13, AAV42-15, AAV42-aa, AAV43-1 , AAV43-12, AAV43-20, AAV43-21, AAV43-23, AAV43-25, AAV43-5, AAV44.44.1, AAV44.2, AAV44.5, AAV223.1, AAV223.2, AAV2 23.4, AAV223.5, AAV223.6, AAV223.7, AAV1-7 / rh.48, AAV1-8 / rh.49, AAV2-15 / rh.62, AAV2-3 / rh.61, AAV2-4 / rh.50, AAV2-5 / rh.51, AAV3.1 / hu.6, AAV3.1 / hu.9, AAV3-9 / rh.52, AAV3-11 / rh.53, AAV4-8 / r11.64, AAV4-9 / rh.54, AAV4-19 / rh.55, AAV5-3 / rh.57, AAV5-22 / rh.58, AAV7.3 / hu.7, AAV16.8 / hu.10, AAV16.12 / hu.11, AAV29.3 / bb.1, AAV29.5 / bb.2, AAV106.1 / hu.37. AAV114.3 / hu.40, AAV127.2 / hu.41, AAV127.5 / hu.42, AAV128.3 / hu.44, AAV130.4 / hu.48. AAV145.1 / hu.53, AAV145.5 / hu.54、AAV145.6 / hu.55、AAV161.10 / hu.60、AAV161.6 / hu.61。AAV33.12 / hu.17、AAV33.4 / hu.15、AAV33.8 / hu.16、AAV52 / hu.19、AAV52.1 / hu.20、AAV58.2 / hu.25、 AAVA3.3,AAVA3.4,AAVA3.5,AAVA3.7,AAVC1,AAVC2,AAVC5,AAV-DJ,AAV-DJ8,AAVF3,AAVF5,AAVH2,AAVrh.72,AAVhu.8,AAVrh.6 8,AAVrh.70,AAVpi.1,AAVpi.3,AAVpi.2,AAVrh.60,AAVrh.44,AAVrh.65,AAVrh.55,AAVrh.47,AAVrh.69,AAVrh.45,AAVrh.59,AAVrh. AVhu.12, AAVH6, AAVLK03, AAVH-1 / hu.1, AAVH-5 / hu.3, AAVLG-10 / rh.40, AAVLG-4 / rh.38, AAVLG-9 / hu.39, AAVN721-8 / rh.43, AAVCh.5, AAVCh.5R1, AAVcy.2, AAVcy.3, AAVcy.4, AAVcy.5, AAVCy.5R1, AAVCy.5R2, AAVCy.5R3, AAVCy.5R4, AAVcy.6, AAVhu.1, AAVhu .2,AAVhu.3,AAVhu.4,AAVhu.5,AAVhu.6,AAVhu.7,AAVhu.9,AAVhu.10,AAVhu.11,AAVhu.13,AAVhu.15,AAVhu.16,AAVhu.17,AAVhu. hu.18,AAVhu.20,AAVhu.21,AAVhu.22,AAVhu.23.2,AAVhu.24,AAVhu.25,AAVhu.27,AAVhu.28,AAVhu.29,AAVhu.29R,AAVhu.31, AAVhu.32,AAVhu.34,AAVhu.35,AAVhu.37,AAVhu.39,AAVhu.40,AAVhu.41,AAVhu.42,AAVhu.43,AAVhu.44,AAVhu.44,AAVhu.44R 1,AAVhu.44R2,AAVhu.44R3,AAVhu.45,AAVhu.46,AAVhu.47,AAVhu.48,AAVhu.48R1,AAVhu.48R2,AAVhu.48R3,AAVhu.49,AAVhu.51,AAVhu.52,AAVhu.54,AAVhu.54,AAVhu..55,AAVhu.56,AAVhu.57,AAVhu.58,AAVhu.60,A AVhu.61,AAVhu.63,AAVhu.64,AAVhu.66,AAVhu.67,AAVhu.14 / 9,AAVhu.t19,AAVrh.2,AAVrh .2R,AAVrh.8,AAVrh.8R,AAVrh.10,AAVrh.12,AAVrh.13,AAVrh.13R,AAVrh.14,AAVrh.17,A AVrh.18,AAVrh.19,AAVrh.20,AAVrh.21,AAVrh.22,AAVrh.23,AAVrh.24,AAVrh.25,AAVrh.3 1, AAVrh.32, AAVrh.33, AAVrh.34, AAVrh.35, AAVrh.36, AAVrh.37, AAVrh.37R2, AAVrh.38, AAVrh.39, AAVrh.40, AAVrh.46, AAVrh.48, AAVrh.48.1, AAVrh.48.1.2, AAVrh.48.2, AAVrh.49 ,AAVrh.51,AAVrh.52,AAVrh.52,AAVrh..53,AAVrh.54,AAVrh.56,AAVrh.57,AAVrh.58,AAVr h.61,AAVrh.64,AAVrh.64R1,AAVrh.64R2,AAVrh.67,AAVrh.73,AAVrh.74,AAVrh8R,AAVrh8R A586R mutation,AAVrh8R R533A mutation,AAAV,BAAV,caprine AAV,bovine AAV,AAVhE1.1,AAVhEr1.5,AAVhER1.14,AAVhEr1.8,AAVhEr1.16,AAVhEr1.18,AAVhEr1.35,AAVhEr1.7,AAVhEr1 .36,AAVhEr2.29,AAVhEr2.4,AAVhEr2.16,AAVhEr2.30,AAVhEr2.31,AAVhEr2.36,AAVhER1.23,AAVhEr3.1,AAV2.5T,AAV-PAEC,AAV-LK01,AAV-LK02,AAV-LK03,AAV-LK04,AAV-LK05,AAV-LK06,AAV-LK07,AAV-LK08,AAV-LK09,AAV-LK10,AAV- LK11, AAV-LK12, AAV-LK13, AAV-LK14, AAV-LK15, AAV-LK16, AAV-LK17, AAV-LK18, AAV-LK19, AAV-PAEC2, AAV-PAEC2, AAV-PAU-PA eC4, AAV-PAEC6, AAV-PAEC7, AAV-PAEC8, AAV-PAEC11, AAV-PAEC12, AAV-2-pre-miRNA-101, AAV-8h, AAV-8b, AAV-h, AAV-b, AAV SM 10-2, AAV Shuffle 100-1, AAV Shuffle 100-3, AAV Shuffle 100-7, AAV Shuffle 10-2, AAV Shuffle 10-6, AAV Shuffle 10-8, AAV Shuffle 100-2, AAV SM 10-1, AAV SM 10-8, AAV SM 100-3, AAV SM 100-10, BNP61 AAV, BNP62 AAV, BNP63 AAV, AAVrh.h.50, AAVrh.43, AAVrh.62, AAVrh.48, AAVhu.19, AAVhu.11, AAVhu.53, AAV4-8 / rh.64, AAVLG-9 / hu.39, AAV54.5 / hu.23, AAV5 4.2 / hu.22, AAV54.7 / hu.24, AAV54.1 / hu.21, AAV54.4R / hu.27, AAV46.2 / hu.28, AAV46.6 / hu.29, AAV128.1 / hu.43, pure AAV(ttAAAV) UPENN AAV 10,Japanese AAV 10serotypes,AAV CBr-7.1,AAV CBr-7.10,AAV CBr-7.2,AAV CBr-7.3,AAV CBr-7.4,AAV CBr-7.5,AAV CBr-7.7,AAV CBr-7.8,AAV CBr-B7.3,AAV CBr-B7.4,AAV CBr-E1,AAV CBr-E2, AAV CBr-E3, AAV CBr-E4.AAV CBr-E5,AAV CBr-e5,AAV CBr-E6,AAV CBr-E7,AAV CBr-E8,AAV CHt-1,AAV CHt-2,AAV CHt-3,AAV CHt-6.1,AAV CHt-6.10,AAV CHt-6.5,AAV CHt-6.6,AAV CHt-6.7,AAV CHt-6.8,AAV CHt-P1,AAV CHt-P2,AAV CHt-P5,AAV CHt-P6,AAV CHt-P8,AAV CHt-P9,AAV CKd-1,AAV CKd-10,AAV CKd-2.AAV CKd-3,AAV CKd-4,AAV CKd-6,AAV CKd-7,AAV CKd-8,AAV CKd-B1,AAV CKd-B2,AAV CKd-B3,AAV CKd-B4,AAV CKd-B5,AAV CKd-B6,AAV CKd-B7,AAV CKd-B8,AAV CKd-H1,AAV CKd-H2,AAV CKd-H3,AAV CKd-H4,AAV CKd-H5,AAV CKd-H6,AAV CKd-N3,AAV CKd-N4,AAV CKd-N9,AAV CLg-F1,AAV CLg-F2,AAV CLg-F3,AAV CLg-F4,AAV CLg-F5,AAV CLg-F6,AAV CLg-F7,AAV CLg-F8,AAV CLv-1,AAV CLv1-1,AAV Clv1-10,AAV CLv1-2,AAV CLv-12,AAV CLv1-3,AAV CLv-13,AAV CLv1-4,AAV Clv1-7,AAV Clv1-8,AAV Clv1-9,AAV CLv-2,AAV CLv-3,AAV CLv-4.AAV CLv-6, AAV CLv-8, AAV CLv-D1, AAV CLv-D2, AAV CLv-D3, AAV CLv-D4, AAV CLv-D5, AAV CLv-D6, AAV CLv-D7, AAV CLv-D8, AAV CLv-E1, AAV CLv-K1, AAV CLv-K3, AAV CLv-K6, AAV CLv-L4, AAV CLv-L5, AAV CLv-L6, AAV CLv-M1, AAV CLv-M11, AAV CLv-M2, AAV CLv-M5, AAV CLv-M6, AAV CLv-M7, AAV CLv-M8, AAV CLv-M9, AAV CLv-R1, AAV CLv-R2, AAV CLv-R3, AAV CLv-R4, AAV CLv-R5, AAV CLv-R6, AAV CLv-R7, AAV CLv-R8, AAV CLv-R9, AAV CSp-1, AAV CSp-10, AAV CSp-11, AAV CSp-2, AAV CSp-3, AAV CSp-4, AAV CSp-6, AAV CSp-7, AAV CSp-8, AAV CSp-8.10, AAV CSp-8.2, AAV CSp-8.4, AAV CSp-8.5, AAV CSp-8.6, AAV CSp-8.7, AAV CSp-8.8, AAV CSp-8.9, AAV CSp-9, AAV.hu.48R3, AAV.VR-355, AAV3B, AAV4, AAV5, AAVF1 / HSC1, AAVF11 / HSC11, AAVF12 / HSC12, AAVF13 / HSC13, AAVF14 / HSC14, AAVF15 / HSC15, AAVF16 / HSC16, AAVF17 / HSC17, AAVF2 / HSC2, AAVF3 / HSC3, AAVF4 / HSC4, AAVF5 / HSC5, AAVF6 / HSC6, AAVF7 / HSC7, AAVF8 / HSC8, AAVF9 / HSC9, AAV-PHP.B (PHP.B), AAV-PHP.A (PHP.A), G2B-26, G2B-13, TH1.1-32 and / or TH1.1-35 and their variants.
[0052] It should be noted that there seems to be a possible error in "AAV CSp-8.4" in the original text, which is likely a typo and is translated as "AAV CSp-8.4" here.Various aspects of the present disclosure relate to AAV capsid proteins. The AAV capsid proteins described herein can have a sequence that is different from the corresponding wild-type AAV capsid protein sequence or a sequence that is different from a reference AAV capsid protein sequence. The AAV capsid protein can include insertions, deletions, or substitutions of one or more nucleotides or one or more amino acids relative to the corresponding wild-type AAV capsid protein sequence or relative to the reference AAV capsid protein sequence. The insertions, deletions, or substitutions of one or more nucleotides or one or more amino acids can be at the 5' end, 3' end, and / or internally within the capsid sequence.
[0053] The AAV capsid disclosed herein includes a targeting sequence (e.g., a 7-mer to 18-mer sequence), which enables the AAV vector to have specific functions, including, in certain embodiments, enabling the AAV capsid to have a stronger ability to cross the blood-brain barrier, tissue specificity, and high transfection efficiency, opening up new design ideas for the future field of gene therapy. In some embodiments, the targeting sequence is inserted into the capsid protein of the AAV vector. In the examples, the targeting sequence (also referred to as an oligomeric peptide for AAV modification in this disclosure) is inserted into a specific site, but it can also be inserted into any region of the capsid protein. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] FIG1 shows a method for generating, optimizing, and screening protein amino acid sequences based on machine learning according to an embodiment of the present disclosure;
[0055] FIG2 shows a flow chart of constructing a data set for training a model generation module from experimental data according to an embodiment of the present disclosure;
[0056] FIG3 shows a flowchart of genetic algorithm sequence transformation steps of a rational design module according to an embodiment of the present disclosure;
[0057] FIG4 shows a flowchart of the steps of sequential transformation of the rational design module distribution estimation algorithm provided by an embodiment of the present disclosure;
[0058] FIG5 shows a flow chart of the steps of transforming the hybrid algorithm sequence of the rational design module provided by an embodiment of the present disclosure;
[0059] FIG6 shows a flow chart of the filter module provided by an embodiment of the present disclosure, wherein the properties of the filter are based on the properties that the user wants to design the amino acid sequence to include, and the filter ranking is sorted from high to low according to the importance of the properties considered by the user:
[0060] FIG7 shows a flow chart of the filter module training process provided by an embodiment of the present disclosure:
[0061] FIG8 shows the steps for verifying the viral amino acid sequence in the target viral amino acid sequence library provided by an embodiment of the present disclosure:
[0062] FIG9 shows the probability density distribution of viral amino acid sequence production fitness when constructing a dataset for training a model using experimental data in an embodiment of the present disclosure;
[0063] FIG10 shows a heat map of the Pearson correlation analysis of the frequency of occurrence of viral amino acid sequences in multiple plasmid replicate experiments and multiple virus replicate experiments when constructing a dataset for training a model according to an embodiment of the present disclosure;
[0064] FIG11 shows the probability density distribution of the organ enrichment of viral amino acid sequences when constructing a dataset for training a model using experimental data in an embodiment of the present disclosure;
[0065] FIG12 shows the probability density distribution of the organ enrichment of viral amino acid sequences when constructing a dataset for training a model using experimental data in an embodiment of the present disclosure;
[0066] FIG13 shows the frequency distribution of negative sample amino acids for fitness production of viral amino acid sequences in an embodiment of the present disclosure;
[0067] FIG14 shows the frequency distribution of amino acids in positive samples of viral amino acid sequence production fitness in an embodiment of the present disclosure;
[0068] FIG15 shows the frequency distribution of amino acids in positive samples of viral amino acid sequence organ enrichment in an embodiment of the present disclosure;
[0069] FIG16 shows the frequency distribution of amino acids in negative samples of viral amino acid sequence organ enrichment according to an embodiment of the present disclosure;
[0070] FIG17 shows the frequency distribution of amino acids in positive samples of viral amino acid sequence organ enrichment according to an embodiment of the present disclosure;
[0071] FIG18 shows the frequency distribution of amino acids in negative samples of viral amino acid sequence organ enrichment according to an embodiment of the present disclosure;
[0072] FIG19 is a schematic diagram showing the effects of using different proportions of the distribution estimation algorithm and the genetic algorithm in the hybrid algorithm in an embodiment of the present disclosure;
[0073] FIG20 shows the distribution estimation algorithm using different excellent individual ratios r in the embodiment of the present disclosure. best Effect diagram;
[0074] FIG. 21 shows the different replication rates r used in the genetic algorithm in the embodiment of the present disclosure. copy Effect diagram;
[0075] FIG. 22 shows the different mutation rates r used in the genetic algorithm in the embodiment of the present disclosure.mutate Effect diagram;
[0076] FIG23 shows the Pearson correlation coefficient of the prediction results of the viral amino acid production fitness prediction module in the filter module for the viral amino acid sequence production fitness in an embodiment of the present disclosure;
[0077] FIG24 shows a Pearson correlation analysis of the viral amino acid sequence production fitness predicted by the viral amino acid sequence production fitness prediction module and the actual viral amino acid sequence production fitness in an embodiment of the present disclosure;
[0078] FIG25 shows a schematic diagram of the probability density distribution of the viral amino acid sequence production fitness predicted by the viral amino acid sequence production fitness prediction module after the viral amino acid sequence generated by the sequence generation module in an embodiment of the present disclosure;
[0079] FIG26 shows a flow chart of the steps for implementing the model architecture for amino acid sequence generation and optimization according to an embodiment of the present disclosure;
[0080] FIG27 shows the ability of the sequences screened by the model to cross the blood-brain barrier when injected into mice according to an embodiment of the present disclosure;
[0081] FIG28 shows the ability of the sequences screened by the model to target the liver when injected into mice according to an embodiment of the present disclosure;
[0082] FIG29 shows the ability of the sequences screened by the model to target other tissues after injection into mice according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0083] To better understand the technical solutions of the present disclosure, the specific embodiments of the present disclosure are further described in detail below in conjunction with the accompanying drawings and examples. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified. The following description and drawings are only for the purpose of initial preferred embodiments and are not to be considered as limiting the scope of the present disclosure.
[0084] In the terms of the present disclosure, expressions such as S1, S2...Sn, unless otherwise specified, generally represent steps executed in sequence, and expressions using Latin letters such as S4a, S4b, etc., indicate that these steps are in a parallel selection relationship and there is no limitation on the order of precedence.
[0085] Example 1
[0086] This disclosure proposes a method for generating, designing, and screening amino acid sequences. As shown in Figure 1, based on a large model of amino acid sequence semantic learning from a generalized dataset, a virus optimization module is designed based on rational design, and pathogenic amino acid sequences are generated and experimentally verified to find amino acid sequences with high specificity and high diversity. The specific steps are as follows:
[0087] S0, constructing a dataset for training the model from experimental data, as shown in Figure 2, includes the following steps:
[0088] S01, viral plasmid library construction steps,
[0089] Randomly generate M (M is 10,000 to 200,000) amino acid sequences of length N (N is 1 to 100). Each amino acid sequence is connected to a specific barcode, and all barcode-connected amino acid sequences are pooled together to construct an amino acid sequence pool. The amino acid sequence library in the amino acid sequence pool is used to replace some sites in the target plasmid to obtain viral plasmids. Different amino acid sequences or replacements of different sites in the target plasmid will produce different viral plasmids, and these different viral plasmids together constitute the viral plasmid library.
[0090] Each amino acid sequence is connected to a specific barcode for subsequent high-throughput sequencing.
[0091] The viral plasmids in the current viral plasmid library are only viral plasmids that may become viruses. They still need to undergo viral vector packaging, virus quality purity testing, and virus in vitro biological activity testing to determine whether the obtained viral plasmids can eventually become viruses.
[0092] S02, viral amino acid sequence production fitness data collection step
[0093] The viral plasmid frequency was calculated by high-throughput sequencing of the plasmid DNA extracted from the viral plasmid. Specifically, the frequency of occurrence of a single viral amino acid sequence in the plasmid was calculated by high-throughput sequencing of the plasmid DNA extracted from the replaced plasmid to obtain the viral plasmid frequency.
[0094] At the same time, after the viral plasmid is transfected into the cells, the virus is purified and the viral DNA is extracted, and then high-throughput calculation is performed to obtain the virus frequency. The specific process is: the viral plasmid is transfected into the cells, the virus is produced in the cells (that is, the virus is spread), and then the virus is purified. After the virus is purified, the virus quality purity is tested, and the viral DNA is extracted and the virus frequency is calculated by high-throughput sequencing; the virus frequency refers to the number of times the same viral amino acid sequence appears in the virus after high-throughput sequencing.
[0095] Finally, the viral plasmid frequency and viral frequency are used to calculate the production fitness of a single viral amino acid sequence. Production fitness is a quantitative representation of the ability of a viral amino acid sequence to generate a virus. A higher production fitness indicates a stronger ability of the amino acid sequence to generate a virus.
[0096] S03 build data set
[0097] The dataset uses viral amino acid sequences as samples, and the production fitness of the corresponding viral amino acid sequences serves as the dataset label. This example primarily uses viral amino acid sequences that have been cleaned of premature stop codons and those with sequencing errors detected during high-throughput sequencing. The samples in the dataset are divided into a training set: validation set: test set ratio of 7:1:2.
[0098] To avoid imbalanced training data and ensure the accuracy and research significance of the model training process, we first need to check the normality of the overall distribution of the viral amino acid sequence production fitness data labels. For unevenly distributed data, we use strategies such as normalization, downsampling, and gradient clipping to balance the data to ensure that the model has no bias during the learning process. Figure 6 shows the probability density distribution of viral amino acid sequence production fitness when constructing the dataset used for model training from experimental data.
[0099] S1, constructs a dataset for pre-training models from a large amount of unlabeled amino acid sequence data (tens of millions of amino acid sequences already in the database) to learn the grammatical and semantic information of amino acid sequences;
[0100] S11: Input the viral amino acid sequences in the dataset into a data preprocessing module to perform feature encoding on the viral amino acid sequences, thereby obtaining viral amino acid sequence features. The viral amino acid sequence features include the characteristics of each amino acid in the viral amino acid sequence. The data preprocessing module can use existing strategies such as word embedding, graph representation, or positional encoding to represent the feature information of the viral amino acid sequence.
[0101] A large language model is initially trained on a pre-training dataset constructed from a large amount of unlabeled amino acid sequence data to obtain a pre-trained large model with semantic extraction capabilities;
[0102] S2, based on the pre-trained model, uses the dataset of the downstream specific design task (viral amino acid sequence) to preliminarily "conditionally generate" amino acid sequences with high synthesizability, including the following steps:
[0103] S21 performs feature encoding on the amino acid sequences in the dataset;
[0104] The viral amino acid sequences in the dataset are input into the data preprocessing module, which performs feature encoding on the viral amino acid sequences to obtain viral amino acid sequence features. These features include the characteristics of each amino acid in the viral amino acid sequence. The data preprocessing module can use existing strategies such as word embedding, graph representation, or positional encoding to represent the feature information of the viral amino acid sequences.
[0105] S22 trains the specific amino acid sequence generation module;
[0106] There are currently 20 known amino acids. The amino acid sequence is composed of the combination of amino acids, but only some of the combinations of amino acids can generate viruses.
[0107] The specific amino acid sequence generation module generates viral amino acid sequences by learning the logical structure and grammatical semantic features of existing amino acid sequences that can generate viruses. Specificity refers to conforming to the logical structure and grammatical semantic features of amino acid sequences that can generate viruses.
[0108] In this embodiment, the specific amino acid sequence generation module adopts a long short-term memory neural network. The training steps for the specific amino acid sequence generation module are as follows: select viral amino acid sequences from the training set according to the set batch number, and the batch number refers to the number of viral amino acid sequences that can be input into the specific amino acid sequence generation module at one time; after the viral amino acid sequence is text-encoded, a number 0 is added to the front of each viral amino acid sequence and the last amino acid is truncated, and all the features of the processed viral amino acid sequence are used as the input of the long short-term memory neural network, and the original viral amino acid sequence after text encoding without other processing is used as the target to be generated by the long short-term memory neural network. The internal workflow of the long short-term memory neural network is that an input viral amino acid sequence will be split into multiple units. amino acids, and uses the previous amino acid as the input of the long short-term memory neural network to generate the amino acid at the next position. The next generated amino acid is used as the input of the long short-term memory neural network to generate the amino acid at the next position, and so on, until an amino acid sequence with the same length as the target amino acid sequence is generated, and the generated amino acid sequence is input into the linear layer for amino acid sequence feature extraction, and the extracted features are input into the activation function softmax to obtain the final generated amino acid sequence; the selected viral amino acid sequence features are compared with the generated amino acid sequence features, the loss function is calculated, and back propagation is performed, and the viral amino acid sequences in the entire training set are cyclically selected according to the batch size to complete the above amino acid generation steps, and the above cycle is repeated M times until the loss function is stable, and the model training parameters are saved. Therefore, when using a data set containing N viral amino acid sequences to train the specific amino acid sequence generation module, it is necessary to go through (N / batch number)*M rounds of iterations before the specific amino acid sequence generation module training is completed;
[0109] Wherein, 1≤N≤50000000, N is a positive integer, 1≤M≤50000, M is a positive integer.
[0110] S23 uses the trained viral amino acid sequence generation and screening device to generate viral amino acid sequences, sets the length of the viral amino acid sequences and the number of sequences to be generated, and then inputs the parameters into a specific amino acid sequence generation module;
[0111] The execution steps of the specific amino acid sequence generation module are as follows: randomly generate an amino acid, characterize the randomly generated amino acid, input the amino acid feature into the trained long short-term memory neural network, and the long short-term memory neural network generates the viral amino acid sequence according to the set amino acid sequence length. Specifically, after the amino acid feature is input into the trained long short-term memory neural network, the long short-term memory neural network will generate the next amino acid according to the current amino acid feature, and the generated amino acid will be input into the long short-term memory neural network to generate the next position amino acid, and so on until the set number of amino acids of the length are generated, and then these generated amino acids are spliced in the dimension of the sequence length, and the above steps are repeated. When the number of viral amino acid sequences reaches the preset generation number, the specific amino acid sequence generation module stops generating the viral amino acid sequence.
[0112] All generated viral amino acid sequences are sent to the viral amino acid sequence production fitness prediction module.
[0113] In addition to the virus-specific sequence generation module described in this embodiment using a long short-term memory neural network to generate a generative adversarial network architecture (GAN) to generate viral amino acid sequences, reliable sequence generation algorithms such as a variational autoencoder (VAE) and a diffusion model can also be used to achieve the generation of virus-specific sequences.
[0114] It's important to note that the variational differentiable autoencoder architecture consists of two parts: an encoder and a decoder. The encoder compresses the raw data into low-dimensional feature vectors using a neural network, while the decoder restores the compressed feature vectors through a neural network and generates data that conforms to the original distribution. While the generated data is highly diverse, the results are often ambiguous. Without the adversarial process of the discriminator, quality assurance is difficult. Furthermore, the posterior distribution is assumed to be a decomposable Gaussian distribution, which is highly hypothetical based on the encoder.
[0115] The diffusion model includes a diffusion process and an inverse diffusion process. The diffusion process, under the conditions of a Markov chain, gradually adds Gaussian noise to the original data, causing it to gradually approach a Gaussian distribution, ultimately obtaining globally diffused data. The inverse diffusion process, on the other hand, uses Gaussian noise to restore the globally diffused data and generate the original data. A generative model based on this architecture can generate relatively robust sequence information, but due to current technological limitations, the generalization ability of the generated data still needs to be improved. Therefore, the virus-specific sequence generation module in this embodiment is the optimal choice.
[0116] As a specific implementation of the diffusion model, the following example can be used: using a neural network to learn how to denoise and demaximize the diffusion process in order to learn the effective distribution of amino acid sequence information, thereby generating highly specific viral amino acid sequences from noise that satisfies a specific prior distribution during the sampling phase. The training steps for the specific amino acid sequence generation module are: selecting viral amino acid sequences from the training set according to the set batch size; label encoding the viral amino acid sequence (i.e., numbering the amino acids 'A', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'K', 'L', 'M', 'N', 'P', 'Q', 'R', 'S', 'T', 'V', 'W', 'Y' from 0 to 19 according to the number of natural amino acids), and then using the labeled viral amino acid sequence as the input of the model. The labeled viral amino acid sequence is defined as Where N is the number of amino acids in a single amino acid sequence, a i is the encoded value of the i-th amino acid label. The purpose of the diffusion model is to learn the spatial distribution p(x) of the amino acid sequence. The internal workflow of the model is as follows: First, the diffusion time step is set to t (t is a positive integer between 0 and 100,000). The diffusion time step is used to limit the number of steps of prediction noise during model training. A diffusion probability model defines two Markov chains of the diffusion process. The forward diffusion process gradually adds noise to x, eventually converting it into complete noise. The reverse generation diffusion process learns how to reverse the forward diffusion process, gradually eliminating some of the noise from the noise to restore the true data x. Forward diffusion can be defined as the following formula:
[0117] α t =1-β t , α t is the weight value of the noise added, the weight value of the noise increases as the time step increases, and z is the noise that obeys the Gaussian distribution. t represents the amino acid sequence at the tth time step, and I represents the identity matrix, that is, a matrix whose elements are all 1. The reverse denoising process can be defined as the following formula:
[0118] μ θ Represents a parameterized neural network to learn the mean and variance of the reverse distribution is a randomly set parameter, so what needs to be learned during training is μ θ .
[0119] We can use variational inference to obtain the variational lower bound (VLB) to optimize the negative log-likelihood as the maximization optimization objective:
[0120] make This is the objective function we want to optimize.
[0121] During model training, the noise added during forward diffusion is used as a training label, and the model is used to predict the noise that needs to be removed during backward diffusion, ensuring that the distribution of the added noise and the noise to be removed are as close as possible. Based on the above principle, all data is divided into multiple batches according to the number of batches. Each batch of original viral amino acid sequence features is input into the model. Random noise is added to the original viral amino acid sequence features sequentially according to the time step, that is, the original data is "noised". The data with random noise added and its corresponding time step are then input into the neural network to predict the noise that needs to be removed during the backward diffusion process at that time step. The internal training steps of the neural network are as follows: the data with random noise added and its corresponding time step are input into the neural network. First, the data with random noise added is input into the linear layer to extract feature information. After the feature information is extracted, the corresponding time step value is added and then the feature is delinearized using the activation function ReLU to obtain the activation feature. This is repeated N times (N is a positive integer from 0 to 1000). After that, the feature is aggregated through a linear layer and the predicted noise is output.
[0122] The predicted noise and the actual noise added are used to calculate the mean squared error (MSE) as the loss function for the neural network backpropagation. Backpropagation is then performed, and the amino acid generation step is completed by cyclically selecting viral amino acid sequences from the entire training set according to the batch size. This cycle is repeated M times until the loss function stabilizes, and the model training parameters are saved. Therefore, when training the specific amino acid sequence generation module using a dataset containing N viral amino acid sequences, training of the specific amino acid sequence generation module is completed after (N / batch size)*M iterations; where 1≤N≤50000000, N is a positive integer, and 1≤M≤50000, M is a positive integer.
[0123] S3, based on the "conditionally generated" amino acid sequence, optimizes the amino acid sequence through the rational design module. Based on the existing positive sample data with specific functions, heuristic algorithms such as annealing algorithm, genetic algorithm, and swarm intelligence optimization algorithm are used to further optimize the amino acid sequence output by the generative model, so that the sequence space evolves in the direction of positive samples. The optimized sequence is used as the input of the downstream filter. The specific steps are as follows:
[0124] S31 obtains the target amino acid sequence output from the sequence generation module, which is called the primary population G0. Each sequence s i ∈G0 is called an individual in the population, and the current population evolution round number t=0.
[0125] S32 establishes population G t Evaluation function for fitness evaluation.
[0126] Import the sequence s′ from the target amino acid sequence database with specific properties in advance i And the corresponding characteristic information, and normalize the corresponding characteristics to obtain the normalized weight w i When all sequences lack corresponding feature information, a binary classification method is used, and all positive samples are set to w + =1, set all negative samples to w - =-1.
[0127] For each individual s in the population i (i=1,2,...,n), and a sequence s′ in the database j (j=1,2,...,m), the sequence similarity score c between them can be calculated by the global similarity algorithm (Needleman-Wunsch algorithm) or the local similarity algorithm (Smith-Waterman algorithm) ij .
[0128] The similarity c between the individual and the sequence in the database ij The corresponding normalized weight w j After multiplication and accumulation, the sum of weighted sequence similarity scores is obtained. This is the fitness score v of the individual i =∑w j c ij .
[0129] According to the above algorithm, all individuals in the population are calculated separately to obtain the fitness level of all individuals in the population.
[0130] S33 evaluates the fitness of population G0 using the method of S32.
[0131] S34 builds a new generation of population G t+1Different swarm intelligence optimization algorithms have different steps. This paper uses swarm intelligence optimization algorithms in the discrete space of categorical variables, such as the genetic algorithm (GA, Genetic Algorithm, as shown in Figure 3), the distribution estimation algorithm (EDA, Estimation of Distribution Algorithm, as shown in Figure 4), and a hybrid algorithm that integrates these algorithms (as shown in Figure 5). The hybrid weight can be EDA:GA = 10~0:0~10 (EDA+GA = 10). The following uses the genetic algorithm, the distribution estimation algorithm, and their hybrid algorithm as examples.
[0132] Genetic algorithms can generate random mutations, which helps to discover the possible positive effects of new amino acids at certain positions. The specific steps for constructing a new generation of populations are as follows:
[0133] The parent population G t Sort by individual fitness scores from high to low G t ={s t,i |i=1,2,...,n;v t,i >v t,i+1}, according to the set replication rate r c The individuals with the highest individual fitness (n c =n×r c ) is copied into the new generation population, i.e. s t+1,i =s t,i (i=1,2,...,n c ). The remaining individuals are generated by crossover and mutation of the parent population: First, the fitness scores of the individuals are normalized: Then, the normalized individual fitness scores are used as weights, and the t The weight is the probability (that is, the probability of selecting sequence i is p i =v′ t,i ) Extract two sequences Unit point or double locus Crossover recombination obtains recombinant individual s′ t+1,i The recombinant individuals follow the set mutation rate r m Randomly mutate the amino acid at each site and add the resulting sequence to the new generation population G t+1 The crossover, recombination and mutation of the parent population are continued until the number of individuals in the new generation population is the same as that in the parent population. t |=|G t+1 |=n.
[0134] The distribution estimation algorithm constructs a probability model based on the distribution of the parent population in each dimension and samples to generate new individuals. Compared with the crossover operation of the genetic algorithm, the distribution estimation algorithm can change the local structure of the parent individuals to a greater extent, helping to discover new and excellent local structures. The specific steps for constructing the new generation population are:
[0135] Let the library of twenty natural amino acids be AA. t A certain number of excellent individuals n b =n×r b The amino acids at each position are counted to obtain the distribution of different amino acids at each position, and the proportion of each amino acid at each position is calculated (i.e., p ij is the frequency of amino acid j at position i). The generation of new individuals is obtained by sampling the probability model: for any position i = 1, 2, ..., l, according to the proportion p of each amino acid at that position ij (j∈AA) is randomly sampled to obtain the amino acid a at that position i (i.e. P(a i =j)=p ij ). Sampling is continued until the corresponding amino acid is extracted at each position and connected to obtain a new amino acid sequence. Count this individual into the new generation population. Repeated sampling until the number of individuals in the new generation population is the same as the number of individuals in the parent population |G t |=|G t+1 |=n.
[0136] The hybrid algorithm integrates the advantages of the previous algorithm. The specific steps for constructing a new generation of population are as follows:
[0137] The proportion of subpopulation individuals generated by each algorithm is pre-set (0% to 100%) to ensure that the sum of the subpopulation individuals is equal. All individuals in the population are sorted from high to low in terms of fitness. If the corresponding conditions are met (for example, the fitness of the optimal individual has reached the requirement, or the number of evolutionary rounds has reached the set value), the algorithm ends and the modified target amino acid sequence is output to the filter. Otherwise, the algorithm returns to S34 to construct a new generation of population.
[0138] S4, the amino acid sequence optimized by the rational design module is filtered through the target property filter. The specific steps are as follows:
[0139] S4a trains viral amino acid sequences to produce fitness prediction modules
[0140] The viral amino acid sequence production fitness prediction module can predict the viral amino acid sequence production fitness based on the viral amino acid sequence. The higher the production fitness, the stronger the ability of the amino acid sequence to generate viruses.
[0141] The training steps of the viral amino acid sequence production fitness prediction module are as follows: the feature information of a viral amino acid sequence is simultaneously input into multiple parallel convolution blocks, the features extracted from multiple convolution blocks are spliced in the hidden layer, the residual module uses a residual connection network to retain the original features with a certain probability while using the convolution layer to update the features, the previously extracted features are input into the residual module, and then layer normalization (LN) is used to perform feature aggregation to obtain aggregated features, the aggregated features are activated using the Sigmoid activation function to obtain sequence information weights, and the aggregated features are delinearized using the relu activation function to obtain activation features, the sequence information weights are multiplied by the activation features to obtain weighted activation information, 1 is subtracted from the previously obtained sequence information weights and multiplied by the previous viral amino acid sequence features to obtain weighted original information, the weighted activation information and the weighted original information are added together as the predicted features, so that the original feature information is passed in with a certain probability while updating the features, the predicted features are input into the temporary backoff method and then into the linear layer, and finally the activation function leakyrelu is used as the output of the viral amino acid sequence production fitness prediction. After Q (1≤Q≤50000, Q is a positive integer) rounds of iteration, the training of the viral amino acid sequence production fitness prediction module is completed.
[0142] After receiving all the generated viral amino acid sequences, the viral amino acid sequence production fitness prediction module uses the trained viral amino acid sequence production fitness prediction module to predict the viral amino acid sequence production fitness of each viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness. The viral amino acid sequences are sorted from high to low according to their production fitness, and sequences with low viral amino acid sequence production fitness are filtered out.
[0143] S4b training module for predicting organ enrichment of viral amino acid sequences
[0144] The viral amino acid sequence organ enrichment prediction module can predict the viral amino acid sequence organ enrichment based on the viral amino acid sequence. The higher the organ enrichment, the stronger the ability to predict the organ enrichment of the amino acid sequence.
[0145] The training steps of the viral amino acid sequence organ enrichment prediction module are as follows: the feature information of a viral amino acid sequence is simultaneously input into multiple parallel convolution blocks, the features extracted from multiple convolution blocks are spliced in the hidden layer, the residual module uses a residual connection network to retain the original features with a certain probability while using the convolution layer to update the features, the previously extracted features are input into the residual module, and then layer normalization (LN) is used to perform feature aggregation to obtain aggregated features. The aggregated features are activated using the Sigmoid activation function to obtain sequence information weights, and the aggregated features are delinearized using the relu activation function to obtain activation features. The sequence information weights are multiplied by the activation features to obtain weighted activation information. The weighted original information is obtained by subtracting 1 from the previously obtained sequence information weight and multiplying it with the previous viral amino acid sequence features. The weighted activation information and the weighted original information are added together as the predicted features, so that the original feature information is passed in with a certain probability while updating the features. The predicted features are input into the temporary backoff method and then into the linear layer. Finally, the activation function leakyrelu is used as the output of the viral amino acid sequence organ enrichment prediction. After Q (1≤Q≤50000, Q is a positive integer) rounds of iteration, the training of the viral amino acid sequence organ enrichment prediction module is completed.
[0146] After receiving all the generated viral amino acid sequences, the viral amino acid sequence organ enrichment degree prediction module uses the trained viral amino acid sequence organ enrichment degree prediction module to predict the viral amino acid sequence organ enrichment degree of each viral amino acid sequence, obtains the corresponding viral amino acid sequence organ enrichment degree, and sorts the viral amino acid sequences from high to low according to the organ enrichment degree, and filters out sequences with low viral amino acid sequence organ enrichment degree.
[0147] S5, a scoring module for the degree of binding to the target based on the amino acid sequence,
[0148] Based on the amino acid sequences obtained by the above module screening, the system scores and counts the degree of binding between each sequence and the target, sorts the degree of binding between the amino acid sequence and the target from high to low, selects the top K (K = 1 to 100,000) amino acid sequences as candidate sequences, and conducts synthesis experiments and verification using the biosynthesis method. If the function of the final candidate sequence does not meet the user's requirements, the generative model can be used to regenerate a batch of sequences and re-execute the above steps until a sequence that satisfies the user is obtained. The number of cycles is n times (n = 1 to 1,000,000);
[0149] S6, verifying the viral amino acid sequence in the target virus amino acid sequence library;
[0150] S61, performing experiments on the viral amino acid sequences in the target viral amino acid sequence library to obtain experimental data, as shown in FIG7 :
[0151] If the number of viral amino acid sequences in the target viral amino acid sequence library is greater than 100, each viral amino acid sequence in the library is connected to a barcode to construct an amino acid sequence pool. If the number of viral amino acid sequences is less than or equal to 100, a viral vector is constructed and packaged separately for each amino acid sequence. After replacing the viral amino acid sequence with the appropriate site of the target plasmid, a viral plasmid is obtained. The viral plasmid frequency is calculated by extracting plasmid DNA from the viral plasmid and performing high-throughput sequencing. The viral plasmid is then transfected into cells, and the virus is purified after virus production in the cells. After virus purification, the virus quality purity is tested. After extracting viral DNA, the virus frequency is calculated by high-throughput sequencing, and the virus biological activity test in vitro is performed. After that, the virus biological activity test in vivo is performed. The viral amino acid sequence production fitness is calculated based on the viral plasmid frequency and the virus frequency. The viral amino acid sequence production fitness is compared with the predicted viral amino acid sequence production fitness through evaluation indicators to verify the effectiveness of the viral amino acid sequence generation and screening device for viral amino acid sequence generation. Finally, a viral vector with high specificity and high viral amino acid sequence production fitness verified by biological experiments is obtained.
[0152] S62, determine the evaluation indicators,
[0153] In terms of selecting evaluation indicators to evaluate the prediction performance of the model, this study involves regression prediction tasks, so the root mean square error, Pearson correlation coefficient, Spearman correlation coefficient and determination coefficient R are selected. 2 To evaluate the performance of the prediction of the production fitness data of the viral nucleocapsid; the root mean square error describes the distance between the predicted value and the true value; the Pearson correlation coefficient and the Spearman correlation coefficient describe the correlation between the predicted value and the true value, where the Pearson correlation coefficient describes the linear correlation between the two values, and the Spearman correlation coefficient is the rank form of the Pearson correlation coefficient, which describes the correlation between the two variables (for example, when one variable increases, the other variable also increases), which is related to the monotonicity of the function; the coefficient of determination R 2 is a dimensionless score describing the effectiveness of the model, which compares the predictions to random guessing based on the mean of the true values.
[0154] The present disclosure discloses a method for generating and screening amino acid sequences based on pre-training learning and rational design according to function, comprising the following steps: S0 pre-trains the model based on a large number of known amino acid sequences to learn the biological "grammar" and "semantic" information of the amino acid sequence; S1 uses the data set constructed in the experimental data to train the viral amino acid sequence generation device based on the model obtained by pre-training; S2 uses the amino acid sequence generated by the upstream model as the evolutionary target, and further transforms the generated amino acid sequence through a rational design process to further design an amino acid sequence with a specific function; S3 uses the sequence obtained by evolution to screen and narrow the range of preferred sequences for experimental verification according to the function of the sequence; S4 uses the sequence obtained by S3 to screen the sequence according to the degree of binding between the sequence and the target, and finally obtains the preferred sequence for experimental verification. The system includes four major parts: cloud computing and supercomputing platform, amino acid sequence design and development laboratory, amino acid sequence generation and screening device, and algorithm result verification laboratory. Its specific connection structure is as follows:
[0155] A dataset for pre-training a model is constructed from a large amount of unlabeled amino acid sequence data (tens of millions of amino acid sequences already in the database) to learn the grammatical and semantic information of the amino acid sequence; a large language model is used to perform preliminary training on the pre-training dataset constructed from a large amount of unlabeled amino acid sequence data to obtain a large pre-trained model with semantic extraction capabilities; based on the pre-trained model, a dataset for a downstream specific design task (viral amino acid sequence) is used to preliminarily "conditionally generate" amino acid sequences with high synthesizability. The parameter setting module is used to set the length and number of viral amino acid sequences to be generated, and the parameters are input into the specific amino acid sequence generation module;
[0156] The specific amino acid sequence generation module generates viral amino acid sequences by learning the logical structure and grammatical semantic features between existing amino acid sequences that can generate viruses; based on existing positive sample data with specific functions, the amino acid sequence output by the generation model is further optimized and transformed using heuristic algorithms such as annealing algorithm, genetic algorithm, and swarm intelligence optimization algorithm, so that the sequence space evolves in the direction of positive samples, and the optimized sequence is used as the input of the downstream filter.
[0157] According to the required amino acid sequence function, the optimized sequence is functionally predicted (viral production fitness and organ enrichment of viral amino acid sequences), and amino acid sequences with poor functions are filtered out. The working process of the specific amino acid sequence generation module is as follows: randomly generate an amino acid, characterize the randomly generated amino acid, input the amino acid feature into the trained long short-term memory neural network, and the long short-term memory neural network generates the viral amino acid sequence according to the set amino acid sequence length. Specifically, after the amino acid feature is input into the trained long short-term memory neural network, the long short-term memory neural network will generate the next amino acid based on the current amino acid feature. The generated amino acid is then input into the long short-term memory neural network to generate the next position amino acid, and so on until the set number of amino acids is generated. These generated amino acids are then spliced in the dimension of the sequence length, and the above steps are repeated. When the number of viral amino acid sequences reaches the preset generation number, the specific amino acid sequence generation module stops generating the viral amino acid sequence.
[0158] The filter module includes a viral amino acid sequence production fitness prediction module and a viral amino acid sequence organ enrichment degree prediction module. The viral amino acid sequence production fitness prediction module, after receiving the viral amino acid sequence, performs viral amino acid sequence production fitness prediction on the viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness. The viral amino acid sequence organ enrichment degree prediction module predicts the viral amino acid sequence organ enrichment degree based on the viral amino acid sequence. The higher the organ enrichment degree, the stronger the enrichment degree of the amino acid sequence in the organ.
[0159] The scoring and ranking module screens the sequences based on the degree of binding between the sequences and the targets based on the sequences obtained by the filter module, and finally obtains the preferred sequences for experimental verification.
[0160] The present disclosure also discloses a system for viral amino acid sequence analysis, as shown in FIG1 , which includes four modules: a cloud computing and supercomputing platform; a viral vector design and development laboratory; a viral amino acid sequence generation and screening device; and an algorithm result verification laboratory.
[0161] The cloud computing and supercomputing platform accepts user or administrator operation instructions based on the I / O interface and assigns corresponding permissions to them. It is responsible for collecting and managing online information data related to virus sequence design, transmitting local experimental data to the storage unit, and allocating the current computing power resources according to the priority order of the computing tasks submitted by the user through the computing unit to execute the corresponding tasks;
[0162] Viral vector design and development laboratory, used to obtain viral amino acid sequences through experiments, as well as the generation fitness and organ enrichment of each viral amino acid sequence;
[0163] Pre-training model: The model is pre-trained based on a large number of known amino acid sequences to learn the biological "grammar" and "semantic" information of amino acid sequences;
[0164] The sequence generation module trains the viral amino acid sequence generation device based on the pre-trained model using the dataset constructed from the experimental data;
[0165] The rational design module, based on the amino acid sequence generated by the upstream model, takes the amino acid sequence with specific function as the evolution target, and further transforms the generated amino acid sequence through the rational design process to further design the amino acid sequence with specific function;
[0166] The filter module, based on the evolved sequences, screens the sequences according to their functions to narrow down the range of preferred sequences for experimental verification;
[0167] The scoring and ranking module, based on the sequences obtained by the filter module, screens the sequences according to the degree of binding between the sequences and the targets, and finally obtains the preferred sequences for experimental verification.
[0168] The algorithm result verification laboratory uses experiments to obtain the production fitness and organ enrichment of the viral amino acid sequences stored in the target viral amino acid sequence library, verifying the effectiveness and accuracy of the model in generating viral amino acid sequences with high production fitness and high organ enrichment.
[0169] The present disclosure has also achieved good prediction accuracy in the prediction of transmembrane peptide sequences, antimicrobial peptide sequences, and antibody sequences. Based on the principles of the prediction and screening methods, the present disclosure can be applied to the prediction of any type of functional amino acid sequence.
[0170] Example 2 Generation of oligomeric peptide sequences for AAV transformation based on AI
[0171] The present invention discloses a machine learning-based amino acid sequence generation, design, and screening method for predicting AAV sequences and their polymorphic site amino acid sequences. The embodiment performs the following experimental steps for generating an amino acid sequence inserted between amino acids 588 and 589 of the AAV9 capsid hypervariable region VIII.
[0172] As shown in Figure 1, the amino acid sequence semantic learning model based on the generalized dataset optimizes the design module of the virus based on rational design, generates and experimentally verifies the pathogenic amino acid sequences, and aims to find amino acid sequences with high specificity and high diversity. The specific steps are as follows:
[0173] S1, constructing a dataset for training the model from experimental data, as shown in Figure 2, includes the following steps:
[0174] S11, steps for constructing viral plasmid library,
[0175] Randomly generate M (M is 10,000 to 200,000) amino acid sequences of length N (N is 1 to 500). Each amino acid sequence is connected to a specific barcode, and all barcode-connected amino acid sequences are pooled together to construct an amino acid sequence pool. The amino acid sequence library in the amino acid sequence pool is used to replace some sites in the target plasmid to generate viral plasmids. Different amino acid sequences or replacements of different sites in the target plasmid will result in different viral plasmids, and these different viral plasmids together constitute the viral plasmid library.
[0176] Each amino acid sequence is connected to a specific barcode for subsequent high-throughput sequencing.
[0177] The viral plasmids in the current viral plasmid library are only viral plasmids that may become viruses. They still need to undergo viral vector packaging, virus quality purity testing, and virus in vitro biological activity testing to determine whether the obtained viral plasmids can eventually become viruses.
[0178] S12, viral amino acid sequence production fitness data collection step
[0179] The viral plasmid frequency was calculated by high-throughput sequencing of the plasmid DNA extracted from the viral plasmid. Specifically, the frequency of occurrence of a single viral amino acid sequence in the plasmid was calculated by high-throughput sequencing of the plasmid DNA extracted from the replaced plasmid to obtain the viral plasmid frequency.
[0180] At the same time, after the viral plasmid is transfected into the cells, the virus is purified and the viral DNA is extracted, and then high-throughput calculation is performed to obtain the virus frequency. The specific process is: the viral plasmid is transfected into the cells, the virus is produced in the cells (that is, the virus is spread), and then the virus is purified. After the virus is purified, the virus quality purity is tested, and the viral DNA is extracted and the virus frequency is calculated by high-throughput sequencing; the virus frequency refers to the number of times the same viral amino acid sequence appears in the virus after high-throughput sequencing.
[0181] Finally, the viral plasmid frequency and viral frequency are used to calculate the production fitness of a single viral amino acid sequence. Production fitness is a quantitative representation of the ability of a viral amino acid sequence to generate a virus. A higher production fitness indicates a stronger ability of the amino acid sequence to generate a virus.
[0182] S13 build dataset
[0183] The dataset uses viral amino acid sequences as samples, and the production fitness of the corresponding viral amino acid sequences serves as the dataset label. This example primarily uses viral amino acid sequences that have been cleaned of premature stop codons and those with sequencing errors detected during high-throughput sequencing. The samples in the dataset are divided into a training set: validation set: test set ratio of 7:1:2.
[0184] To avoid imbalanced training data and ensure the accuracy and research significance of the model training process, we first need to check the normality of the overall distribution of the viral amino acid sequence production fitness data labels. For unevenly distributed data, we use strategies such as normalization, downsampling, and gradient clipping to balance the data to ensure that the model has no bias during the learning process. Figure 6 shows the probability density distribution of viral amino acid sequence production fitness when constructing the dataset used for model training from experimental data.
[0185] S2, constructing a dataset for pre-training models from a large amount of unlabeled amino acid sequence data (tens of millions of amino acid sequences already in the database) to learn the grammatical and semantic information of amino acid sequences;
[0186] S21: Input the viral amino acid sequences in the dataset into a data preprocessing module to perform feature encoding on the viral amino acid sequences, thereby obtaining viral amino acid sequence features. The viral amino acid sequence features include the characteristics of each amino acid in the viral amino acid sequence. The data preprocessing module can use existing strategies such as word embedding, graph representation, or positional encoding to represent the feature information of the viral amino acid sequences.
[0187] A large language model is initially trained on a pre-training dataset constructed from a large amount of unlabeled amino acid sequence data to obtain a pre-trained large model with semantic extraction capabilities;
[0188] S3, based on the pre-trained model, uses the dataset of downstream specific design tasks (viral amino acid sequences) to preliminarily "conditionally generate" amino acid sequences with high synthesizability, including the following steps:
[0189] S31 performs feature encoding on the amino acid sequences in the dataset;
[0190] The viral amino acid sequences in the dataset are input into the data preprocessing module, which performs feature encoding on the viral amino acid sequences to obtain viral amino acid sequence features. These features include the characteristics of each amino acid in the viral amino acid sequence. The data preprocessing module can use existing strategies such as word embedding, graph representation, or positional encoding to represent the feature information of the viral amino acid sequences.
[0191] S32 trains the specific amino acid sequence generation module;
[0192] There are currently 20 known amino acids. The amino acid sequence is composed of the combination of amino acids, but only some of the combinations of amino acids can generate viruses.
[0193] The specific amino acid sequence generation module generates viral amino acid sequences by learning the logical structure and grammatical semantic features of existing amino acid sequences that can generate viruses. Specificity refers to conforming to the logical structure and grammatical semantic features of amino acid sequences that can generate viruses.
[0194] In this embodiment, the specific amino acid sequence generation module adopts a long short-term memory neural network. The training steps for the specific amino acid sequence generation module are as follows: select viral amino acid sequences from the training set according to the set batch number, and the batch number refers to the number of viral amino acid sequences that can be input into the specific amino acid sequence generation module at one time; after the viral amino acid sequence is text-encoded, a number 0 is added to the front of each viral amino acid sequence and the last amino acid is truncated, and all the features of the processed viral amino acid sequence are used as the input of the long short-term memory neural network, and the original viral amino acid sequence after text encoding without other processing is used as the target to be generated by the long short-term memory neural network. The internal workflow of the long short-term memory neural network is that an input viral amino acid sequence will be split into multiple units. amino acids, and uses the previous amino acid as the input of the long short-term memory neural network to generate the amino acid at the next position. The next generated amino acid is used as the input of the long short-term memory neural network to generate the amino acid at the next position, and so on, until an amino acid sequence with the same length as the target amino acid sequence is generated, and the generated amino acid sequence is input into the linear layer for amino acid sequence feature extraction, and the extracted features are input into the activation function softmax to obtain the final generated amino acid sequence; the selected viral amino acid sequence features are compared with the generated amino acid sequence features, the loss function is calculated, and back propagation is performed, and the viral amino acid sequences in the entire training set are cyclically selected according to the batch size to complete the above amino acid generation steps, and the above cycle is repeated M times until the loss function is stable, and the model training parameters are saved. Therefore, when using a data set containing N viral amino acid sequences to train the specific amino acid sequence generation module, it is necessary to go through (N / batch number)*M rounds of iterations before the specific amino acid sequence generation module training is completed;
[0195] Wherein, 1≤N≤50000000, N is a positive integer, 1≤M≤50000, M is a positive integer.
[0196] S33 uses the trained viral amino acid sequence generation and screening device to generate viral amino acid sequences, sets the length of the viral amino acid sequences and the number of sequences to be generated, and then inputs the parameters into a specific amino acid sequence generation module;
[0197] The execution steps of the specific amino acid sequence generation module are as follows: randomly generate an amino acid, characterize the randomly generated amino acid, input the amino acid feature into the trained long short-term memory neural network, and the long short-term memory neural network generates the viral amino acid sequence according to the set amino acid sequence length. Specifically, after the amino acid feature is input into the trained long short-term memory neural network, the long short-term memory neural network will generate the next amino acid according to the current amino acid feature, and the generated amino acid will be input into the long short-term memory neural network to generate the next position amino acid, and so on until the set number of amino acids of the length are generated, and then these generated amino acids are spliced in the dimension of the sequence length, and the above steps are repeated. When the number of viral amino acid sequences reaches the preset generation number, the specific amino acid sequence generation module stops generating the viral amino acid sequence.
[0198] All generated viral amino acid sequences are sent to the viral amino acid sequence production fitness prediction module.
[0199] In addition to the virus-specific sequence generation module described in this embodiment using a long short-term memory neural network to generate an adversarial network architecture to generate the viral amino acid sequence, a variable differential autoencoder (VAE) and a reliable sequence generation algorithm such as a diffusion model can also be used to achieve the generation of virus-specific sequences. However, the variable differential autoencoder architecture includes two parts: an encoder and a decoder. The encoder compresses the original data into a low-dimensional feature vector by using a neural network, and the decoder restores the compressed feature vector through a neural network and generates data that conforms to the original distribution. Although the generated data diversity is high, its generation result is usually relatively vague. Without the adversarial process of the discriminator, the quality is difficult to guarantee, and there are problems such as the posterior distribution being assumed to be a decomposable Gaussian distribution, which is based on the assumption of the encoder. The diffusion model includes a diffusion process and an inverse diffusion process. The diffusion process, under the condition of a Markov chain, continuously adds Gaussian noise to the original data, so that the original data gradually tends to a Gaussian distribution, and finally obtains global diffusion data, while the inverse diffusion process restores the global diffusion data through Gaussian noise and generates the original data. The generative model based on this architecture can generate relatively robust sequence information, but due to current technological limitations, the generalization ability of the generated data still needs to be improved. Therefore, the virus-specific sequence generation module in this embodiment is the most suitable choice.
[0200] S4, based on the "conditionally generated" amino acid sequence, optimizes the amino acid sequence through the rational design module. Based on the existing positive sample data with specific functions, heuristic algorithms such as annealing algorithm, genetic algorithm, and swarm intelligence optimization algorithm are used to further optimize the amino acid sequence output by the generation model, so that the sequence space evolves in the direction of positive samples. The optimized sequence is used as the input of the downstream filter. The specific steps are as follows:
[0201] S41 obtains the target amino acid sequence output from the sequence generation module, which is called the primary population G0. Each sequence s i ∈G0 is called an individual in the population, and the current population evolution round number t=0.
[0202] S42 establishes population G t Evaluation function for fitness evaluation.
[0203] Import the sequence s′ from the target amino acid sequence database with specific properties in advance i And the corresponding characteristic information, and normalize the corresponding characteristics to obtain the normalized weight w i When all sequences lack corresponding feature information, a binary classification method is used, and all positive samples are set to w + =1, set all negative samples to w - =-1.
[0204] For each individual s in the population i (i=1,2,...,n), and a sequence s′ in the database j (j=1,2,...,m), the sequence similarity score c between them can be calculated by the global similarity algorithm (Needleman-Wunsch algorithm) or the local similarity algorithm (Smith-Waterman algorithm) ij .
[0205] The similarity c between the individual and the sequence in the database ij The corresponding normalized weight w j After multiplication and accumulation, the sum of weighted sequence similarity scores is obtained. This is the fitness score v of the individual i =∑w j c ij .
[0206] According to the above algorithm, all individuals in the population are calculated separately to obtain the fitness level of all individuals in the population.
[0207] S43 evaluates the fitness of population G0 using the method of S42.
[0208] S44 builds a new generation of population G t+1 Here, different swarm intelligence optimization algorithms have different steps. This paper adopts swarm intelligence optimization algorithms under discrete space of categorical variables, such as genetic algorithm, distribution estimation algorithm, etc., as well as hybrid algorithms that integrate these algorithms. The following uses genetic algorithm, distribution estimation algorithm and their hybrid algorithms as examples.
[0209] Genetic algorithms can generate random mutations, which helps to discover the possible positive effects of new amino acids at certain positions. The specific steps for constructing a new generation of populations are as follows:
[0210] The parent population G t Sort by individual fitness scores from high to low G t ={s t,i |i=1,2,...,n;v t,i >v t,i+1}, according to the set replication rate r c The individuals with the highest individual fitness (n c =n×r c ) is copied into the new generation population, i.e. s t+1,i =s t,i (i=1,2,...,n c ). The remaining individuals are generated by crossover and mutation of the parent population: First, the fitness scores of the individuals are normalized: Then, the normalized individual fitness scores are used as weights, and the t The weight is the probability (that is, the probability of selecting sequence i is p i =v′ t,i ) Extract two sequences Unit point or double locus Crossover recombination obtains recombinant individual s′ t+1,i The recombinant individuals follow the set mutation rate r m Randomly mutate the amino acid at each site and add the resulting sequence to the new generation population G t+1 The crossover, recombination and mutation of the parent population are continued until the number of individuals in the new generation population is the same as that in the parent population. t |=|G t+1 |=n.
[0211] The distribution estimation algorithm constructs a probability model based on the distribution of the parent population in each dimension and samples to generate new individuals. Compared with the crossover operation of the genetic algorithm, the distribution estimation algorithm can change the local structure of the parent individuals to a greater extent, helping to discover new and excellent local structures. The specific steps for constructing the new generation population are:
[0212] Let the library of twenty natural amino acids be AA. t A certain number of excellent individuals n b =n×r b The amino acids at each position are counted to obtain the distribution of different amino acids at each position, and the proportion of each amino acid at each position is calculated (i.e., p ij is the frequency of amino acid j at position i). The generation of new individuals is obtained by sampling the probability model: for any position i = 1, 2, ..., l, according to the proportion p of each amino acid at that position ij (j∈AA) is randomly sampled to obtain the amino acid a at that position i (i.e. P(a i =j)=p ij ). Sampling is continued until the corresponding amino acid is extracted at each position and connected to obtain a new amino acid sequence. Count this individual into the new generation population. Repeated sampling until the number of individuals in the new generation population is the same as the number of individuals in the parent population |G t |=|G t+1 |=n.
[0213] The hybrid algorithm integrates the advantages of the previous algorithm. The specific steps for constructing a new generation of population are as follows:
[0214] The proportion of subpopulation individuals generated by each algorithm is pre-set (0% to 100%) to ensure that the number of subpopulation individuals is Then, all individuals in the population are sorted from high to low in terms of fitness. If the corresponding conditions are met (for example, the fitness of the optimal individual has reached the requirement, or the number of evolutionary rounds has reached the set value), the algorithm ends and the modified target amino acid sequence is output to the filter. Otherwise, it returns to S44 to construct a new generation of population.
[0215] S5, the amino acid sequence optimized by the rational design module is filtered through the target property filter. The specific steps are as follows:
[0216] S51 training virus amino acid sequence to produce fitness prediction module
[0217] The viral amino acid sequence production fitness prediction module can predict the viral amino acid sequence production fitness based on the viral amino acid sequence. The higher the production fitness, the stronger the ability of the amino acid sequence to generate viruses.
[0218] The training steps of the viral amino acid sequence production fitness prediction module are as follows: the feature information of a viral amino acid sequence is simultaneously input into multiple parallel convolution blocks, the features extracted from multiple convolution blocks are spliced in the hidden layer, the residual module uses a residual connection network to retain the original features with a certain probability while using the convolution layer to update the features, the previously extracted features are input into the residual module, and then layer normalization (LN) is used to perform feature aggregation to obtain aggregated features, the aggregated features are activated using the Sigmoid activation function to obtain sequence information weights, and the aggregated features are delinearized using the relu activation function to obtain activation features, the sequence information weights are multiplied by the activation features to obtain weighted activation information, 1 is subtracted from the previously obtained sequence information weights and multiplied by the previous viral amino acid sequence features to obtain weighted original information, the weighted activation information and the weighted original information are added together as the predicted features, so that the original feature information is passed in with a certain probability while updating the features, the predicted features are input into the temporary backoff method and then into the linear layer, and finally the activation function leakyrelu is used as the output of the viral amino acid sequence production fitness prediction. After Q (1≤Q≤50000, Q is a positive integer) rounds of iteration, the training of the viral amino acid sequence production fitness prediction module is completed.
[0219] After receiving all the generated viral amino acid sequences, the viral amino acid sequence production fitness prediction module uses the trained viral amino acid sequence production fitness prediction module to predict the viral amino acid sequence production fitness of each viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness. The viral amino acid sequences are sorted from high to low according to their production fitness, and sequences with low viral amino acid sequence production fitness are filtered out.
[0220] S52 training virus amino acid sequence organ enrichment prediction module
[0221] The viral amino acid sequence organ enrichment prediction module can predict the viral amino acid sequence organ enrichment based on the viral amino acid sequence. The higher the organ enrichment, the stronger the ability to predict the organ enrichment of the amino acid sequence.
[0222] The training steps of the viral amino acid sequence organ enrichment prediction module are as follows: the feature information of a viral amino acid sequence is simultaneously input into multiple parallel convolution blocks, the features extracted from multiple convolution blocks are spliced in the hidden layer, the residual module uses a residual connection network to retain the original features with a certain probability while using the convolution layer to update the features, the previously extracted features are input into the residual module, and then layer normalization (LN) is used to perform feature aggregation to obtain aggregated features. The aggregated features are activated using the Sigmoid activation function to obtain sequence information weights, and the aggregated features are delinearized using the relu activation function to obtain activation features. The sequence information weights are multiplied by the activation features to obtain weighted activation information. The weighted original information is obtained by subtracting 1 from the previously obtained sequence information weight and multiplying it with the previous viral amino acid sequence features. The weighted activation information and the weighted original information are added together as the predicted features, so that the original feature information is passed in with a certain probability while updating the features. The predicted features are input into the temporary backoff method and then into the linear layer. Finally, the activation function leakyrelu is used as the output of the viral amino acid sequence organ enrichment prediction. After Q (1≤Q≤50000, Q is a positive integer) rounds of iteration, the training of the viral amino acid sequence organ enrichment prediction module is completed.
[0223] After receiving all the generated viral amino acid sequences, the viral amino acid sequence organ enrichment degree prediction module uses the trained viral amino acid sequence organ enrichment degree prediction module to predict the viral amino acid sequence organ enrichment degree of each viral amino acid sequence, obtains the corresponding viral amino acid sequence organ enrichment degree, and sorts the viral amino acid sequences from high to low according to the organ enrichment degree, and filters out sequences with low viral amino acid sequence organ enrichment degree.
[0224] S53, a scoring module for the degree of binding to the target based on the amino acid sequence,
[0225] Based on the amino acid sequences obtained by the above module screening, the system scores and statistically analyzes the degree of binding between each sequence and the target, sorts the amino acid sequences from high to low in terms of binding, selects the top K (K = 1 to 100,000) amino acid sequences as candidate sequences, and conducts synthesis experiments and verification using biosynthesis methods;
[0226] S6, verifying the viral amino acid sequence in the target virus amino acid sequence library;
[0227] S61, performing experiments on the viral amino acid sequences in the target viral amino acid sequence library to obtain experimental data, as shown in FIG9 :
[0228] If the number of viral amino acid sequences in the target viral amino acid sequence library is greater than 100, each viral amino acid sequence in the library is connected to a barcode to construct an amino acid sequence pool. If the number of viral amino acid sequences is less than or equal to 100, a viral vector is constructed and packaged separately for each amino acid sequence. After replacing the viral amino acid sequence with the appropriate site of the target plasmid, a viral plasmid is obtained. The viral plasmid frequency is calculated by extracting plasmid DNA from the viral plasmid and performing high-throughput sequencing. The viral plasmid is then transfected into cells, and the virus is purified after virus production in the cells. After virus purification, the virus quality purity is tested. After extracting viral DNA, the virus frequency is calculated by high-throughput sequencing, and the virus biological activity test in vitro is performed. After that, the virus biological activity test in vivo is performed. The viral amino acid sequence production fitness is calculated based on the viral plasmid frequency and the virus frequency. The viral amino acid sequence production fitness is compared with the predicted viral amino acid sequence production fitness through evaluation indicators to verify the effectiveness of the viral amino acid sequence generation and screening device for viral amino acid sequence generation. Finally, a viral vector with high specificity and high viral amino acid sequence production fitness verified by biological experiments is obtained.
[0229] S62, performing in vivo experimental verification on the viral amino acid sequence in the target viral amino acid sequence library to obtain experimental data, as shown in FIG27 :
[0230] If the number of viral amino acid sequences in the target viral amino acid sequence library is greater than 100, each viral amino acid sequence in the library is connected to a barcode to construct an amino acid sequence pool. If the number of viral amino acid sequences is less than or equal to 100, viral vector construction and packaging are performed separately for each amino acid sequence; then the biological activity of the virus in vivo is detected. Histological analysis shows the distribution of the virus in different organs in the animal body. Experiments have found that some newly designed AAV capsids have strong ability to cross the blood-brain barrier, strong tissue specificity, and high transfection efficiency. Ultimately, a viral vector that has been verified by biological experiments and has high specificity, high viral amino acid sequence production adaptability, high tissue specificity, high transfection efficiency, and can cross the blood-brain barrier is obtained.
[0231] S63, performing in vitro experimental verification on the viral amino acid sequence in the target viral amino acid sequence library to obtain experimental data, as shown in FIG28:
[0232] If the number of viral amino acid sequences in the target viral amino acid sequence library is greater than 100, each viral amino acid sequence in the library is connected to a barcode to construct an amino acid sequence pool. If the number of viral amino acid sequences is less than or equal to 100, a viral vector is constructed and packaged separately for each amino acid sequence. Afterwards, the binding ability of the virus and the target protein is verified in vitro, and the binding ability of the virus and the target protein is detected by surface plasmon resonance (SPR) to verify the effectiveness of the viral amino acid sequence generation and screening device for viral amino acid sequence generation. Ultimately, a viral vector with high specificity, high viral amino acid sequence production adaptability, and strong binding ability to the target protein, which has been verified by biological experiments, is obtained.
[0233] S64, determine the evaluation indicators,
[0234] In terms of selecting evaluation indicators to evaluate the prediction performance of the model, this study involves regression prediction tasks, so the root mean square error, Pearson correlation coefficient, Spearman correlation coefficient and determination coefficient R are selected. 2 To evaluate the performance of the prediction of the production fitness data of the viral nucleocapsid; the root mean square error describes the distance between the predicted value and the true value; the Pearson correlation coefficient and the Spearman correlation coefficient describe the correlation between the predicted value and the true value, where the Pearson correlation coefficient describes the linear correlation between the two values, and the Spearman correlation coefficient is the rank form of the Pearson correlation coefficient, which describes the correlation between the two variables (for example, when one variable increases, the other variable also increases), which is related to the monotonicity of the function; the coefficient of determination R 2 is a dimensionless score describing the effectiveness of the model, which compares the prediction to random guessing based on the mean of the true values;
[0235] The present disclosure discloses a method for generating and screening amino acid sequences based on pre-training learning and rational design according to function, comprising the following steps: S1 pre-training the model based on a large number of known amino acid sequences to learn the biological "grammar" and "semantic" information of the amino acid sequence; S2 training the viral amino acid sequence generation device using the data set constructed in the experimental data based on the model obtained by pre-training; S3 based on the amino acid sequence generated by the upstream model, with the amino acid sequence of a specific function as the evolutionary goal, further transforming the generated amino acid sequence through a rational design process to further design an amino acid sequence with a specific function; S4 based on the sequence obtained by evolution, screening according to the function of the sequence to narrow the range of preferred sequences for experimental verification; S5 based on the sequence obtained in S4, screening the sequence according to the degree of binding between the sequence and the target, and finally obtaining the preferred sequence for experimental verification. The system comprises four major parts: cloud computing and supercomputing platform, amino acid sequence design and development laboratory, amino acid sequence generation and screening device, and algorithm result verification laboratory. Its specific connection structure is as follows:
[0236] A dataset for pre-training a model is constructed from a large amount of unlabeled amino acid sequence data (tens of millions of amino acid sequences already in the database) to learn the grammatical and semantic information of the amino acid sequence; a large language model is used to perform preliminary training on the pre-training dataset constructed from a large amount of unlabeled amino acid sequence data to obtain a large pre-trained model with semantic extraction capabilities; based on the pre-trained model, a dataset for a downstream specific design task (viral amino acid sequence) is used to preliminarily "conditionally generate" amino acid sequences with high synthesizability. The parameter setting module is used to set the length and number of viral amino acid sequences to be generated, and the parameters are input into the specific amino acid sequence generation module;
[0237] The specific amino acid sequence generation module generates viral amino acid sequences by learning the logical structure and grammatical semantic features between existing amino acid sequences that can generate viruses; based on existing positive sample data with specific functions, the module uses heuristic algorithms such as annealing algorithm, genetic algorithm, and swarm intelligence optimization algorithm to further optimize the amino acid sequence output by the generation model, so that the sequence space evolves in the direction of positive samples, and the optimized sequence is used as the input of the downstream filter.
[0238] According to the required amino acid sequence function, the optimized sequence is functionally predicted (viral production fitness and organ enrichment of viral amino acid sequences), and amino acid sequences with poor functions are filtered out. The working process of the specific amino acid sequence generation module is as follows: randomly generate an amino acid, characterize the randomly generated amino acid, input the amino acid feature into the trained long short-term memory neural network, and the long short-term memory neural network generates the viral amino acid sequence according to the set amino acid sequence length. Specifically, after the amino acid feature is input into the trained long short-term memory neural network, the long short-term memory neural network will generate the next amino acid based on the current amino acid feature. The generated amino acid is then input into the long short-term memory neural network to generate the next position amino acid, and so on until the set number of amino acids is generated. These generated amino acids are then spliced in the dimension of the sequence length, and the above steps are repeated. When the number of viral amino acid sequences reaches the preset generation number, the specific amino acid sequence generation module stops generating the viral amino acid sequence.
[0239] The filter module includes a viral amino acid sequence production fitness prediction module and a viral amino acid sequence organ enrichment degree prediction module. The viral amino acid sequence production fitness prediction module, after receiving the viral amino acid sequence, performs viral amino acid sequence production fitness prediction on the viral amino acid sequence to obtain the corresponding viral amino acid sequence production fitness. The viral amino acid sequence organ enrichment degree prediction module predicts the viral amino acid sequence organ enrichment degree based on the viral amino acid sequence. The higher the organ enrichment degree, the stronger the enrichment degree of the amino acid sequence in the organ.
[0240] The scoring and ranking module screens the sequences based on the degree of binding between the sequences and the targets based on the sequences obtained by the filter module, and finally obtains the preferred sequences for experimental verification.
[0241] The present disclosure also discloses a system for viral amino acid sequence analysis based on machine learning, as shown in FIG1 , which includes four modules: a cloud computing and supercomputing platform; a viral vector design and development laboratory; a viral amino acid sequence generation and screening device; and an algorithm result verification laboratory.
[0242] The cloud computing and supercomputing platform accepts user or administrator operation instructions based on the I / O interface and assigns corresponding permissions to them. It is responsible for collecting and managing online information data related to virus sequence design, transmitting local experimental data to the storage unit, and allocating the current computing power resources according to the priority order of the computing tasks submitted by the user through the computing unit to execute the corresponding tasks;
[0243] Viral vector design and development laboratory, used to obtain viral amino acid sequences through experiments, as well as the generation fitness and organ enrichment of each viral amino acid sequence;
[0244] Pre-training model: The model is pre-trained based on a large number of known amino acid sequences to learn the biological "grammar" and "semantic" information of amino acid sequences;
[0245] The sequence generation module trains the viral amino acid sequence generation device based on the pre-trained model using the dataset constructed from the experimental data;
[0246] The rational design module, based on the amino acid sequence generated by the upstream model, takes the amino acid sequence with specific function as the evolution target, and further transforms the generated amino acid sequence through the rational design process to further design the amino acid sequence with specific function;
[0247] The filter module, based on the evolved sequences, screens the sequences according to their functions to narrow down the range of preferred sequences for experimental verification;
[0248] The scoring and ranking module, based on the sequences obtained by the filter module, screens the sequences according to the degree of binding between the sequences and the targets, and finally obtains the preferred sequences for experimental verification.
[0249] The algorithm result verification laboratory uses the viral amino acid sequences stored in the target viral amino acid sequence library to obtain the production fitness, organ enrichment and target protein binding ability of the corresponding viral amino acid sequences through experiments, verifying the effectiveness and accuracy of the model in generating viral amino acid sequences with high production fitness and high organ enrichment. For the target protein here, we chose host cell proteins expressed in specific organs (in this embodiment, the affinity of the modified AAV and the host cell protein LY6C1 expressed on the mouse blood-brain barrier was measured and recorded as Pred_Tar_Fit score). The stronger the interaction between the viral amino acid sequence and the protein target, the stronger its ability to target a specific organ.
[0250] Through the above examples, the list of oligomeric peptide sequences for AAV transformation obtained in the present disclosure can be referred to in Tables 1 and 2 described below. According to the ranking given by the scoring module of this embodiment, in theoretical calculations, the higher the Pred_Tar_Fit score, the stronger the blood-brain barrier penetration ability should be. Table 2 shows the top-ranked sequences, and the visible pattern is that the sequences have GYSS at the C-terminus.
[0251] In the table, Prod_Fit and Pred_VS_Score represent the production fitness scores, which indicate the ability of the capsid protein to form the virus, and Target_Enri represents the prediction result score of the virus target protein binding ability.
[0252] Table 1
[0253] Table 2
[0254] Example 3 Wet experiment 1: Several sequences were randomly selected to synthesize AAV sequences and verify the inclusion effect of capsid formation
[0255] According to the sequence corresponding to the experimental number in Table 3, based on the well-known conventional experimental method, the screening sequence was inserted into the capsid protein sequence amino acids 588 and 589 using the pAAV-2 / 9n plasmid, and the core plasmid and auxiliary plasmid with mCherry fluorescence were used to perform three-plasmid transfection using HEK293 cells, and the virus was purified by gradient density centrifugation. (For more specific operation methods, please refer to Variants of the adeno-associated virus serotype 9 with enhanced penetration of the blood-brain barrier in rodents and primates, Nat Biomed Eng. 2022 Nov; 6(11): 1257-1271).
[0256] The amino acid sequence of the capsid protein is selected from SEQ ID NO: 1 or SEQ ID NO: 2:
[0257] Subsequently, viral titers were measured by qPCR, and it was found that some sequences selected by the algorithm had strong packaging capacity compared to wild-type AAV9, which is of great value in industrial production. AAV.ALICE-N2 and AAV.ALICE-N6 have high production suitability (high viral titers).
[0258] Table 3
[0259] Example 4 Wet experiment 2: Confirmation of blood-brain barrier penetration effect
[0260] AAV.ALICE-N2 and AAV.ALICE-N6 samples were selected for animal experiments to observe organ enrichment.
[0261] Mice were anesthetized with isoflurane and the virus was injected into the retroorbital vein using a 29G insulin needle. After surgery, erythromycin eye ointment was applied to the injection site to prevent infection.
[0262] Three weeks after surgery, the heart was perfused with PBS and 4% paraformaldehyde. The heart, liver, spleen, lungs, kidneys, and head, spine, and legs were removed and soaked in 4% paraformaldehyde overnight. The next day, the brain, eyeballs, optic nerves, spinal cord, dorsal root ganglia, and sciatic ganglia were removed. All tissues were then soaked in a 15% sucrose-PBS solution for one day. On the third day, the solution was replaced with a 30% sucrose-PBS solution for one day. On the fourth day, the tissues were embedded in OCT embedding medium and stored frozen at -80°C.
[0263] Then the brain, heart, liver, spleen, lungs and kidneys were cut into 60um float slices and stored in antifreeze at -20℃. The spinal cord, muscles, eyeballs and nerves were cut into 20um patches and stored at -80℃.
[0264] Wash three times with PBS, stain with 0.5 μg / ml DAPI-PBS solution, and then wash again with PBS. After bleaching, mount the slide flatly on a glass slide. First, aspirate the PBS, rinse briefly with pure water, and once dry, mount the slide with anti-fluorescence decay mounting medium and a clean coverslip. The mounting procedure is similar to the subsequent steps after bleaching.
[0265] Whole-tissue images were taken using a VS120 microscope, while selected tissue sections were photographed using a Nikon N1R confocal microscope. The results showed that the sequences selected by the algorithm showed significant enrichment in the brain compared to AAV9, with minimal enrichment in tissues such as the heart, spleen, lung, kidney, muscle, and spinal cord. This demonstrates that the sequences selected by the algorithm have strong blood-brain barrier cross-potentiality and strong tissue specificity (see Figures 27-29).
[0266] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present disclosure. It should be understood that the above are only specific embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent substitutions, etc. made within the spirit and principles of the present disclosure should be included in the scope of protection of the present disclosure. Industrial Applicability
[0267] The technology disclosed in the present invention can provide a general method and system for virtual screening and prediction of amino acid sequences as needed. Based on this method and system, excellent results have been achieved in the AAV virus capsid screening process. The key sequences predicted by virtual screening have been confirmed to improve the performance of AAV virus capsids. The method and amino acid sequence provided by the present invention have broad industrial application prospects.
Claims
1. A method for generating and screening viral amino acid sequences, characterized in that: The specific steps are as follows: S1 pre-training step, building a data set for pre-training model based on the unlabeled amino acid sequence database to learn the grammatical semantic information of the amino acid sequence; S2 conditional generation step, based on the pre-trained model, uses the viral amino acid sequence dataset to preliminarily generate amino acid sequences. S3 rational design step, based on the "conditionally generated" amino acid sequence, optimizes the amino acid sequence through the rational design module, and further optimizes the amino acid sequence output by the generation model based on the existing positive sample data with the required function using the heuristic algorithm, so that the sequence space evolves in the direction of the positive sample and optimizes the sequence, including the following steps: S31 obtains the amino acid sequence as population G t , t is 0 or a positive integer. In the first generation of population, t is 0, i.e., G0. S32 establishes population G t An evaluation function that performs fitness evaluation, where fitness represents the similarity to a target amino acid sequence with the desired properties, S33 for population G t The S32 method is used to evaluate the fitness. S34 uses a heuristic algorithm to t Each amino acid site is changed to construct a new generation of population G t+1 , G t To G t+1 Defined as one round of evolution, S35 loops through S33 and S34, presets the maximum number of evolutionary rounds and the required value of individual fitness, stops the loop after reaching the preset value, sorts all individuals in the population from high to low in fitness, and outputs the optimized and modified target amino acid sequence. In the S4 filtering step, the function of the viral amino acid sequence in the target viral amino acid sequence library is predicted according to the required amino acid sequence function, and the amino acid sequence with poor function is filtered out.
2. The method for generating and screening viral amino acid sequences according to claim 1, characterized in that: The heuristic algorithm in step S34 is selected from one of annealing algorithm, genetic algorithm, swarm intelligence optimization algorithm or a combination thereof.
3. The method for generating and screening viral amino acid sequences according to claim 1, characterized in that: Before S1 The following steps are included: S0 dataset construction step, the step of constructing a dataset for training the model from experimental data, and / or, S4 also includes one of the following steps: S5 scoring and sorting step, the amino acid sequences are scored and sorted according to the degree of binding between the viral amino acid sequences and the target, and the viral amino acid sequences with high scores and specific functions are selected. If the function of the preferred sequence finally obtained does not meet the user's needs, a batch of sequences can be regenerated using the generation model and the above steps can be re-executed until a sequence that satisfies the user is obtained. S6 Experimental Verification Module The viral amino acid sequences with specific functions output by the scoring and ranking module were experimentally synthesized and their functionality was verified.
4. The method for generating and screening viral amino acid sequences according to claim 2, characterized in that: S0 includes the following steps: S01, a viral plasmid library construction step, randomly generating M amino acid sequences of length N, connecting each amino acid sequence to a specific barcode, and gathering all the amino acid sequences connected to the barcodes together to construct an amino acid sequence pool; using the amino acid sequence library in the amino acid sequence pool to replace part of the target plasmid sites to obtain viral plasmids, different amino acid sequences or replacement of different sites of the target plasmid will obtain different viral plasmids, and different viral plasmids together constitute a viral plasmid library, wherein M is preferably 10,000 to 200,000, and N is preferably 1 to 100; S02, a step of collecting fitness data of viral amino acid sequence production, extracting plasmid DNA from viral plasmids and performing high-throughput sequencing to calculate the frequency of viral plasmids, specifically: extracting plasmid DNA from replaced plasmids and calculating the frequency of occurrence of a single viral amino acid sequence in the plasmid by high-throughput sequencing to obtain the frequency of viral plasmids; S03 constructs a data set, which uses viral amino acid sequences as samples and the production fitness of the corresponding viral amino acid sequences as labels of the data set; the samples in the data set are divided according to the ratio of training set: validation set: test set = 5~8:1~3:1~2.
5. The method for generating and screening viral amino acid sequences according to claim 1, characterized in that: In step S1, the viral amino acid sequence in the data set is input into the data preprocessing module to perform feature encoding on the viral amino acid sequence to obtain the viral amino acid sequence features; The viral amino acid sequence characteristics include the characteristics of each amino acid in the viral amino acid sequence; The data preprocessing module uses existing word embedding, graph representation or position encoding strategies to characterize the characteristic information of the virus amino acid sequence; The large language model is used for preliminary training on a pre-training dataset constructed from a large amount of unlabeled amino acid sequence data to obtain a pre-trained large model with semantic extraction capabilities.
6. The method for generating and screening viral amino acid sequences according to claim 1, characterized in that: The S2 step includes the following steps: S21 performs feature encoding on the amino acid sequences in the data set; Inputting the virus amino acid sequence in the data set into the data preprocessing module to perform feature encoding on the virus amino acid sequence to obtain the virus amino acid sequence features; S22 trains the specific amino acid sequence generation module; Learn the logical structure and grammatical semantic features of the existing amino acid sequences that can generate viruses to generate viral amino acid sequences; S23 uses the trained virus amino acid sequence generation and screening device to generate virus amino acid sequences, sets the length of the virus amino acid sequence and the number of sequences to be generated, and randomly generates amino acid sequences. The generation of virus-specific sequences is achieved using any generation algorithm selected from a long short-term memory neural network to generate an adversarial network architecture, a variational autoencoder VAE, and a diffusion model Diffusion Model.
7. The method for generating and screening viral amino acid sequences according to claim 1, characterized in that: Step S4 includes the step of filtering and screening using at least one prediction module in the following S4a and S4b. S4a trains viral amino acid sequences to produce fitness prediction modules. The virus amino acid sequence production fitness prediction module predicts the virus amino acid sequence production fitness based on the virus amino acid sequence. The higher the production fitness, the stronger the ability of the amino acid sequence to generate viruses. After receiving all the generated viral amino acid sequences, the viral amino acid sequence production fitness prediction module uses the trained viral amino acid sequence production fitness prediction module to predict the viral amino acid sequence production fitness of each viral amino acid sequence, obtains the corresponding viral amino acid sequence production fitness, sorts the viral amino acid sequences from high to low according to the production fitness, and filters out the sequences with low viral amino acid sequence production fitness; S4b trains a virus amino acid sequence organ enrichment prediction module. The virus amino acid sequence organ enrichment prediction module predicts the virus amino acid sequence organ enrichment based on the virus amino acid sequence. The higher the organ enrichment, the stronger the ability of the amino acid sequence organ enrichment. After receiving all the generated viral amino acid sequences, the viral amino acid sequence organ enrichment prediction module uses the trained viral amino acid sequence organ enrichment prediction module to predict the viral amino acid sequence organ enrichment for each viral amino acid sequence, obtains the corresponding viral amino acid sequence organ enrichment, sorts the viral amino acid sequences from high to low according to their organ enrichment, and filters out sequences with low viral amino acid sequence organ enrichment.
8. The method for generating and screening viral amino acid sequences according to claim 1, characterized in that: In step S5, the scoring module based on the degree of binding between the amino acid sequence and the target performs scoring statistics on the degree of binding between each sequence and the target based on the amino acid sequence obtained by screening, and ranks and scores the degree of binding between the amino acid sequence and the target from high to low. In step S6, The following steps are included: S61, conducting experiments on the virus amino acid sequences in the target virus amino acid sequence library to obtain experimental data, An amino acid sequence pool is constructed, and a viral vector is constructed and packaged separately for each amino acid sequence. After the virus is purified, the virus quality purity is tested. After the viral DNA is extracted, the virus frequency is calculated by high-throughput sequencing, and the in vitro biological activity of the virus is tested. Then, the in vivo biological activity of the virus is tested. The production fitness of the viral amino acid sequence is calculated based on the viral plasmid frequency and the viral frequency, and compared with the predicted viral amino acid sequence production fitness through evaluation indicators to verify the effectiveness of the viral amino acid sequence generation and screening device in generating viral amino acid sequences. S62, determine the evaluation index, The root mean square error, Pearson correlation coefficient, Spearman correlation coefficient and determination coefficient R were selected. 2 To evaluate the performance of predictions based on production fitness data of viral nucleocapsids.
9. The method for generating and screening viral amino acid sequences according to claim 1, characterized in that: In step S32, a sequence s′ is imported from a target amino acid sequence database with desired properties. i And the corresponding characteristic information, and normalize the corresponding characteristics to obtain the normalized weight w i When all sequences lack corresponding feature information, a binary classification method is used, and all positive samples are set to w + =1, set all negative samples to w - = -1, for each individual s in the population i (i=1,2,...,n), and a sequence s′ in the database j (j=1,2,...,m), and calculate the sequence similarity score c between them through the global similarity algorithm (Needleman-Wunsch algorithm) or the local similarity algorithm (Smith-Waterman algorithm) ij , the similarity c between the individual and the sequence in the database ij The corresponding normalized weight w j After multiplication and accumulation, the sum of the weighted sequence similarity scores is obtained, which is the fitness score v of the individual i =∑w j c ij .
10. The method for generating and screening viral amino acid sequences according to claim 1, characterized in that: The heuristic algorithm in step S34 is a swarm intelligence optimization algorithm, which includes a genetic algorithm, a distribution estimation algorithm, or a hybrid algorithm integrating the genetic algorithm and the distribution estimation algorithm.
11. The method for generating and screening viral amino acid sequences according to claim 1, characterized in that: The heuristic algorithm in step S34 is a genetic algorithm in the swarm intelligence optimization algorithm.
12. A system for generating and screening amino acid sequences, characterized in that: It includes four major modules: cloud computing and supercomputing platform, vector design and development laboratory, amino acid sequence generation and screening device, and algorithm result verification laboratory; among which: The cloud computing and supercomputing platform accepts the operation instructions of users or administrators based on the I / O interface and assigns corresponding permissions to them. It is responsible for collecting and managing online information data related to sequence design, transmitting local experimental data to the storage unit, and allocating the current computing power resources according to the priority order of the computing tasks submitted by users through the computing unit to execute the corresponding tasks; Vector design and development laboratory, used to obtain amino acid sequences through experiments, as well as the fitness of each amino acid sequence and the degree of organ enrichment; Pre-trained model: pre-train the model based on a large number of known amino acid sequences to learn the biological "grammar" and "semantics" information of amino acid sequences; A sequence generation module, based on the pre-trained model, uses a data set constructed in the experimental data to train the amino acid sequence generation device; The rational design module is based on the amino acid sequence generated by the upstream model, takes the amino acid sequence with specific function as the evolution target, and further transforms the generated amino acid sequence through the rational design process to further design the amino acid sequence with specific function; The filter module, based on the evolved sequences, screens according to the functions of the sequences to narrow down the range of preferred sequences for experimental verification; The scoring and ranking module screens the sequences based on the degree of binding between the sequences and the targets based on the sequences obtained by the filter module, and finally obtains the preferred sequences for experimental verification; The algorithm result verification laboratory uses experiments to obtain the production fitness and organ enrichment of the amino acid sequences stored in the target amino acid sequence library, and verifies the effectiveness and accuracy of the model in generating amino acid sequences with high production fitness and high organ enrichment.
13. An adeno-associated virus (AAV) vector capsid, characterized in that: It comprises a continuous amino acid sequence having a length of 7-mer to 18-mer and a C-terminal sequence satisfying GYSS.
14. The AAV vector capsid according to claim 13, characterized in that The continuous amino acid sequence has a length of 7-mer to 8-mer; preferably, the amino acid sequence of the oligomeric peptide is MMRGYSS or TGFGYSS.
15. The AAV vector capsid according to claim 13 or 14, characterized in that The oligomeric peptide containing the continuous amino acid sequence is a part of the AAV vector capsid protein.
16. The AAV vector capsid according to any one of claims 13 to 15, characterized in that The continuous amino acid sequence is inserted into the AAV vector capsid at a position corresponding to between 588-589 of the amino acid sequence shown in SEQ ID NO: 1 or 2.
17. An oligomeric peptide for modifying the capsid protein of an AAV vector, characterized in that: It has a length of 7-mer to 18-mer and comprises a GYSS sequence, and preferably the sequence at the C-terminus satisfies GYSS.
18. The oligomeric peptide for capsid protein modification of AAV vector according to claim 17, characterized in that: It has a length of 7-mer to 8-mer; preferably, the amino acid sequence of the oligomeric peptide is MMRGYSS or TGFGYSS.
19. The oligomeric peptide for capsid protein modification of AAV vector according to claim 17 or 18, characterized in that: It is used to be inserted into the AAV vector capsid at a position corresponding to the amino acid sequence shown in SEQ ID NO: 1 or 2 between 588-589.
20. A method for modifying an AAV vector capsid, characterized in that: An amino acid sequence having a length of 7-mer to 18-mer and containing the sequence GYSS is inserted into the position between 588-589 of the amino acid sequence shown in SEQ ID NO: 1 or 2 in the AAV vector capsid. Preferably, the inserted sequence is a continuous amino acid sequence satisfying GYSS at the C-terminus.
21. The method for modifying the AAV vector capsid according to claim 20, characterized in that: The continuous amino acid sequence is MMRGYSS or TGFGYSS.
Citation Information
Patent Citations
Protein structure prediction method based on conformational diversity sampling
CN110556161A
Gene sequence optimization method , device and equipment and medium
CN111883208A
Antibacterial peptide prediction method and device based on protein pre-training representation learning
CN112614538A
Redirection of tropism of aav capsids
CN113166208A
Method and apparatus for evolutionary data driven design of protein and other sequence defined biomolecules using machine learning
CN114651064A