Insecticidal protein variant and application thereof

By screening amino acid mutations using pre-trained protein language model and structural model, protein variants with high insecticidal activity against Fatty and Bollworm were obtained, which solved the problem of narrow insecticidal spectrum of Cry1Da and achieved the expansion of insecticidal spectrum.

CN120209103APending Publication Date: 2025-06-27HAINAN LIKEN BIOTECHNOLOGY CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411856450.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing Cry1Da insecticidal protein has a narrow insecticidal profile, which makes it limited in commercial applications, and it is difficult to control local characteristics when generating novel protein amino acid sequences using generative adversarial networks.

Method used

By pre-training the protein language model, structural model and protein residue prediction tool, the effects of amino acid mutations on protein activity were calculated, and protein variants with high insecticidal activity on Fatty and Bollworm were screened out.

Benefits of technology

Three protein variants with high insecticidal activity against both Fallia meadow and Bollworm were obtained, and were used to develop biopesticides or cultivate insect-resistant plants, expanding the insecticidal spectrum.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005191935890000061
    Figure BDA0005191935890000061
  • Figure HDA0005191935900000011
    Figure HDA0005191935900000011
Patent Text Reader

Abstract

The invention discloses an insecticidal protein variant and application thereof, and belongs to the field of genetic engineering. On the basis of active site analysis and insecticidal activity detection, a batch of insecticidal protein variants with outstanding effects are obtained through screening. The insecticidal protein variants have application value in the field of preparation of insecticides or cultivation of genetic engineering plants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses an insecticidal protein variant and its application, belonging to the field of genetic engineering. Background Art

[0002] The insecticidal proteins derived from Bacillus thuringiensis are the target proteins commonly used in the development of transgenic crops and biological pesticides. Cry1Da is a Cry-type insecticidal protein and has not been commercially applied on a large scale due to its narrow insecticidal spectrum.

[0003] The functions of protein diversity come from the combined sequences of 20 amino acid molecules. Each type of protein has a unique amino acid sequence, which endows a specific three-dimensional structure. The amino acid sequence is the main functional expression of the information stored in DNA in the form of genes.

[0004] The Generative Adversarial Network (GAN) is a deep learning model. When using the generative adversarial network to train and generate the amino acid sequences of new proteins, a large amount of real amino acid sequence data is required. However, specific categories (such as insecticidal proteins) or rare amino acid sequences may be rare in the training data, so it may cause the sequences generated by the adversarial network to deviate from the real distribution. Therefore, although the generative adversarial network can obtain overall realistic amino acid sequences, it is difficult to control the local features of the generated sequences (for example, amino acids at specific positions).

[0005] The Transformer architecture, as a deep learning model, is widely used in the field of natural language processing. Through the self-attention mechanism, this architecture enables the model to process all elements in the sequence simultaneously, greatly improving the computational efficiency. The language models of the Transformer architecture are usually pre-trained with a large amount of text data and then fine-tuned on specific tasks. Through unsupervised large-scale text generation and supervised language understanding or generation tasks, the model can learn rich language representations and thus achieve excellent performance in various downstream tasks. The present invention introduces the language model of the Transformer architecture to generate the amino acid sequences of proteins with specific functions and properties. Summary of the Invention

[0006] In order to solve the above problems, the present invention adopts the following technical solutions:

[0007] The present invention provides a protein, characterized in that the protein has any one of the following mutations based on the amino acid sequence shown in SEQ ID NO.1: (1) C424R; (2) C424A + M494A; (3) C424A + M494S.

[0008] The present invention also provides a nucleic acid molecule, characterized in that the nucleic acid molecule encodes the above-mentioned protein.

[0009] The present invention also provides a vector, characterized in that the vector contains the above-mentioned nucleic acid molecule.

[0010] The present invention also provides a recombinant cell, characterized in that the recombinant cell contains the above-mentioned nucleic acid molecule or the above-mentioned vector.

[0011] In some embodiments, the above-mentioned recombinant cell is a microbial cell or a non-renewable plant or animal cell.

[0012] The present invention also provides the use of the above-mentioned protein, or nucleic acid molecule, or vector, or recombinant cell in any one of the following:

[0013] (1) Against Spodoptera frugiperda or Helicoverpa armigera;

[0014] (2) Preparing a preparation against Spodoptera frugiperda or Helicoverpa armigera;

[0015] (3) Cultivating plants resistant to Spodoptera frugiperda and Helicoverpa armigera.

[0016] The beneficial effects of the present invention are as follows: Based on the prediction of protein active sites and mutation analysis using artificial intelligence algorithms, combined with batch activity detection and screening, the present invention has obtained 3 protein variants with high insecticidal activity against both Spodoptera frugiperda and Helicoverpa armigera, which can be used for the development of biological pesticides or the cultivation of new insect-resistant plants. Description of the Drawings

[0017] Figure 1 Prediction and mutation analysis of Cry1Da active sites. Detailed Embodiments

[0018] The following definitions and methods are provided to better define the present application and to guide those of ordinary skill in the art in the practice of the present application. Unless otherwise specified, terms are understood in accordance with the conventional usage of those of ordinary skill in the relevant art. All patent documents, academic papers, industry standards, and other publicly published materials cited herein are incorporated herein by reference in their entirety.

[0019] The following examples are used to illustrate the present invention, but not to limit the scope of the present invention. Without departing from the spirit and essence of the present invention, any modification or substitution made to the method, steps or conditions of the present invention shall fall within the scope of this application. Unless otherwise specified, the examples are carried out under conventional experimental conditions, such as the Molecular Cloning Laboratory Manual by Sambrook et al. (Sambrook J & Russell DW, Molecular cloning: a laboratory manual, 2001), or according to the conditions recommended in the manufacturer's instructions. Unless otherwise specified, the chemical reagents used in the examples are all commercially available conventional reagents, and the technical means used in the examples are conventional means well known to those skilled in the art.

[0020] Example 1 Active Site and Mutation Analysis of a Novel Protein

[0021] Cry1Da is a Cry insecticidal protein that has not been commercially applied on a large scale due to its narrow insecticidal spectrum. Chinese Patent CN106852147B discloses a point mutant protein of Cry1Da, which includes 3 amino acid mutations, namely S282V, Y316S, and I368P. The mutations of these 3 amino acids can enhance the resistance to Helicoverpa armigera while maintaining the activity of Cry1Da against Spodoptera frugiperda, thus expanding the insecticidal spectrum of the Cry1Da protein. This indicates that amino acid mutation is an effective way to expand the insecticidal spectrum of Cry proteins. However, since there are 606 amino acids in the core insecticidal region of Cry1Da, it is impossible to predict which amino acid to mutate and which amino acid each amino acid should mutate into to achieve the technical effect of expanding the insecticidal spectrum. A large screening cost is required for large-scale amino acid mutation screening.

[0022] In order to more efficiently achieve the effect of expanding the insecticidal spectrum through amino acid mutation, the present invention intends to use a pre-trained protein language model (PLMs), a structural model, and a protein residue prediction tool to calculate the impact of amino acid mutation at each different site on protein activity, select the mutation method that can bring a greater positive impact to verify the protein activity, so as to screen and obtain a novel mutant protein with better effects.

[0023] The pre-trained language model uses SaProt, the structural model uses the ProteinGAN tool, the protein residue probability prediction is implemented by the energy model ABACUS-R, and finally learning to rank (LTR) is used to obtain the mutations that can improve protein function.

[0024] Among them, the general process of constructing the pre-trained protein language model is as follows:

[0025] 1. Collect the amino acid sequences of proteins from the databases UniParc, UniprotKB, and Pfam;

[0026] 2. Using the amino acid sequences of proteins obtained in the above step, build a language model based on the Transformer architecture with the open-source machine learning library PyTorch, and pre-train the language model;

[0027] 3. Collect the amino acid sequences of all proteins and the properties of different types of proteins from the database ProThermDB;

[0028] 4. Compose the properties of the protein into control labels, and use the control labels and the amino acid sequences of the protein as a fine-tuning data set; where the control labels include the family to which it belongs, solubility, activity, etc.

[0029] 5. Set hyperparameters for model fine-tuning to obtain a fine-tuned protein language model;

[0030] 6. Amino acid sequence generation, input the control label into the protein language model, and output the amino acid sequence of a protein with specific functions and properties.

[0031] The protein residue prediction tool automatically designs all or part of the amino acid sequence of a protein according to a preset target backbone structure. Using a pre-trained deep learning neural network encoder, the three-dimensional local structural environment of a single central residue is encoded as a real-valued vector. At the same time, the encoder-decoder is pre-trained, and the decoder is used to decode this vector into the side chain type of the central residue. The input of the encoder contains the side chain type information of other residues that are spatially adjacent to the central residue. In sequence design, starting from an arbitrarily set initial sequence, this encoder / decoder is applied to different central residues, and the side chain type of the central residue is updated according to the decoded output of the local environment of the central residue in the current sequence context; by repeatedly iterating this process, an amino acid sequence is finally generated as the design result.

[0032] Among them, the input of the encoder includes the local environment of the central residue, including the backbone conformation of the central residue, and the position and orientation of the residues that are spatially adjacent to the central residue relative to the central residue, the sequence position and side chain type relative to the central residue. The adjacent residues refer to multiple residues that are closest to the central residue in distance and surround the central residue spatially in a given backbone structure. Specifically, in a given backbone structure, the 20 residues with the closest Cα atoms of other residues to the Cα atom of the central residue are sorted by distance. The decoder was trained using the Adam optimizer implemented in PyTorch. The PDB structures used for training and testing were selected from tens of thousands of X-ray structures using the PISCES server.

[0033] After the construction of the above-mentioned pre-trained language model and structural model, protein residue probability prediction and ranking learning, the active sites and variant scores of the function of Cry1Da protein are obtained (see Figure 1 ).

[0034] Example 2 Insecticidal Activity Test of Protein Variants

[0035] According to the ranking results in Example 1, the following variants are selected for verification of insecticidal activity: Cry1DaV1 (C424A), Cry1DaV2 (C424D), Cry1DaV3 (C424G), Cry1DaV4 (C424I), Cry1DaV5 (C424L), Cry1DaV6 (C424N), Cry1DaV7 (C424P), Cry1DaV8 (C424R), Cry1DaV9 (C424S), Cry1DaV10 (C424T), Cry1DaV11 (C424A + M494A), Cry1DaV12 (C424A + M494G), Cry1DaV13 (C424A + M494S), Cry1DaV14 (C424A + M494T).

[0036] Design a nucleic acid sequence encoding the above-mentioned original amino acid sequence of Cry1Da (SEQ ID NO.1), where the codons are set to the preference of Escherichia coli (K12 strain), and avoid XhoI and HindIII restriction sites to obtain a nucleic acid sequence (SEQ ID NO.2). Manually synthesize this nucleic acid molecule and clone it between the XhoI and HindIII sites of the vector pET28a expression vector to obtain a protein expression vector. Then, according to the above variants, site-directed mutagenesis is performed on the SEQ ID NO.2 sequence in the above protein expression vector to obtain protein expression vectors of all variants. Transfer the above vectors into Escherichia coli BL21 cell line and perform protein expression. The specific steps are as follows:

[0037] A single colony was inoculated into 0.5 mL of LB liquid medium and cultured at 37 °C for 4 h until the medium became turbid. Then, 100 μL of the bacterial solution was taken and IPTG (Isopropyl-β-D-thiogalactoside) was added to a final concentration of 0.8 mM. At the same time, 100 μL of the bacterial solution was taken as a negative control and cultured for another 4 h. Then, 25 μL of loading buffer was added to 100 μL of the bacterial solution to prepare the sample for electrophoresis. According to the comparison between the negative control and the result of IPTG induction, it was determined whether there was expression. For the samples with expression, the remaining 20 μL was taken and inoculated into 2 mL of LB liquid medium and cultured at 37 °C for 12 - 16 h as the seed solution. The seed solution was then inoculated into 250 mL of LB liquid medium until OD600 = 0.5 - 0.6, and then IPTG (Isopropyl-β-D-thiogalactoside) was added to a concentration of 0.8 mM and cultured for another 4 h under the same conditions. The culture solution was centrifuged at 5000 g for 10 minutes to precipitate Escherichia coli cells, and then the supernatant was discarded and the precipitate was collected. 30 mL of 20 mM Tris - 50 mM NaCl buffer was added to the precipitate and sonicated. After centrifugation, the supernatant was detected for the presence of recombinant protein.

[0038] Furthermore, the recombinant protein obtained from the above experiment was subjected to an insecticidal activity test. Specifically:

[0039] Bioassay was carried out by the surface coating method. First, about 1 mL of non-solidified artificial diet (about 0.5 g) was added to a 24-well plate and gently shaken to spread the diet evenly on the bottom of the well. After the diet solidified, protein solutions with different concentrations (10 μL / well) were added, and after addition, it was gently shaken to evenly spread the liquid medicine on the surface of the diet. It was naturally air-dried in a fume hood for 1 h. The experiment was set with 6 gradient concentrations (0.01, 0.1, 0.5, 1, 2, 5 μg / g) and a blank control (buffer). 24 newly hatched larvae (hatching time was 2 - 12 h) of Spodoptera frugiperda or Helicoverpa armigera reared artificially were inoculated for each treatment, and 3 replicates were set. They were placed in an insect rearing room at a temperature of 25 ± 2 °C, a photoperiod of 14:10 (L:D) h, and a relative humidity of 50 - 70% for culture. After 7 days, the mortality rate was investigated. A larva was considered dead if it did not move when gently touched with a brush at the tail, and a larva that did not develop to the second instar was also considered dead.

[0040] The mortality rate and corrected mortality rate were calculated according to the following formula, and the LC50 value was calculated using graphpad.

[0041]

[0042] The LC50 values of the original Cry1Da and V1, V3, V4, V6, V10, V11, V12, V13, and V14 against Spodoptera frugiperda are all less than 0.5 μg / cm 2 , and the LC50 values of other variants are all greater than 0.5 μg / cm 2 . In addition, the LC50 values of Cry1Da V8, V11, and V13 against Helicoverpa armigera are all less than 1 μg / cm 2 , and the LC50 values of other variants and the original Cry1Da protein are all greater than 1 μg / cm 2 .

[0043] Therefore, these three variants, Cry1Da V8 (C424R), Cry1Da V11 (C424A + M494A), and Cry1Da V13 (C424A + M494S), have applications in the fields of controlling Spodoptera frugiperda or Helicoverpa armigera, or preparing agents for controlling Spodoptera frugiperda or Helicoverpa armigera, or cultivating plants resistant to Spodoptera frugiperda or Helicoverpa armigera. Among them, the C424R mutation is obtained by changing the codon TGC to CGC, the C424A mutation is obtained by changing the codon TGC to GCC, the M494A mutation is obtained by changing the codon ATG to GCG, and the M494S mutation is obtained by changing the codon ATG to TCG.

[0044] Although the present invention has been described in detail above with general descriptions and specific embodiments, on the basis of the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of the present invention claimed.

Claims

1. A protein, characterized in that The protein undergoes any of the following mutations based on the amino acid sequence shown in SEQ ID NO.1: (1) C424R; (2) C424A+M494A; (3) C424A+M494S.

2. A nucleic acid molecule, characterized in that The nucleic acid molecule encodes the protein according to claim 1.

3. A carrier, characterized in that The vector comprises the nucleic acid molecule of claim 2.

4. A recombinant cell, characterized in that The recombinant cell contains the nucleic acid molecule according to claim 2 or the vector according to claim 3.

5. The recombinant cell according to claim 4, characterized in that The recombinant cell is a microbial cell or a non-renewable plant or animal cell.

6. Use of the protein according to claim 1, or the nucleic acid molecule according to claim 2, or the vector according to claim 3, or the recombinant cell according to claim 4 in any of the following: (1) Resistance to fall armyworm or cotton bollworm; (2) preparing an anti-fall armyworm or cotton bollworm preparation; (3) Cultivate plants resistant to fall armyworm and cotton bollworm.

Citation Information

Patent Citations

  • Lepidoptera active Cry1Da1 amino acid sequence variant protein

    CN106852147B