A method for determining protein mutation stability and its application in peptide splicing
Through sliding window slices and amino acid word frequency statistics, the problem of protein mutation stability prediction in the prior art depends on the tertiary structure, and efficient and accurate mutation stability judgment and peptide splicing are achieved.
Patent Information
- Application Number
- CN202310930596.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-27
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-07-27
AI Technical Summary
Existing methods for predicting protein mutation stability rely on changes in the tertiary structure or physical chemical properties of proteins, resulting in large workload and low accuracy.
By using sliding window slices of amino acid sequences, the word frequency of amino acid fragments is calculated, the stability of protein mutations is judged using statistical methods, and the optimal matching junction point is selected in the polypeptide splicing.
It realizes efficient and accurate mutation stability judgment and peptide splicing that do not rely on protein structure information, improving the calculation speed and accuracy.
Smart Images

Figure CN116959568B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to a method for determining protein mutation stability and its application in polypeptide splicing. Background Art
[0002] In the field of protein mutation stability prediction, it is basically dependent on the tertiary structure of the protein or other indicators based on sequence calculations to better predict its stability.
[0003] The existing prediction methods include: DDGun, ACDC-NN-Seq, SAAFEC-Seq, MuPro, INPS-Seq and I-Mutant.
[0004] In the DDGun method, Ludovica Montanucci et al. first obtained the sequence from the ATOM field of the PDB coordinate file for each protein identified by PDBID, if its variant data was available. A multiple sequence alignment to the Uniprot database was established by the hhblits program (using default parameters). Subsequently, an evolutionary score for a single site mutation was calculated, which includes three items in total: (1) Blosum evolutionary score (sBl), which uses the Blosum62 substitution matrix to calculate the difference in substitution scores between wild-type and mutant residues. (2) Sequence-based score (sSk) is a statistical potential developed by Skolnick et al. This score is given by the difference in pairwise statistical potential calculated between wild-type and mutant residues and their nearest neighbors in the sequence (within a two-residue window on each side). (3) The hydrophobicity score (sHp), measured by the Kyte-Doulittle hydrophobicity scale, measures the hydrophobicity difference between wild-type and mutant residues. (4) The structure-based score (sSt) takes into account the structural environment of the variant, captured by the pairwise statistical potential of Bastolla-Vendruscolo. The score is calculated for the wild-type amino acid and its structural neighborhood (defined as a region with a radius of centered on the mutation site). The scores for each single-site substitution were finally combined, and a linear combination was selected, with weights chosen based on how closely each score correlated with the known ΔΔG of single-site variants in the high-quality dataset from Yang and colleagues.
[0005] In the ACDC-NN-Seq method, Corrado Pancotti et al. pre-trained the model using an artificial dataset called IvankovDDGun and then applied transfer learning to two experimental datasets, S2648 and Varibench. The IvankovDDGun dataset consists of 10,000 protein sequences with known changes in stability after mutation. The S2648 dataset contains 2,648 protein sequences with experimental changes in stability, while the Varibench dataset contains 1,000 protein sequences with known changes in stability upon mutation. The authors evaluated the model's performance on the experimental datasets using a 5-fold cross-validation method.
[0006] In the SAAFEC-Seq method, Gen Li et al. used the following scores to predict the sequence: (1) Pseudo position specific score matrix (PsePSSM). (2) Neighbor mutation conservation score, considering a stretch of residues, XXXCXXX, where C is the mutation site and X refers to the adjacent amino acid. We selected the rows and neighbors belonging to the mutation site from the PSSM to obtain 20×7 conservation score features. (3) Sequence proximity features, five amino acids were selected from the left and right sides of the mutation site as sequence information. Each label may have 20 possibilities, representing 20 different amino acids. (4) Physicochemical property features, including: net volume, net hydrophobicity, mutation type, net flexibility, chemical properties, size, polarity, hydrogen bonding, and label hydrophobicity. Finally, to predict ΔΔG for a given mutation, a regression model was developed by using knowledge-based features representing evolutionary information and the physicochemical environment around the mutation site. To build a reliable and robust model, 100 five-fold cross validations were performed. The selection of training and test sets was randomly repeated 100 times, and the average PCC and MSE were considered. The model was trained on 80% of the 2,648 mutations in the assembled dataset and tested on the remaining 20%.
[0007] In the MuPro method, Jianlin Cheng et al. use SVM to nonlinearly map the input vector to the feature space and use linear methods in the feature space for regression or classification, thereby providing nonlinear function approximation. The three parameters for each task (penalty ratio, γ, and C for classification; tube width, γ, and C for regression) are optimized on the training data. For each cross-validation fold, we use the LOOCV (leave-one-out crossvalidation) procedure to optimize these parameters. In the LOOCV process, for a training dataset with N data points, one data point is retained in each round and the model is trained on the remaining N-1 data points. The model is then tested on the retained data points. This process is repeated N times until all data points have been tested once, and then the overall accuracy of the training dataset is calculated.
[0008] In the INPS-Seq method, P Fariselli et al. used a radial basis function kernel based on an SVM approach. Specifically, INPS uses substitution scores derived from the BLOSUM62 matrix, differences in alignment scores between the native and variant sequences, hydrophobicity, evolutionary information, and other factors to predict stability.
[0009] In the I-Mutant method, Capriotti E et al. used a radial basis function kernel based on the SVM method and input 42 features, including temperature, Ph, 20 mutation feature codes and 20 spatial residue environment feature codes (if there is protein structure), or the nearest sequence neighbor (if there is only protein sequence).
[0010] Most of the aforementioned prediction methods rely on the protein's tertiary structure or extract changes in indices caused by differences in physicochemical properties such as volume, hydrophobicity, and polarity before and after mutation to determine the protein's mutational stability. Extracting these indices is labor-intensive and difficult, and has low accuracy. Therefore, a statistically simple, highly accurate method for determining protein mutational stability and its application to peptide splicing are urgently needed. Summary of the Invention
[0011] In view of this, the present invention provides a method for determining protein mutation stability with simple statistics and high accuracy, and its application in polypeptide splicing.
[0012] The present invention provides a method for determining protein mutation stability, comprising the following steps:
[0013] S1, slice the amino acid sequence in the database through a sliding window to obtain amino acid segments with a step size of 1;
[0014] S2. Statistically sum the amino acid fragments in step S1 and calculate the word frequency of each amino acid fragment to prepare a comparison database;
[0015] S3. Preprocess the amino acid sequence for which mutation stability calculation is required. After determining the mutation site based on the mutation information, extract a local amino acid fragment for mutation. Slice the two amino acid fragments before and after the mutation using a sliding window of length K to obtain k-mers. Use the comparison database from step S2 to perform statistics on all k-mers, and determine whether the protein is stable based on the difference in word frequency before and after the mutation.
[0016] Preferably, the database is a 530 million non-redundant amino acid sequence database.
[0017] Preferably, before slicing in step S1, amino acid sequences with a length of less than 50 bp are filtered out.
[0018] Preferably, before slicing in step S1, uppercase and lowercase characters other than the twenty amino acid abbreviations are filtered out.
[0019] Preferably, step S2 adopts a dictionary nested hierarchical file storage method, and the name of a single dictionary is named according to the first three or four characters of the amino acid fragment; step S2 creates a double-layer nested dictionary for character statistics, the key of the first layer dictionary is named using the first three characters of the amino acid fragment, and the value of the first layer dictionary contains all amino acid slices with the same first three characters; the second layer dictionary is used to count the word frequency information of all amino acid fragments with the same first three characters; the text document is named and saved according to the key of the first layer dictionary, and all values corresponding to the key of the first layer dictionary are saved in each text document.
[0020] Preferably, when pre-processing the amino acid sequence for which mutation stability calculation is required in step S3, only the amino acid sequence near the mutation site is considered.
[0021] Preferably, the amino acid range selection near the mutation site is determined based on the sliding window length k, where the sliding window length is k, starting from the mutation site, taking (k-1) amino acid sequences forward and (k-1) amino acid sequences backward.
[0022] Preferably, the sliding window length of step S1 is 3-8.
[0023] The present invention also provides an application of a method for determining protein mutation stability in polypeptide splicing, which further comprises the following steps:
[0024] S4, extracting the beginning and end of the peptide segments to be spliced, and using the word frequency statistics method in step S3, by comparing the word frequencies of the amino acid fragments, selecting the two adjacent peptide segments with the most combinations as the optimal peptide segment combination, and determining the optimal matching connection point;
[0025] S5. Select a peptide from the set of peptides to be matched as the first peptide segment of the complete polypeptide sequence. Select an optimal peptide segment from the remaining peptide segments to be matched according to the optimal matching connection point determined in step S4 and splice it with the first peptide segment. For the next peptide segment to be spliced, similarly select the optimal peptide segment and splice it with the spliced polypeptide sequence until the last peptide segment to be matched is spliced to obtain a complete polypeptide sequence.
[0026] S6. The peptide segments to be matched that have not been fully sequenced are sequentially used as the first peptide segment of the complete polypeptide sequence, and then the splicing process of step S5 is performed to obtain a corresponding number of complete polypeptide sequences.
[0027] Beneficial effects of the present invention:
[0028] 1. It does not rely on the secondary structure, tertiary structure and changes in the physicochemical properties of proteins. It only needs to use the amino acid sequence for calculation and statistics. The statistics are simple, accurate and interpretable.
[0029] 2. Based on the information of the mutation site, only the amino acid sequence near the mutation site is considered to improve the calculation and statistical speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is an example graph of word frequencies for saving amino acid fragments.
[0031] Figure 2 This is an example diagram of a comparison database created based on a non-redundant amino acid sequence database.
[0032] Figure 3 Schematic diagram of the polypeptide sequence splicing method.
[0033] Figure 4 This is a comparison chart of the statistical method in Example 1 and its confusion matrix results based on sequence prediction.
[0034] Figure 5 Schematic diagram of protein expression identification by Western blot. DETAILED DESCRIPTION
[0035] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0036] Example 1
[0037] This embodiment provides a method for calculating protein mutation stability based on statistical methods, comprising the following steps:
[0038] S1. A database containing 530 million non-redundant amino acid sequences was used. This database was obtained from the National Center for Biotechnology Information (NCBI) BLAST database, a widely used database containing amino acid sequences from multiple species. These sequences were obtained from a variety of sources, including known protein databases, genome projects, and protein research. To reduce redundant information and improve alignment efficiency and accuracy, the database adopted a non-redundant strategy, retaining only one copy of similar or identical amino acid sequences. This database contains some non-standard amino acid abbreviations, including six abbreviations in addition to the 20 standard amino acid abbreviations, as well as some lowercase abbreviations to indicate modified positions within the amino acid sequence. During preprocessing of the raw database, this portion was filtered using a code. Sequences less than 50 base pairs were also filtered out, as these sequences were peptide sequences rather than amino acid sequences. While ensuring a sufficient data volume, the raw data was cleaned and filtered appropriately. A sliding window was used to segment the remaining amino acid sequences. The sliding window length was 8, with a step size of 1, meaning that the window was slid from beginning to end in steps of 1 character, resulting in 8-character segments.
[0039] S2. For each amino acid fragment obtained, classify them according to the first three characters. By creating a nested dictionary, the key of the first layer dictionary is the first three characters of the amino acid fragment, and the corresponding value is all fragments matching the prefix. For each prefix, the word frequency of the corresponding amino acid fragment needs to be counted. To this end, the amino acid fragment is used as the key in the second layer dictionary, and the corresponding value is the number of times the amino acid fragment appears, such as Figure 1 When saving nested dictionaries, the key of the first-level dictionary is the file name, and the value is the content of the corresponding text document, as shown in the following example. Figure 2 As shown, a comparison database was obtained.
[0040] S3. When preprocessing the amino acid sequences in the test set that require mutation stability calculations, the mutation site of the amino acid sequence is locally segmented for calculation. Only local information is considered, and amino acid information at other positions is ignored, thereby reducing calculation time. The length of the segmented amino acid fragment is based on the length already in the comparison database. For example, when using a comparison database of length 8, the length of the test segmented fragment is 15, with the mutated amino acid in the middle. The method in S2 is used to further segment the amino acid fragments before and after the mutation in the test set. For example, as shown in Table 1, a pair of amino acid fragments of length 15 in the test set are slid using a window of length 8 and step size 1, resulting in two groups of eight 8-mers. These two groups of eight 8-mers are then used to search the comparison database for the corresponding word frequencies. The word frequencies are then summed to obtain the number of occurrences of the amino acid at the mutation site before and after the mutation. If the number of occurrences increases after the mutation, it is considered stable; conversely, if the number of occurrences decreases, it is considered unstable.
[0041] Table 1 Statistics of the number of amino acid fragments near the mutation site before and after mutation
[0042]
[0043] To verify the performance of this embodiment under the condition of unbalanced data sets, a confusion matrix was made for comparison between the method in Example 1 and other similar methods. Figure 3 As shown, this embodiment uses the S669 test set for performance testing, with an accuracy of 74.95%; the I-Mutant method has an accuracy of 74.55%; the MuPro method has an accuracy of 73.69%; the SAAFEC-Seq method has an accuracy of 73.50%; the DDGun method has an accuracy of 71.70%; the ACDC-NN-Seq method has an accuracy of 74.44%; and the INPS-Seq method has an accuracy of 73.84%. Considering the prediction accuracy of whether the mutation is stable or unstable, Example 1 has the highest accuracy. No more protein information is required for calculation and the calculation speed is fast. It only takes a few minutes to simulate an amino acid sequence. Traditional MD simulation of an amino acid sequence takes about 24 hours.
[0044] Example 2
[0045] This example provides, based on Example 1, an application of a method for determining protein mutation stability in polypeptide splicing, further comprising the following steps:
[0046] S4, extracting the beginning and end of the peptide segments to be spliced, and using the word frequency statistics method in step S3, by comparing the word frequencies of the amino acid fragments, selecting the two adjacent peptide segments with the most combinations as the optimal peptide segment combination, and determining the optimal matching connection point;
[0047] S5. Select a peptide from the set of peptides to be matched as the first peptide segment of the complete polypeptide sequence. Select an optimal peptide segment from the remaining peptide segments to be matched according to the optimal matching connection point determined in step S4 and splice it with the first peptide segment. For the next peptide segment to be spliced, similarly select the optimal peptide segment and splice it with the spliced polypeptide sequence until the last peptide segment to be matched is spliced to obtain a complete polypeptide sequence.
[0048] S6. The peptide segments to be matched that have not been fully sequenced are sequentially used as the first peptide segment of the complete polypeptide sequence, and then the splicing process of step S5 is performed to obtain a corresponding number of complete polypeptide sequences.
[0049] In summary, Example 2 is to gradually construct a complete polypeptide sequence by selecting the optimal peptide segment and splicing it with the spliced polypeptide sequence. The remaining peptide segments are placed as the first of the sequence and the process is repeated. Figure 4 As shown, a is taken as the first peptide segment, and the remaining bp peptide segments are spliced in sequence according to the optimal matching points to obtain a complete polypeptide sequence.
[0050] In this embodiment, regarding the selection of the optimal peptide segment, the length of the sliding window is first defined according to the length of the fragment in the comparison database. For example, if the comparison database uses a fragment of length 6, it is necessary to intercept the last 5 characters of the previous segment and the first 5 characters of the next segment, and use a sliding window of length 6 and step size 1 for sliding slicing. The obtained 5 fragments are used to compare the number of queries in the database, and then the 5 times are added to obtain the total number of times. This operation is performed on all the fragments to be matched in S2 to find the optimal peptide segment to be matched.
[0051] In order to verify the splicing effect of the polypeptide, 16 peptide segments to be spliced are provided, namely: LTDDQVDEIIR, IDLGAVNNVDK, EYDESGPSIVHR, ELPDGQVITIGNER, ADVQVFGNPGAK, EDIPEIGEK, ATNPLDSNPGTIR, CGNIQGITKPAIR, HDIIEVDPDTK, ATLVSEVR, FEDIPLVK, STSGDTHLGGEDFDNR, VGNSTAIQELFK, LDSDLEGHPTPR, QPNNEGNLLNAIAK, and AADGGLDIPHSTK.
[0052] Select two complete polypeptide sequences op1 spliced according to Example 2
[0053] (AGDDAPRARRLRWPKRLDGQGRPRVWLGRAGDDAPRAHGAGVGERPGNLKSAKVLPV PQKGILGLSKLLGDKEAPLNPKSDRDLLGPNNQYLPKGWVDYPIISVYWKTTANIEDRR PIIVYWKTTLGDLIHRGMCCSRVKVYIAGKIEELEEELEAERFSVVPSPKHVHNGRTKI QLLEEDLERYRIPEYAKACFRYPPAKCSMANAVRCTGPATCKCGSGKWWGHKLDTK) and op2
[0054] (TTANIEDRRRLRWPKRLDGQGRPRVWLGRAGDDAPRARAGDDAPRAHGAGVGERPGN LKSAKVLPVPQKGILGLSKLLGDKSDRDLLGPNNQYLPKEAPLNPKGWVDYPIISVYWK PIIVYWKGMCCSRTTLGDLIHRVKVYIAGKIEELEEELEAERHVHNGRTKFSVVPSPKIQLLEEDLERYRIPECSMANAVRYAKACFRCTGPATCKCGSGKYPPAKWWGHKLDTK);
[0055] Select three randomly spliced polypeptide sequences control1
[0056] (AGDDAPRHVHNGRTKCSMANAVRAHGAGVGERIQLLEEDLERSDRDLLGPNNQYLPKEAPLNPKVLPVPQKFSVVPSPKPGNLKSAKYPPAKCTGPATCKCGSGKYAKACFRGILGLSKLLGDKPIISVYWKRLRWPKVKVYIAGKAGDDAPRARRLDGQGRPRVWLGRTTANIEDRRGMCCSRGWVDYWWGHKLDTKIEELEEELEAERPIIVYWKTTLGDLIHRYRIPE).
[0057] Control2
[0058] (PIISVYWKRLRWPKAGDDAPRARAGDDAPRRLDGQGRPRVWLGRVLPVPQKTTANIEDRRTTLGDLIHREAPLNPKSDRDLLGPNNQYLPKGILGLSKLLGDKAHGAGVGERIEEL EEELEAERPGNLKSAKVKVYIAGKFSVVPSPKPIIVYWKIQLLEEDLERYRIPEGWVDYYPPAKYAKACFRHVHNGRTKCTGPATCKCGSGKGMCCSRCSMANAVRWWGHKLDTK).
[0059] Control3
[0060] (PIIVYWKRLRWPKAGDDAPRARAGDDAPRRLDGQGRPRVWLGRVLPVPQKTTANIEDRTTLGDLIHREAPLNPKSDRDLLGPNNQYLPKGILGLSKLLGDKAHGAGVGERIEELE EELEAERPGNLKSAKVKVYIAGKFSVVPSPKPIISVYWKIQLLEEDLERYRIPEGWVDYYPPAKYAKACFRHVHNGRTKCTGPATCKCGSGKGMCCSRCSMANAVRWWGHKLDTK).
[0061] Marker is used to locate the target protein. Proteins with different molecular weights have different band positions. Western blot was performed to identify protein expression of op1, op2, control1, control2, control3 and marker. The results are as follows: Figure 5 As shown, the protein bands of op1 and op2 match those of the marker and can be secreted from the organism to the outside of the organism. The protein expression of control1, control2, and control3 is poor. They are directly degraded after production in the microorganism and cannot be released outside the microorganism for normal expression.
[0062] The above is a specific implementation of the present invention, and its description is relatively specific and detailed, but it should not be understood as limiting the scope of the present invention. For those skilled in the art, various modifications and improvements can be made without departing from the concept of the present invention, and these obvious alternative forms are all within the scope of protection of the present invention.
Claims
1. A method for determining protein mutation stability, characterized in that: The following steps are involved: S1, slice the amino acid sequence in the database through a sliding window to obtain amino acid segments with a step size of 1; S2. Statistically sum the amino acid fragments in step S1 and calculate the word frequency of each amino acid fragment to prepare a comparison database; S3. Preprocess the amino acid sequence for which mutation stability calculation is required. After determining the mutation site based on the mutation information, extract a local amino acid fragment for mutation. Slice the two amino acid fragments before and after the mutation using a sliding window of length K to obtain k-mers. Use the comparison database from step S2 to perform statistics on all k-mers, and determine whether the protein is stable based on the difference in word frequency before and after the mutation.
2. The method for determining protein mutation stability according to claim 1, wherein: The database is a database of 530 million non-redundant amino acid sequences.
3. The method for determining protein mutation stability according to claim 2, wherein: Before slicing in step S1, amino acid sequences with a length of less than 50 bp are filtered out.
4. The method for determining protein mutation stability according to claim 3, wherein: Before slicing in step S1, uppercase and lowercase characters other than the abbreviations of the twenty amino acids are filtered out.
5. The method for determining protein mutation stability according to claim 1, wherein: When step S3 pre-processes the amino acid sequence that needs to be calculated for mutation stability, only the amino acid sequence near the mutation site is considered.
6. The method for determining protein mutation stability according to claim 1, wherein: The selection of the amino acid range near the mutation site is determined by the length of the sliding window. When the sliding window length is k, starting from the mutation site, (k-1) amino acid sequences are taken forward and (k-1) amino acid sequences are taken backward.
7. The method for determining protein mutation stability according to claim 1, wherein: The sliding window length of step S1 is 3-8.
8. An application of the method for determining protein mutation stability according to any one of claims 1 to 7 in polypeptide splicing, characterized in that: The following steps are also included: S4, extracting the beginning and end of the peptide segments to be spliced, and using the word frequency statistics method in step S3, by comparing the word frequencies of the amino acid fragments, selecting the two adjacent peptide segments with the most combinations as the optimal peptide segment combination, and determining the optimal matching connection point; S5. Select a peptide from the set of peptides to be matched as the first peptide segment of the complete polypeptide sequence. Select an optimal peptide segment from the remaining peptide segments to be matched according to the optimal matching connection point determined in step S4 and splice it with the first peptide segment. For the next peptide segment to be spliced, similarly select the optimal peptide segment and splice it with the spliced polypeptide sequence until the last peptide segment to be matched is spliced to obtain a complete polypeptide sequence. S6. The peptide segments to be matched that have not been fully sequenced are sequentially used as the first peptide segment of the complete polypeptide sequence, and then the splicing process of step S5 is performed to obtain a corresponding number of complete polypeptide sequences.
Citation Information
Patent Citations
Method for the determination of sequence variants of polypeptides
CN102762987A
Method and system for predicting amino acid mutation
CN106650314A