Protein ensemble prediction method adopting asymptotic search MSA
Through the asymptotic search method, clustering and sequence alignment, finding the sequence center, merging sub-MSA and inserting GAP, the problem of low efficiency of protein dynamic ensemble prediction in the prior art is solved, and efficient and accurate protein ensemble prediction is achieved.
Patent Information
- Application Number
- CN202510048792.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-13
AI Technical Summary
It is difficult for the prior art to efficiently predict protein dynamic ensembles, especially when computing resources are limited, traditional methods have problems such as high computational costs and short time scales.
Asymptotic search multi-sequence alignment (MSA) method is used to generate MSA for target protein sequences, cluster and sequence alignment, find the sequence center, repeat the above process, merge sub-MSA, insert appropriate GAP, and finally enter AlphaFold2 to generate protein ensemble.
Effectively utilize MSA coevolution information, retain the initial coevolution information and add the associated coevolution information, which improves the efficiency and accuracy of protein ensemble prediction and is suitable for scenarios with limited computing resources.
Smart Images

Figure CN119964645A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the fields of bioinformatics and computer application, and in particular to a protein ensemble prediction method using asymptotic search MSA. Background Art
[0002] Proteins are the main executors of life activities, components of cells, and have a wide range of functions. They play a vital role in almost all biological processes. However, proteins should not be viewed as static single structures, but as a conformational ensemble containing multiple accessible states, which reveal the relationship between proteins and biological processes such as ligand binding, enzyme catalysis, and signal transduction. Therefore, obtaining the conformational ensemble of proteins is crucial to clarify protein functions.
[0003] Among the current prediction methods for protein ensembles, traditional experimental methods such as X-ray crystallography and nuclear magnetic resonance (NMR) can provide static structural information of proteins, but usually cannot capture dynamic behaviors. In recent years, molecular dynamics simulation (MD) has been applied to dynamic structure prediction. This method can explore the movement of proteins under different conditions and capture their dynamic changes, but it has difficult-to-solve problems such as high computational cost and short time scale. With the rapid development of machine learning, methods that combine machine learning and MD to predict protein ensembles have also achieved some results, such as idpGAN, ProMD, etc., but these methods also have the problems faced by the above-mentioned MD methods. In today's protein structure prediction, AlphaFold2 has made significant progress in the field of static protein structure prediction, providing a solid foundation for many protein structure prediction tasks. However, the result predicted by AlphaFold2 is the final stable static structure, and does not take into account the dynamic capture of protein structure. Therefore, many methods have emerged to modify the AlphaFold2 pipeline to predict the dynamic structure of proteins, such as AF-Cluster, SPEACH_AF, AFsample2, AlphaFLOW, etc. However, among these methods, only AlphaFLOW predicts protein ensembles, while other methods only predict multiple conformations of proteins, and AlphaFLOW also has problems such as large computing resource requirements and slow speed. Therefore, how to maintain efficiency to predict protein ensembles with limited resources has always been a major challenge in biology. Summary of the invention
[0004] In order to overcome the shortcomings of the prior art and to effectively utilize MSA co-evolution information to predict the dynamic ensemble of proteins, the present invention proposes a protein ensemble prediction method using asymptotic search of MSA, wherein multiple clusters are obtained by clustering the MSA obtained by searching the current protein sequence, and the cluster center in each cluster is obtained, that is, the sequence with the highest sequence similarity score, and the corresponding MSA is searched using the sequence with the highest score, and the above operation is performed again to obtain a new MSA, and all the obtained MSAs are merged and clustered, and finally, each cluster obtained by clustering is input into AlphaFold2 to predict the protein ensemble.
[0005] The technical solution adopted by the present invention to solve its technical problem is:
[0006] A protein ensemble prediction method using asymptotic search MSA, the method comprising the following steps:
[0007] 1) Given the sequence information of the target protein;
[0008] 2) Based on the sequence information of the given target protein, use MMSeqs to generate a multiple sequence alignment MSA;
[0009] 3) Clustering is performed on the generated MSA, and the sequences whose GAP ratio is greater than the set threshold are removed during clustering;
[0010] 4) For the sub-MSAs after clustering, the BLAST sequence alignment method is used to find the sequence center;
[0011] 5) Use the sequence center in the sub-MSA to generate the corresponding MSA again using MMSeqs;
[0012] 6) For the generated new set of MSAs, repeat the process from step 2) to step 5);
[0013] 7) Merge multiple sub-MSAs obtained by the cycle into a new MSA, ensuring that each initial sequence corresponds to only one integrated MSA;
[0014] 8) Remove duplicate sequences in database searches;
[0015] 9) For sequences in MSA with different sequence lengths from the original sequence, the MAFFT multiple sequence alignment tool was used to insert appropriate GAPs;
[0016] 10) Clustering is performed using the method in step 3) to retain sequences whose GAP ratio is greater than the set threshold;
[0017] 11) Input the processed multiple MSAs into AlphaFold2 one by one to generate a protein conformation ensemble.
[0018] Further, the process of 3) is as follows:
[0019] 3.1) Use DBSCAN to cluster MSA and obtain multiple sub-MSAs. DBSCAN is a commonly used density-based clustering algorithm. It finds clusters in data by density and can identify noise. The two core conditions of DBSCAN clustering are ε (epsilon) and minPts, which represent the maximum distance of the neighborhood and the minimum number of points in the ε neighborhood;
[0020] 3.2) For MSAs that cannot be successfully clustered by the DBSCAN clustering method, the Gaussian mixture model GMM expectation maximization EM clustering is used.
[0021] Preferably, the process of 3.2) is:
[0022] 3.2.1) Calculate for each data point x i The probability of belonging to the kth cluster is:
[0023]
[0024] Among them, γ ik Represents data point x i The posterior probability of belonging to the kth cluster is used to measure the probability of each cluster k for the data point x i The probability explained by the Gaussian distribution of π k represents the prior probability of each Gaussian distribution, that is, the proportion of data points belonging to this cluster. N is the probability density function of the Gaussian distribution, μ k is the mean vector of the k-th Gaussian distribution, ∑k is the covariance matrix of the k-th Gaussian distribution, N(x i |μ k ,∑k) is x i The probability density function value on the k-th Gaussian distribution. For all x i The total probability density explained by all K Gaussian distributions;
[0025] 3.2.2) Update parameters: Using the probability calculated in step 3.2.1), update the mean, covariance and weight of each cluster:
[0026] Mean:
[0027]
[0028] Among them, μ k is the mean vector of the kth Gaussian distribution. For all data points x i The weighted sum of , where the weight is the posterior probability that it belongs to the kth cluster.
[0029] Covariance:
[0030]
[0031] Where ∑k is the covariance matrix of the kth Gaussian distribution. It is the weighted covariance sum of all data points, where the weight is the posterior probability that it belongs to the kth cluster.
[0032] Weight:
[0033]
[0034] Among them, π k is the weight of the kth Gaussian component. N is the total number of samples, used for normalization.
[0035] 3.2.3) Repeat steps 3.2.1) and 3.2.2) until the model parameters, i.e., mean, covariance, and weight, converge, i.e., the change in parameters is very small, or the preset number of iterations is reached.
[0036] Further, the process of 4) is as follows:
[0037] 4.1) Run BLASTP on each pair of sequences to align and calculate their Bitscore;
[0038] 4.2) Summarize the alignment scores of each query sequence and calculate its average alignment score with other sequences;
[0039] 4.3) The one with the highest score is considered as the sequence center.
[0040] Preferably, the process of 4.1) is:
[0041] 4.1.1) BLASTP is a tool for protein sequence alignment and is part of the BLAST tool family. It compares the input protein sequence with the protein sequences in the database to find local similarities and calculate their alignment scores;
[0042] 4.1.2) Bitscore is an important scoring indicator in BLAST comparison results, which is used to measure the similarity between two sequences;
[0043]
[0044] Among them, S ′is the standardized alignment score, which is used to measure the significance of the alignment between two sequences. S is the raw alignment score, which reflects the alignment quality of the two sequences. λ is the statistical scaling factor, which makes the alignment score have a uniform scale, and K is the statistical correction factor, which standardizes the significance of the alignment so that the alignment results in different backgrounds are comparable.
[0045] Furthermore, the process of 9) is as follows:
[0046] 9.1) First read all input A3M files and extract sequences;
[0047] 9.2) Then, use MAFFT to align these sequences with adjusted lengths. MAFFT will automatically add necessary GAPs during the alignment process based on the differences in the sequences. If there are gaps between the sequences at certain positions, MAFFT will insert GAPs at these positions so that all sequences can be aligned at these positions.
[0048] 9.3) The alignment results output by MAFFT will contain inserted GAPs, which represent that the sequence has no corresponding amino acid or nucleotide at that position. The aligned sequences will be written into the output file to ensure that all sequences are of the same length and aligned.
[0049] Preferably, in 3) and 10), the threshold is set to 25%.
[0050] The technical idea of the present invention is as follows: First, given the target protein sequence, use MMSeqs to generate MSA. Then, use the DBSCAN method to cluster each MSA to generate multiple sub-MSAs; for MSAs that fail to cluster, use Gaussian mixture model (GMM) to perform expectation maximization (EM) clustering. Next, use the BLAST sequence alignment method to select the sequence with the highest similarity from each sub-MSA, and use MMSeqs again to generate a new MSA, merge all sub-MSAs into a new MSA, remove completely identical sequences, and use MAFFT to add GAPs to sequences with inconsistent lengths. The processed MSA is clustered again and finally input into AlphaFold2 to generate a protein ensemble. The present invention provides a more diverse protein conformation ensemble.
[0051] The beneficial effects of the present invention are as follows: first, the MSA obtained by the first sequence search retains the basic co-evolution information; second, the MSA is searched by finding the sequence with the highest sequence alignment score, and combined with the original MSA to form a new MSA, thereby adding a variety of co-evolution information. In general, the initial co-evolution information is retained, and the co-evolution information associated with the initial co-evolution information is added. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 This is the overall flow chart of the method.
[0053] Figure 2 It is a protein dynamic ensemble predicted by using the protein ensemble prediction method of asymptotic search MSA for 6uof_A protein. DETAILED DESCRIPTION
[0054] The present invention will be further described below in conjunction with the accompanying drawings.
[0055] Reference Figure 1 and Figure 2 , a protein ensemble prediction method using asymptotic search MSA, the method comprising the following steps:
[0056] 1) Given the sequence information of the target protein;
[0057] 2) Based on the sequence information of the given target protein, use MMSeqs to generate a multiple sequence alignment MSA;
[0058] 3) For the generated MSA, clustering is performed, and the sequences whose GAP ratio is greater than the set threshold (taken as 25%) are removed during clustering. The process is as follows;
[0059] 3.1) Use DBSCAN to cluster MSA and obtain multiple sub-MSAs. DBSCAN is a commonly used density-based clustering algorithm. It uses density to find clusters in data and can identify noise. The two core conditions for DBSCAN clustering are ε (epsilon) and minPts, which mainly represent the maximum distance of the neighborhood and the minimum number of points in the ε neighborhood;
[0060] 3.2) For MSAs that cannot be successfully clustered by the DBSCAN clustering method, the Gaussian mixture model GMM expectation maximization EM clustering is used; the process is:
[0061] 3.2.1) Calculate for each data point x i The probability of belonging to the kth cluster is:
[0062]
[0063] Among them, γ ik Represents data point x i The posterior probability of belonging to the kth cluster is used to measure the probability of each cluster k for the data point x i The probability explained by the Gaussian distribution of π k represents the prior probability of each Gaussian distribution, that is, the proportion of data points belonging to this cluster. N is the probability density function of the Gaussian distribution, μ k is the mean vector of the k-th Gaussian distribution, ∑k is the covariance matrix of the k-th Gaussian distribution, N(xi |μ k ,∑k) is x i The probability density function value on the k-th Gaussian distribution. For all x i The total probability density explained by all K Gaussian distributions.
[0064] 3.2.2) Update parameters: Using the probability calculated in step 3.2.1), update the mean, covariance and weight of each cluster:
[0065] Mean:
[0066]
[0067] Among them, μ k is the mean vector of the kth Gaussian distribution. For all data points x i The weighted sum of , where the weight is the posterior probability that it belongs to the kth cluster.
[0068] Covariance:
[0069]
[0070] Where ∑k is the covariance matrix of the kth Gaussian distribution. It is the weighted covariance sum of all data points, where the weight is the posterior probability that it belongs to the kth cluster.
[0071] Weight:
[0072]
[0073] Among them, π k is the weight of the kth Gaussian component. N is the total number of samples, used for normalization.
[0074] 3.2.3) Repeat steps 3.2.1) and 3.2.2) until the model parameters (mean, covariance, weight) converge, that is, the change in parameters is very small, or the preset number of iterations is reached.
[0075] 4) For the sub-MSA after clustering, the BLAST sequence alignment method is used to find the sequence center; the process is as follows:
[0076] 4.1) Run BLASTP on each pair of sequences (such as sequence 1 and sequence 2, sequence 1 and sequence 3, etc.) to compare and calculate their Bitscores; the process is as follows:
[0077] 4.1.1) BLASTP is a tool for protein sequence alignment and is part of the BLAST tool family. It compares the input protein sequence with the protein sequences in the database to find local similarities and calculate their alignment scores;
[0078] 4.1.2) Bitscore is an important scoring indicator in BLAST comparison results, which is used to measure the similarity between two sequences;
[0079]
[0080] Among them, S ′ It is a normalized alignment score that measures the significance of the alignment between two sequences.
[0081] S is the raw alignment score, which reflects the alignment quality of the two sequences. λ is the statistical scaling factor, which makes the alignment scores have a uniform scale. K is the statistical correction factor, which standardizes the significance of the alignment, making the alignment results comparable in different backgrounds.
[0082] 4.2) Summarize the alignment scores of each query sequence and calculate its average alignment score with other sequences;
[0083] 4.3) The one with the highest score is considered as the sequence center;
[0084] 5) Use the sequence center in the sub-MSA to generate the corresponding MSA again using MMSeqs;
[0085] 6) For the generated new set of MSAs, repeat the process from step 2) to step 5);
[0086] 7) Merge multiple sub-MSAs obtained by the cycle into a new MSA, ensuring that each initial sequence corresponds to only one integrated MSA;
[0087] 8) Remove duplicate sequences in database searches to reduce redundancy;
[0088] 9) For sequences in MSA whose length is different from that of the original sequence, the MAFFT multiple sequence alignment tool is used to insert appropriate GAPs; the process is as follows:
[0089] 9.1) First read all input A3M files and extract sequences;
[0090] 9.2) Then, use MAFFT to align these sequences with adjusted lengths. MAFFT will automatically add necessary GAPs in the alignment process based on the differences in the sequences. For example, if there are deletions (such as insertions or deletions) between the sequences at certain positions, MAFFT will insert GAPs at these positions so that all sequences can be aligned at these positions.
[0091] 9.3) The alignment results output by MAFFT will contain inserted GAPs, which represent that the sequence has no corresponding amino acid or nucleotide at that position. The aligned sequences will be written into the output file to ensure that all sequences are of the same length and aligned;
[0092] 10) Clustering is performed using the methods of step 3.1) and step 3.2) to retain sequences with a GAP ratio greater than a set threshold (25%);
[0093] 11) Input the processed multiple MSAs into AlphaFold2 one by one to generate a protein conformation ensemble.
[0094] This implementation case uses the transcriptional regulatory protein 6uof_A with a sequence length of 119 as an implementation case. The experimental results are as follows Figure 2 A protein ensemble prediction method using asymptotic search MSA comprises the following steps:
[0095] 1) Given the sequence information of the target protein;
[0096] 2) Based on the sequence information of the given target protein, use MMSeqs to generate a multiple sequence alignment MSA;
[0097] 3) Clustering is performed on the generated MSA, and the sequences with GAP ratio greater than 25% are removed during clustering;
[0098] The process is as follows:
[0099] 3.1) Use DBSCAN to cluster MSA and obtain multiple sub-MSAs. DBSCAN is a commonly used density-based clustering algorithm that uses density to find clusters in data and can identify noise. The two core conditions for DBSCAN clustering are ε (epsilon) and
[0100] minPts, represents the maximum distance of the neighborhood and the minimum number of points in the ε neighborhood;
[0101] 3.2) For MSAs that cannot be successfully clustered by the DBSCAN clustering method, the Gaussian mixture model (GMM) expectation maximization (EM) clustering is used; the process is as follows:
[0102] 3.2.1) Calculate for each data point x i The probability of belonging to the kth cluster is:
[0103]
[0104] Among them, γ ik Represents data point x iThe posterior probability of belonging to the kth cluster is used to measure the probability of each cluster k for the data point x i The probability explained by the Gaussian distribution of π k represents the prior probability of each Gaussian distribution, that is, the proportion of data points belonging to this cluster. N is the probability density function of the Gaussian distribution, μ k is the mean vector of the k-th Gaussian distribution, ∑k is the covariance matrix of the k-th Gaussian distribution, N(x i |μ k ,∑k) is x i The probability density function value on the k-th Gaussian distribution. For all x i The total probability density explained by all K Gaussian distributions.
[0105] 3.2.2) Update parameters: Using the probability calculated in step 3.2.1), update the mean, covariance and weight of each cluster:
[0106] Mean:
[0107]
[0108] Among them, μ k is the mean vector of the kth Gaussian distribution. For all data points x i The weighted sum of , where the weight is the posterior probability that it belongs to the kth cluster.
[0109] Covariance:
[0110]
[0111] Where ∑k is the covariance matrix of the kth Gaussian distribution. It is the weighted covariance sum of all data points, where the weight is the posterior probability that it belongs to the kth cluster.
[0113] Weight:
[0114]
[0115] Among them, π k is the weight of the kth Gaussian component. N is the total number of samples, used for normalization.
[0116] 3.2.3) Repeat steps 3.2.1) and 3.2.2) until the model parameters, i.e., mean, covariance,
[0117] The weights converge, that is, the change in parameters is very small, or the preset number of iterations is reached.
[0118] 4) For the sub-MSA after clustering, the BLAST sequence alignment method is used to find the sequence center; the process is as follows:
[0119] 4.1) Run BLASTP on each pair of sequences (such as sequence 1 and sequence 2, sequence 1 and sequence 3, etc.) to compare and calculate their Bitscores; the process is as follows:
[0120] 4.1.1) BLASTP is a tool for protein sequence alignment and is part of the BLAST tool family. It compares the input protein sequence with the protein sequences in the database to find local similarities and calculate their alignment scores;
[0121] 4.1.2) Bitscore is an important scoring indicator in BLAST comparison results, which is used to measure the similarity between two sequences;
[0122]
[0123] Among them, S ′ It is a normalized alignment score that measures the significance of the alignment between two sequences.
[0124] S is the raw alignment score, which reflects the alignment quality of the two sequences. λ is the statistical scaling factor, which makes the alignment scores have a uniform scale. K is the statistical correction factor, which standardizes the significance of the alignment, making the alignment results comparable in different backgrounds.
[0125] 4.2) Summarize the alignment scores of each query sequence and calculate its average alignment score with other sequences;
[0126] 4.3) The one with the highest score is considered as the sequence center;
[0127] 5) Use the sequence center in the sub-MSA to generate the corresponding MSA again using MMSeqs;
[0128] 6) For the generated new set of MSAs, repeat the process from step 2) to step 5);
[0129] 7) Merge multiple sub-MSAs obtained by the cycle into a new MSA, ensuring that each initial sequence corresponds to only one integrated MSA;
[0130] 8) Remove duplicate sequences in database searches to reduce redundancy;
[0131] 9) For sequences in MSA whose length is different from that of the original sequence, the MAFFT multiple sequence alignment tool is used to insert appropriate GAPs; the process is as follows:
[0132] 9.1) First read all input A3M files and extract sequences;
[0133] 9.2) Then, use MAFFT to align these adjusted length sequences. MAFFT will automatically add necessary GAPs in the alignment process based on the differences in the sequences. For example, if there are deletions (such as insertions or deletions) between the sequences at certain positions, MAFFT will insert GAPs at these positions so that all sequences can be aligned at these positions;
[0134] 9.3) The alignment results output by MAFFT will contain inserted GAPs, which represent that the sequence has no corresponding amino acid or nucleotide at that position. The aligned sequences will be written into the output file to ensure that all sequences are of the same length and aligned;
[0135] 10) Clustering was performed using the methods of steps 3.1) and 3.2) to retain sequences with a GAP ratio greater than 25%;
[0136] 11) Input the processed multiple MSAs into AlphaFold2 one by one to generate a protein conformation ensemble.
[0137] The above is a result of an example of the present invention. Obviously, the present invention is not only suitable for the above embodiment, but also can be implemented with various changes without departing from the basic spirit of the present invention and without exceeding the content involved in the essential content of the present invention.
Claims
1. A protein ensemble prediction method using asymptotic search MSA, characterized in that: The method comprises the following steps: 1) Given the sequence information of the target protein; 2) Based on the sequence information of the given target protein, use MMSeqs to generate a multiple sequence alignment MSA; 3) Clustering is performed on the generated MSA, and the sequences whose GAP ratio is greater than the set threshold are removed during clustering; 4) For the sub-MSAs after clustering, the BLAST sequence alignment method is used to find the sequence center; 5) Use the sequence center in the sub-MSA to generate the corresponding MSA again using MMSeqs; 6) For the generated new set of MSAs, repeat the process from step 2) to step 5); 7) Merge multiple sub-MSAs obtained by the cycle into a new MSA, ensuring that each initial sequence corresponds to only one integrated MSA; 8) Remove duplicate sequences in database searches; 9) For sequences in MSA with different sequence lengths from the original sequence, the MAFFT multiple sequence alignment tool was used to insert appropriate GAPs; 10) Clustering is performed using the method in step 3) to retain sequences whose GAP ratio is greater than the set threshold; 11) Input the processed multiple MSAs into AlphaFold2 one by one to generate a protein conformation ensemble.
2. A protein ensemble prediction method using asymptotic search MSA as claimed in claim 1, characterized in that: The process of 3) is as follows: 3.1) Use DBSCAN to cluster MSA to obtain multiple sub-MSAs. The two core conditions for clustering judgment of DBSCAN are ε and minPts, which represent the maximum distance of the neighborhood and the minimum number of points in the ε neighborhood; 3.2) For MSAs that cannot be successfully clustered by the DBSCAN clustering method, the Gaussian mixture model GMM expectation maximization EM clustering is used.
3. A protein ensemble prediction method using asymptotic search MSA as claimed in claim 2, characterized in that: The process of 3.2) is as follows: 3.2.1) Calculate for each data point x i The probability of belonging to the kth cluster is: Among them, γ ik Represents data point x i The posterior probability of belonging to the kth cluster is used to measure the probability of each cluster k for the data point x i The probability explained by the Gaussian distribution of π k represents the prior probability of each Gaussian distribution, that is, the proportion of data points belonging to the cluster, N is the probability density function of the Gaussian distribution, μ k is the mean vector of the k-th Gaussian distribution, ∑k is the covariance matrix of the k-th Gaussian distribution, N(x i |μ k ,∑k) is x i The probability density function value on the k-th Gaussian distribution. For all x i The total probability density explained by all K Gaussian distributions. 3.2.2) Update parameters: Using the probability calculated in step 3.2.1), update the mean, covariance and weight of each cluster: Mean: Among them, μ k is the mean vector of the kth Gaussian distribution. For all data points x i The weighted sum of , where the weight is the posterior probability that it belongs to the kth cluster. Covariance: Where ∑k is the covariance matrix of the kth Gaussian distribution. It is the weighted covariance sum of all data points, where the weight is the posterior probability that it belongs to the kth cluster. Weight: Among them, π k is the weight of the kth Gaussian component. N is the total number of samples, used for normalization. 3.2.3) Repeat steps 3.2.1) and 3.2.2) until the model parameters, i.e., mean, covariance, and weight, converge, i.e., the change in parameters is very small, or the preset number of iterations is reached.
4. A protein ensemble prediction method using asymptotic search of MSA as claimed in any one of claims 1 to 3, characterized in that: The process of 4) is as follows: 4.1) Run BLASTP on each pair of sequences to align and calculate their Bitscore; 4.2) Summarize the alignment scores of each query sequence and calculate its average alignment score with other sequences; 4.3) The one with the highest score is considered as the sequence center.
5. A protein ensemble prediction method using asymptotic search MSA as claimed in claim 4, characterized in that: The process of 4.1) is as follows: 4.1.1) BLASTP is a tool for protein sequence alignment and is part of the BLAST tool family. It compares the input protein sequence with the protein sequences in the database to find local similarities and calculate their alignment scores; 4.1.2) Bitscore is an important scoring indicator in BLAST comparison results, which is used to measure the similarity between two sequences; Among them, S ′ is the standardized alignment score, which is used to measure the significance of the alignment between two sequences. S is the raw alignment score, which reflects the quality of the alignment between the two sequences. λ is the statistical scaling factor, which makes the alignment score have a uniform scale. K is the statistical correction factor, which standardizes the significance of the alignment so that the alignment results are comparable in different backgrounds.
6. A protein ensemble prediction method using asymptotic search of MSA as claimed in any one of claims 1 to 3, characterized in that: The process of 9) is as follows: 9.1) First read all input A3M files and extract sequences; 9.2) Then, use MAFFT to align these sequences with adjusted lengths. MAFFT will automatically add necessary GAPs during the alignment process based on the differences in the sequences. If there are gaps between the sequences at certain positions, MAFFT will insert GAPs at these positions so that all sequences can be aligned at these positions. 9.3) The alignment results output by MAFFT will contain inserted GAPs, which represent that the sequence has no corresponding amino acid or nucleotide at that position. The aligned sequence will be written into the output file. Make sure all sequences are of the same length and aligned.
7. A protein ensemble prediction method using asymptotic search of MSA as claimed in any one of claims 1 to 3, characterized in that: In the above 3) and 10), the threshold is set to 25%.
Citation Information
Patent Citations
Method for predicting multiple conformations and conformation transformation paths of protein
CN118692554A
Generative deep learning model-based allosteric protein conformation ensemble prediction method
CN118866118A
Protein structure prediction system
US20180260517A1
Method to construct protein structures
US6490532B1
Computer-based strategy for peptide and protein conformational ensemble enumeration and ligand affinity analysis
WO2002073193A1