Method for phage classification based on Markov model
Through the phage classification method based on the Markov model, the problem of difficulty in classifying short fragments and low-homology genomes in the existing technology is solved, and high-accuracy automated classification is achieved.
Patent Information
- Application Number
- CN202211639837.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-20
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-12-20
AI Technical Summary
Existing phage classification methods have difficulty in effectively classifying short fragments and low-homology genomes, and the accuracy of automated classification is insufficient.
A phage classification method based on the Markov model was adopted. By constructing a protein library and calculating the peptide state transition probability, the Markov model was used to evaluate the likelihood value of the unclassified genome, and the credibility was determined by combining Gaussian distribution fitting.
Accurate classification of short genomic fragments and low-homology genomes was achieved, with a prediction accuracy of up to 98%, improving the automated credibility of classification.
Smart Images

Figure CN116013417B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virome taxonomy, and in particular to a method for phage classification based on a Markov model. Background Art
[0002] Viruses are the largest untapped reservoir of genetic diversity on Earth. With the widespread application of metagenomic sequencing, new viral genome sequences are accumulating rapidly. Bacteriophages play an important role in balancing global ecosystems by regulating the abundance of bacteria in the natural environment. Bacteriophages are also closely related to human health. Changes in the abundance and composition of bacteriophages have been found to be associated with ulcerative colitis, Crohn's disease, and diabetes. After identifying and assembling viral genome sequences from metavirome data, taxonomic classification of sequences is the foundation of virome research. However, the newly identified viruses contain a large number of novel viral sequences and short genome fragments ranging from a few hundred bases to a few thousand bases in length. These genome sequences pose a challenge to the classification of viromes.
[0003] Currently, the most common method is the Blast-based search method. First, all proteins encoded by a genome are annotated using Prodigal. Then, the Blast method is used to search for the most similar known protein for each protein in the genome from a library of known sequences. Finally, a vote is performed based on the family or genus to which each best match belongs to determine the family or genus of the queried genome. This method has the advantages of low false positives and high accuracy; however, its disadvantage is that it is difficult to predict sequences from short fragments and distantly homologous genomes.
[0004] Another representative classification method is vConTACT, which constructs a network based on gene sharing between genomes, and then clusters the phage genomes into virus clusters according to the network to classify them. Its advantage is that it can automatically and reliably give genus-level classifications and can be applied to large metagenomic data sets; its disadvantage is that smaller genomes or genome fragments are difficult to classify, and there is a large amount of virus sequence space that cannot be classified by it.
[0005] Therefore, the present application designs a method for phage classification based on the Markov model. Summary of the Invention
[0006] The present invention provides a method for phage classification based on a Markov model, the purpose of which is to solve the above-mentioned problems existing in the background technology.
[0007] To achieve the above objectives, the present invention provides a novel method for automated phage classification, which is suitable for the taxonomic classification of genome fragments assembled from metagenomic or macrovirome data.
[0008] An embodiment of the present invention provides a method for phage classification based on a Markov model, comprising the following steps:
[0009] S1. Use Prodigal to translate NCBI phage genomes into protein sequences and establish a protein library T. Construct a protein database for each taxonomic unit at a certain taxonomic level (order, family, or genus, etc.). Taking the family level as an example, integrate the genomes in the existing database into a separate genome database by family. Then, use Prodigal to perform protein annotation and convert it into a protein library T for each family.
[0010] S2. Take the protein sequence of the protein library T and calculate the next adjacent amino acid (N) in the peptide state of a peptide segment with a length of k. xk+1 ). N x1 ...N xk For a peptide segment of length x1...x k The number of proteins in the protein library T, N x1 ...N x+1k For a peptide segment of length x1...x k+1 The number of proteins in the protein library T is used to obtain the state transition probability matrix of the k-order Markov model, and α is an adjustable pseudo count, as shown in formula (1):
[0011]
[0012] The unclassified genome was translated into protein using Prodigal to obtain the protein sequence V = y1...y N , calculate the log-likelihood value LL between the protein sequence of the unclassified genome and the protein library T according to formula (1), i+k is the peptide segment of length k when the starting point is i, y i+k-1 is a peptide segment of length k-1 when the starting point is i. Substituting these two peptide segments into the state transition probability matrix P obtained in formula (1) T Then, all the probability values obtained for the protein sequence V of the unclassified genome are summed up and the average is taken to obtain the log-likelihood value LL, as shown in formula (2):
[0013]
[0014] Fitting the LL to the mean and variance of a Gaussian distribution to obtain the Gaussian distribution and P value of the LL of the Markov model;
[0015] S3. The genome corresponding to the Markov model with the highest LL score is the predicted phage genome. To assess the confidence of the prediction, using the family level as an example, for a Markov model constructed for a particular family, the log-likelihood (LL) of the model is calculated for viral genomes from other families. These LL values are then fitted to the mean and variance of a Gaussian distribution to obtain a Gaussian distribution for the model's log-likelihood (LL). In application, after obtaining the LL value for unclassified genomes, a P value is calculated based on the model's Gaussian distribution. The P value can then be used to determine the confidence threshold.
[0016] Furthermore, when the P value is <0.01, the prediction accuracy is as high as 98%.
[0017] The above solution of the present invention has the following beneficial effects:
[0018] 1. After classifying the genomes in the virus library according to taxonomy, the present invention models all protein sequences in each class to obtain a model of each class, which can better evaluate the overall similarity between the viral genome to be classified and all genomes of each class of viruses.
[0019] 2. The present invention enables more accurate taxonomic classification assignment of shorter genomic fragments (e.g., several thousand base pairs in length);
[0020] 3. The present invention can predict and assign accurate taxonomic classifications to genomic segments that have low homology with known viral genomes. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 It is the coverage and accuracy of the test set under different thresholds of the embodiment of the present invention;
[0023] Figure 2 The embodiment of the present invention is for comparing genomic sequences of different lengths;
[0024] Figure 3 is a comparison of the predictive capabilities of the embodiments of the present invention for novel viruses;
[0025] Figure 4 4 is a flowchart of a method for phage classification based on a Markov model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0027] Unless otherwise defined, all technical terms used hereinafter have the same meanings as those generally understood by those skilled in the art. The technical terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the scope of protection of the present invention.
[0028] Unless otherwise specified, various raw materials, reagents, instruments and equipment used in the present invention can be purchased from the market or prepared by existing methods.
[0029] The present invention aims at solving the existing problems and provides a phage classification method based on a Markov model.
[0030] An embodiment of the present invention provides a method for phage classification based on a Markov model, comprising the following steps:
[0031] S1. Use Prodigal to translate NCBI phage genomes into protein sequences and establish a protein library T. Construct a protein database for each taxonomic unit at a certain taxonomic level (order, family, or genus, etc.). Taking the family level as an example, integrate the genomes in the existing database into a separate genome database by family. Then, use Prodigal to perform protein annotation and convert it into a protein library T for each family.
[0032] S2. Take the protein sequence of the protein library T and calculate the next adjacent amino acid (N) in the peptide state of a peptide segment with a length of k. xk+1 ). N x1 ...N xk For a peptide segment of length x1...x k The number of proteins in the protein library T, N x1 ...N x(k+1) For a peptide segment of length x1...x k+1 The number of proteins in the protein library T is used to obtain the state transition probability matrix of the k-order Markov model, and α is an adjustable pseudo count, as shown in formula (1):
[0033]
[0034] The unclassified genome was translated into protein using Prodigal to obtain the protein sequence V = y1...y N , calculate the log-likelihood value LL between the protein sequence of the unclassified genome and the protein library T according to formula (1), i+k is the peptide segment of length k when the starting point is i, y i+k-1is a peptide segment of length k-1 when the starting point is i. Substituting these two peptide segments into the state transition probability matrix P obtained in formula (1) T Then, all the probability values obtained for the protein sequence V of the unclassified genome are summed up and the average is taken to obtain the log-likelihood value LL, as shown in formula (2):
[0035]
[0036] Fitting the LL to the mean and variance of a Gaussian distribution to obtain the Gaussian distribution and P value of the LL of the Markov model;
[0037] S3. The genome corresponding to the Markov model with the highest LL score is the predicted phage genome. To assess the confidence of the prediction, using the family level as an example, for a Markov model constructed for a particular family, the log-likelihood (LL) of the model is calculated for viral genomes from other families. These LL values are then fitted to the mean and variance of a Gaussian distribution to obtain a Gaussian distribution for the model's log-likelihood (LL). In application, after obtaining the LL value for unclassified genomes, a P value is calculated based on the model's Gaussian distribution. The P value can then be used to determine the confidence threshold.
[0038] Furthermore, when the P value is <0.01, the prediction accuracy is as high as 98%.
[0039] Example
[0040] S1. Build a database using 4357 phage viruses from NCBI's RefSeq as reference genomes. Based on the classification information provided by RefSeq, there are 46 phage families. After translating each genome into protein sequences using Prodigal, a protein library T of 46 families was obtained by classifying them according to family.
[0041] S2. Calculate the state transition probability matrix of the k-order Markov model. Calculate the state transition probability matrix for the protein library of each family obtained in the previous step according to formula (1). 46 Markov models are obtained. The log-likelihood value LL is calculated for each model and other viral proteins, and a Gaussian distribution is fitted for credibility assessment.
[0042] S3. Construct a test data set and translate 24,000 phage genomes with family tag information downloaded from Genbank into proteins using Prodigal.
[0043] S4. Predict the classification of the test data set. For each protein in the 24,000 phage genomes, the Markov model is used to calculate the likelihood value LL and P-value. The family of the model corresponding to the highest likelihood value LL is selected as the predicted family.
[0044] S5. Compared with the actual classification, the Accuracy (A) and recall (R) curves are used to compare the coverage and accuracy of the test set at different thresholds. In the overall results of the test set, the Markov Model and Blast are comparable, with the area under the AR curve (AR-AUC) being 0.98.
[0045] from Figure 1 It can be seen that the Markov Model is better at predicting short genomic fragments.
[0046] Figure 2 For genomic sequences of different lengths, the area under the curve (AR-AUC) of the Markov model is higher than that of the benchmark method blast.
[0047] Figure 3 The AR-AUC under different AAI similarity levels between the test set and the known database was compared. It is obvious that the Markov Model has better predictive ability for novel viruses than the benchmark method Blast.
[0048] In this example, a P-value of 0.01 is selected as the threshold for credible results. Predictions with a P-value lower than 0.01 are considered credible. The performance indicators are shown in Table 1.
[0049] Table 1
[0050]
[0051] The above solution of the present invention has the following beneficial effects:
[0052] 1. The present invention enables more accurate taxonomic classification assignment of shorter genomic fragments (e.g., several thousand base pairs in length);
[0053] 2. The present invention can predict and assign accurate taxonomic classifications to genomic fragments that have low homology with known viral genomes.
[0054] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for phage classification based on a Markov model, characterized in that: The steps include: S1. Use Prodigal to translate the NCBI phage genome into protein sequences and create protein library T; S2. Take the protein sequence of the protein library T and calculate the next adjacent amino acid (N) in the peptide state of a peptide segment with a length of k. xk+1 ) conditional probability; N x1 ...N xk For a peptide segment of length x1...x k The number of proteins in the protein library T, N x1 ...N x(k+1) For a peptide segment of length x1...x k+1 The number of proteins in the protein library T is used to obtain the state transition probability matrix of the k-order Markov model, and α is an adjustable pseudo count, as shown in formula (1): The unclassified genome was translated into protein using Prodigal to obtain the protein sequence V = y1...y N , calculate the log-likelihood value LL between the protein sequence of the unclassified genome and the protein library T according to formula (1), i+k is the peptide segment of length k when the starting point is i, y i+k-1 is a peptide segment of length k-1 when the starting point is i. Substituting these two peptide segments into the state transition probability matrix P obtained in formula (1) T Then, all the probability values obtained for the protein sequence V of the unclassified genome are summed up and the average is taken to obtain the log-likelihood value LL, as shown in formula (2): Fitting the LL to the mean and variance of a Gaussian distribution to obtain the Gaussian distribution and P value of the LL of the Markov model; S3. The genome of the Markov model corresponding to the highest LL score is the predicted phage genome.
2. The method for phage classification based on the Markov model according to claim 1, characterized in that: When the P value is < 0.01, the prediction accuracy is as high as 98%.
Citation Information
Patent Citations
Phage host prediction method, apparatus and device, and storage medium
CN113658633A