An analysis method for predicting bacteriophage hosts with high throughput based on second-generation sequencing technology
By performing quality control and assembly on second-generation sequencing data, combining multiple methods to predict phage hosts, and evaluating them through purity and consistency indicators, the problems of efficient utilization of second-generation sequencing data and accurate assessment of phage hosts were solved, achieving efficient and accurate analysis results.
Patent Information
- Application Number
- CN202211393619.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-08
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2042-11-08
AI Technical Summary
Existing technologies struggle to efficiently utilize next-generation sequencing data to simultaneously assemble the genomes of bacteria and viruses and accurately assess the interaction between bacteriophages and their hosts, resulting in high research costs and inconsistent analytical results.
By performing quality control, filtering, splicing, binning, and redundancy removal on second-generation sequencing data, bacterial and viral genomes are identified. Species annotation is performed using hidden Markov models and supervised machine learning. Multiple methods are combined to predict bacteriophage hosts, and the accuracy of the prediction results is evaluated through purity and consistency indicators.
It enables efficient use of a single set of sequencing data for comprehensive analysis, reduces research costs, and improves the accuracy and consistency of analysis results. It is suitable for non-bioinformatics researchers to independently perform high-throughput sequencing data analysis.
Smart Images

Figure CN115662516B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of second-generation sequencing technology and the field of phage host prediction, in particular to a high-throughput analysis method for predicting phage hosts based on second-generation sequencing technology. BACKGROUND
[0002] In 2005, Roche launched the first second-generation sequencer, Roche 454, and life sciences entered the era of high-throughput sequencing. Subsequently, with the launch of the Illumina series of sequencing platforms, the price of second-generation sequencing was greatly reduced, promoting the popularization of high-throughput sequencing in various research fields of life sciences. Although third-generation sequencing technology has emerged, it is still the most mainstream conventional research method at present due to the high cost of sequencing and imperfect analysis software, and is widely used in scientific research. Second-generation sequencing (NGS) is also known as high-throughput sequencing (HTS), which is a DNA sequencing technology based on PCR and gene chips, and it introduces reversible termination ends, thereby realizing sequencing by synthesis. In second-generation sequencing, a single DNA molecule must be amplified into a gene cluster composed of the same DNA, and then replicated synchronously to enhance the intensity of the fluorescent signal to read the DNA sequence. With the increase of read length, the cooperativity of gene cluster replication decreases, leading to a decrease in base sequencing quality, which strictly limits the read length of second-generation sequencing (50-250 bp, with a maximum of 500 bp). Therefore, second-generation sequencing has the characteristics of high throughput and short read length.
[0003] Due to the rapid development of sequencing technology, metagenomics and viral metagenomics have emerged, and their research objects are mainly bacteria and viruses, analogues and the genetic information they carry in microbial communities. Traditional microbial research relies on laboratory culture, and the emergence of metagenomics and viral metagenomics fills the gap in the research of microorganisms that cannot be cultured in the laboratory. Phages are a class of viruses that can invade bacteria and cause them to lyse, and they are the most abundant biological species in the Earth's biosphere. As a mobile genetic element, they can also transmit genetic material between bacteria. Therefore, phages play an important role in regulating bacterial biomass, maintaining biodiversity, horizontal gene transfer, and biochemical cycles in the entire biosphere. The host range of phages is very narrow and specific, mainly at the genus or species level, so they can be used for precise regulation of microbial communities. Studying the interaction between phages and microbial communities, i.e., the host of phages, can more easily find strains that play an important role in health and disease, providing new targets and tools for disease treatment and drug development.
[0004] There are many tools for phage host prediction, each with different focuses and advantages, and the prediction results often differ greatly. How to effectively evaluate and screen multiple prediction results to obtain more accurate phage and host interaction relationships is a key problem in the field of bioinformatics that has been concerned and tried to solve. Although individual metagenomic or viromic analysis based on second-generation sequencing technology is relatively mature, how to save research costs and efficiently use a set of sequencing data to complete bacterial and viral genome assembly and evaluate and screen more accurate interaction relationships between the two has become an urgent need. SUMMARY
[0005] The purpose of the present application is to provide an analysis method for high-throughput prediction of phage hosts based on second-generation sequencing technology. The present application provides a complete process for obtaining phage and bacterial genomes from second-generation sequencing data and accurately evaluating phage hosts, enabling researchers to efficiently use a set of sequencing data to obtain more comprehensive analysis results and allowing non-bioinformatics professionals to independently complete high-throughput sequencing data analysis. This achieves the purposes of optimizing researchers' work efficiency, improving the reuse of second-generation sequencing data, and reducing research costs. The present application proposes a reliable analysis method for high-throughput prediction of phage hosts based on second-generation sequencing technology, which is simple to implement and widely applicable, and solves the technical problem of simultaneously completing bacterial and viral genome assembly in the prior art.
[0006] According to the purpose of the present application, an analysis method for high-throughput prediction of phage hosts based on second-generation sequencing technology is provided, comprising the following steps:
[0007] (1) Quality control, filtering, assembly, binning, and de-redundancy of raw sequencing data to obtain non-redundant bacterial microbial assembly genomes MAGs;
[0008] In step (1), the de-redundancy of MAGs is performed, and the specific steps are as follows:
[0009] S1: Filtering genomes with a length <50kb;
[0010] S2: Identifying genes in MAGs based on a dynamic programming gene finding algorithm for prokaryotes and translating the corresponding protein sequences;
[0011] S3: Using the single-copy nature of genes to effectively align the completeness and contamination of the genome, and filtering low-quality bacterial genomes with a sequence completeness <80% or a contamination >10%;
[0012] S4: Primary and secondary clustering based on genome distance and average nucleotide identity, and selecting the longest genome within the same cluster as the optimal genome;
[0013] (2) The bacterial genes obtained from step (1) are used to identify single-copy marker genes and construct a phylogenetic tree using a hidden Markov model, and finally compared with known bacterial and archaeal phylogenetic trees for species annotation;
[0014] (3) Quality control, filtering, and assembly of raw sequencing data to obtain viral contigs sequences, and quality control of viral contigs to obtain high-quality viral contigs;
[0015] The quality control analysis of viral contigs in step (3) includes the following specific steps:
[0016] S1: Filter contigs with a length of <1.5 kb;
[0017] S2: Compare the sequence with the viral genome to estimate the integrity, 0-5% error matching is considered as high-quality contig, 5-10% error matching is medium-quality contig, and more than 10% error matching is low-quality contig and needs to be filtered, finally retaining high and medium-quality viral contigs;
[0018] (4) Calculate the average nucleotide identity (ANI) of viral contigs, retain contigs with ANI>95%, compare the amino acid level genes of contigs with the viral subset in TrEMBL database, and thus perform species annotation at the family level of viruses, perform species annotation at the genus level based on K-mer characteristics through supervised machine learning method, and finally according to the annotation results of family and genus level of viruses, complete the annotation of other classification levels of viral contigs from the known taxonomy library;
[0019] (5) Based on bacterial MAGs and viral contigs, at least three different methods are used to predict phage hosts, and the prediction results are evaluated for accuracy from the purity and agreement indicators;
[0020] The method for predicting phage hosts in step (5) includes any three methods selected from the following four methods or using the following four methods:
[0021] Method 1: Prediction method of phage and its host relationship based on CRISPR-Cas system;
[0022] Method 2: Prediction method of active phage from bacterial genomes based on sequence similarity alignment and machine learning classification of genetic characteristics;
[0023] Method 3: Prediction method of phage host based on dynamic programming algorithm;
[0024] Method 4: a method for predicting bacteriophage hosts based on virus and its host oligonucleotide frequency;
[0025] The precision evaluation in step (5) is specifically as follows:
[0026] S1: purity index evaluation: the index is an evaluation index for measuring the consistency of a single predicted bacteriophage host; the hosts of a virus are extracted, and the proportion of the most common host is counted at different species levels, and the specific calculation formula is as follows:
[0027]
[0028] Suppose there are n virus contigs, and the predicted host of a contig is N, wherein i∈(1, n), j∈(1, N), r∈(1, 7), V ir represents the proportion of the jth host at the rth species level of the ith virus, m ir is the maximum value of V ir , that is, the proportion of the most common host;
[0029] The m ir of n at the rth species level is obtained, and the average value is purity, when putiry is greater than 50%, the host prediction result of the virus is reserved;
[0030] S2: agreement index evaluation: the index is an index for measuring the consistency of the predicted hosts between two methods in the method for predicting bacteriophage hosts; viruses existing in both prediction methods are screened, each virus has corresponding m ir at the same species level, and whether the two are the same is compared; the proportion of all viruses with the same m ir is the agreement index; all prediction methods are compared in pairs, if the agreement index of two methods is higher than 5%, the prediction host result of the method is reserved;
[0031] The bacteriophage hosts in the prediction results of purity index evaluation and agreement index evaluation are reserved at the same time, which are determined as bacteriophage hosts.
[0032] Preferably, the quality control and filtering of the raw sequencing data in step (1) specifically comprises: removing adapter sequences, removing reads with a proportion of N greater than 10%; performing base quality analysis based on base composition and quality distribution, removing low-quality reads with a quality value Q≤30; and removing reads with a base number greater than 85% aligned to a sample-derived host genome based on a short sequence alignment algorithm, to finally obtain clean reads of high quality.
[0033] Preferably, the assembly in step (1) specifically comprises: obtaining contigs based on a K-mer iterative de Bruijn graph assembly algorithm, and filtering out short sequences with a length less than 2.5 kb.
[0034] Preferably, the binning in step (1) specifically comprises: performing iterative binning using a k-medoids clustering algorithm to obtain bins.
[0035] Preferably, the quality control and filtering of the raw sequencing data in step (3) specifically comprises: removing adapter sequences, removing reads with a proportion of N greater than 10%; performing base quality analysis based on base composition and quality distribution, removing low-quality reads with a quality value Q≤30; and removing reads with a base number greater than 85% aligned to a bacterial contaminant genome based on a short sequence alignment algorithm, to finally obtain clean reads of high quality.
[0036] Preferably, the assembly in step (3) specifically comprises: obtaining contigs based on a K-mer iterative de Bruijn graph assembly algorithm, and filtering out short sequences with a length less than 1.5 kb.
[0037] Overall, compared with the prior art, the above technical solutions conceived by the present application mainly have the following technical advantages:
[0038] (1) The present application provides a complete process for obtaining phage and bacterial genomes from second-generation sequencing data and accurately evaluating phage hosts, so that researchers can efficiently utilize a set of sequencing data to obtain more comprehensive analysis results, and non-biological information professionals can independently complete high-throughput sequencing data analysis. The present application optimizes the work efficiency of researchers, improves the reuse of second-generation sequencing data, and reduces the cost of scientific research. The present application provides a reliable high-throughput phage host prediction analysis method based on second-generation sequencing technology, which is simple to implement and widely applicable.
[0039] (2) The present invention has more efficient use of analytical data, a more comprehensive analytical process, and more accurate analytical results. It solves the problem of the current phage host prediction methods being numerous and the prediction process being not standardized, and provides convenience and technical support for researchers. Attached Figure Description
[0040] Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0042] The specific processes of some embodiments of the present invention are as follows:
[0043] S1: Perform quality control and filtering, splicing and assembly, binning and redundancy removal on the raw sequencing data to obtain non-redundant, high-quality bacterial microbial assembled genomes (MAGs).
[0044] Sequencing data filtering (bacteria): Adapter sequences were removed, and reads with an N ratio greater than 10% were removed; base quality analysis was performed based on base composition and quality distribution, removing low-quality reads with a quality value Q ≤ 30; after quality control, reads were aligned to the source genome of the sample using a short sequence alignment algorithm, removing reads with more than 85% of bases aligned to the source genome, ultimately obtaining high-quality clean reads; the raw sequencing data was first evaluated using FastQC, and then filtered using Trimmomatic and Bowtie2 software based on the evaluation results;
[0045] Concatenation and assembly: Contigs are obtained based on the K-mer iterative de Bruijn graph assembly algorithm. Short sequences with a length of less than 2.5kb are filtered out. Specifically, the software Megahit is used for implementation.
[0046] Binning: Iterative binning is performed using the k-medoids clustering algorithm to obtain bins, specifically implemented using the software MetaBAT2.
[0047] The redundancy removal process described in step S1 to obtain high-quality MAGs involves the following steps:
[0048] a. Filter genomes with a length <50kb;
[0049] b. Identify genes in MAGs and translate corresponding protein sequences based on prokaryotic-based dynamic programming gene-finding algorithm;
[0050] c. Estimate genome completeness and contamination effectively using single copy genes, preferably filter low quality bacterial genomes (completeness <80% or contamination >10%);
[0051] d. Primary and secondary clustering by genome distance estimation and average nucleotide identity, select the longest genome in the same cluster as the optimal genome.
[0052] Specifically implemented using software dRep.
[0053] S2: From the bacterial genes obtained in step S1, identify single copy marker genes using hidden Markov model and construct phylogenetic tree, and finally compare with known bacterial and archaeal phylogenetic tree for species annotation, specifically using software GTDB-Tk to implement;
[0054] S3: Quality control and filtering, assembly of raw sequencing data, obtaining viral contigs sequence, quality control of viral contigs to obtain high-quality viral contigs;
[0055] The specific process of the step S3 of quality control and filtering, assembly of raw sequencing data, obtaining viral contigs sequence is as follows:
[0056] Sequencing data filtering (virus): remove adapter sequences, remove reads with N ratio greater than 10%; base quality analysis based on base composition and quality distribution, remove low-quality reads with quality value Q≤30; quality-controlled reads based on short sequence alignment algorithm, remove reads with more than 85% of bases aligned to contaminant genomes such as bacteria, finally obtain high-quality clean reads. Specifically, use fastqc to first evaluate the original sequencing data, and use software Trimmomatic and software bowtie2 for filtering according to the evaluation results;
[0057] Assembly: contigs are obtained based on K-mer iterative de Bruijn graph assembly algorithm, preferably filter out short sequences with length below 1.5kb, specifically using software megahit to implement.
[0058] The quality control analysis of viral contigs in step S3 is as follows:
[0059] a. Filter contigs with length <1.5kb, specifically implemented using shell language;
[0060] b. Estimate the integrity of contigs by comparing the sequence to the complete viral genome, 0-5% error match is considered high-quality contigs, 5-10% error match is considered medium-quality contigs, and more than 10% error match is considered low-quality contigs and needs to be filtered, finally, high-quality and medium-quality viral contigs are retained, and CheckV is used to achieve the above steps.
[0061] S4: Calculate the average nucleotide identity (ANI) of viral contigs, and retain contigs with ANI>95%, FastANI is used to achieve the above steps; obtain family-level annotations by aligning amino acid-level genes with a subset of viruses in the TrEMBL database, demovir is used to achieve the above steps; perform genus-level species annotation of viral genomes based on K-mer features through supervised machine learning methods, VirusTaxo is used to achieve the above steps; based on the annotation results of family and genus levels of viruses, other classification level annotations of viral sequences are completed from the known taxonomy library, R or python language is used to achieve the above steps.
[0062] S5: Based on bacterial MAGs and viral contigs, a variety of methods are used to predict phage hosts, and the prediction results are evaluated for accuracy from the purity index and the agreement index.
[0063] The variety of methods for predicting phage hosts in step S5 includes:
[0064] a. A method for predicting the relationship between phages and their hosts based on the CRISPR-Cas system, and CRISPRCasFinder is used to achieve the above steps;
[0065] b. A method for predicting active phages from bacterial genomes based on sequence similarity alignment and genetic feature machine learning classification, and Prophage Hunter is used to achieve the above steps;
[0066] c. A method for predicting phage hosts based on a dynamic programming algorithm, and blastn is used to achieve the above steps;
[0067] d. A method for predicting phage hosts based on viral and host oligonucleotide frequencies, and VirHostMatcher is used to achieve the above steps.
[0068] The accuracy evaluation in step S5 includes the following steps:
[0069] a、purity index: the index is an evaluation index for measuring the consistency of the predicted phage host. The host of a virus is extracted, and the proportion of the most common host is counted at different species levels (domain, class, order, family, genus, species), and the specific calculation formula is:
[0070]
[0071] Suppose there are n virus contigs, and the predicted host of a contig is N, wherein i∈(1, n), j∈(1, N), r∈(1, 7), V ir represents the proportion of the jth host of the ith virus at the rth species level, m ir is the maximum value of V ir , that is, the proportion of the most common host.
[0072] The obtained n m ir at the rth species level is taken as the purity, preferably, when the purity is greater than 50%, the host prediction result of the virus is reserved;
[0073] b、agreement index: the index is an index for measuring the consistency of the predicted host between two methods. The viruses existing in both prediction methods are screened, and each virus has a corresponding m ir at the same species level, and whether the two are the same is compared. The proportion of all viruses with the same m ir is the agreement index. All prediction methods are compared with each other, preferably, if the agreement index of a method with other methods is less than 5%, the prediction host result of the method is not reserved.
[0074] Specifically, python or R language is used to realize.
[0075] In summary, the present application develops an integrated analysis tool based on second-generation sequencing to obtain more comprehensive and accurate phage host analysis results, thereby solving the problem that the phage host prediction methods are various and the prediction process is not quite standardized, and the analysis result is more accurate.
[0076] Those skilled in the art can easily understand that the above description is only the preferred embodiment of the present application, and is not used to limit the present application, and any modification, equivalent replacement and improvement made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. An analysis method for predicting bacteriophage hosts in high throughput based on a next-generation sequencing technique, characterized by, Comprising the following steps: (1) Quality control, filtering, assembly, binning and dereplication of raw sequencing data to obtain non-redundant bacterial microbial assembly genomes MAGs; The dereplication in step (1) obtains MAGs, and the specific steps are: S1: filtering genomes with length <50kb; S2: identifying genes in MAGs based on a dynamic programming gene finding algorithm for prokaryotes and translating the corresponding protein sequences; S3: using the single copy nature of genes to effectively align the completeness and contamination of the genome, filtering low-quality bacterial genomes with sequence completeness <80% or contamination >10%; S4: primary and secondary clustering by genome distance and average nucleotide identity, selecting the longest genome in the same cluster as the optimal genome; (2) From the bacterial genes obtained in step (1), identify single-copy marker genes using a hidden Markov model and construct a phylogenetic tree, and finally compare with known bacterial and archaeal phylogenetic trees for species annotation; (3) Quality control, filtering, assembly of raw sequencing data to obtain viral contigs sequences, and quality control of viral contigs to obtain high-quality viral contigs; The quality control analysis of viral contigs in step (3) has the following specific steps: S1: filtering contigs with length <1.5kb; S2: comparing the sequence with the viral genome to estimate the integrity, 0-5% error matching is considered as high-quality contig, 5-10% error matching is medium-quality contig, and more than 10% error matching is low-quality contig and needs to be filtered, finally retaining high-quality and medium-quality viral contigs; (4) Calculate the average nucleotide identity ANI of viral contigs, retain contigs with ANI>95%, compare the genes at the amino acid level of contigs with the viral subset in the TrEMBL database, and thus perform species annotation at the family level of viruses, perform species annotation at the genus level based on K-mer characteristics through supervised machine learning method, and finally according to the annotation results at the family and genus levels of viruses, complete the annotation of other classification levels of viral contigs from the known taxonomy library; (5) Based on bacterial MAGs and viral contigs, at least three different methods are used to predict phage hosts, and the prediction results are evaluated for accuracy from the purity index and the consistency index; The method for predicting phage hosts in step (5) includes selecting three methods from the following four methods or using the following four methods: Method 1: phage and its host relationship prediction method based on CRISPR-Cas system; Method 2: predicting active phages from bacterial genomes based on sequence similarity alignment and genetic feature machine learning classification; Method 3: phage host prediction method based on dynamic programming algorithm; Method 4: phage host prediction method based on machine learning classification of genetic features. Method 4: a method for predicting bacteriophage hosts based on virus and its host oligonucleotide frequency; The accuracy evaluation in step (5) is specifically as follows: S1: purity index evaluation: the index is an evaluation index for measuring the consistency of a single predicted bacteriophage host; the hosts of a virus are extracted, and the proportion of the most common host is counted at different species levels, and the specific calculation formula is: Assuming there are n viral contigs, a contig has N predicted hosts, where i ∈ (1, n), j ∈ (1, N), r ∈ (1, 7), V ir represents the proportion of the jth host of the rth species level of the ith virus, m ir is the maximum value of V ir , i.e., the most common host proportion; The n m ir The average value is purity, when putiry is greater than 50%, the host prediction result of the virus is retained; S2: agreement index evaluation: the index is an index for measuring the consistency of the predicted hosts between two methods in the method for predicting the hosts of bacteriophages; viruses existing in both prediction methods are screened, and each virus has corresponding m ir values in the same species water; ir Compare whether they are the same; the proportion of all viruses with the same m ir values is the agreement index; all prediction methods are compared in pairs, and if the agreement index of the two methods is higher than 5%, the prediction host result of the method is retained; The bacteriophage hosts remaining in the prediction results of the purity index evaluation and the agreement index evaluation are determined as bacteriophage hosts.
2. The method of claim 1, wherein the method is a high-throughput method for predicting a phage host based on a next-generation sequencing technology. The quality control and filtering of the original sequencing data in step (1) is specifically as follows: removing adapter sequences, removing reads with a proportion of N greater than 10%; base quality analysis based on base composition and quality distribution, removing low-quality reads with a quality value Q≤30; The quality-controlled reads are based on a short sequence alignment algorithm, and reads with a base number exceeding 85% aligned to the sample source host genome are removed, and finally clean reads of high quality are obtained.
3. The method for analyzing a host of a phage based on a next generation sequencing technology according to claim 1 or 2 or the same, characterized in that, The splicing assembly in step (1) is specifically as follows: a de Bruijn graph assembly algorithm based on K-mer iteration is used to obtain contigs, and short sequences with a length below 2.5 kb are filtered out.
4. The method of claim 1 or 2, wherein the method is based on a next-generation sequencing technique. The binning in step (1) is specifically as follows: using a k-medoids clustering algorithm for iterative binning to obtain bins.
5. The method for analyzing a host of a phage based on a next-generation sequencing technology according to claim 1 or 2, wherein, In step (3), the quality control and filtering of the original sequencing data are specifically as follows: removing adapter sequences, removing reads with a proportion of N greater than 10%; base quality analysis based on base composition and quality distribution, removing low-quality reads with a quality value Q≤30; The quality-controlled reads are based on a short sequence alignment algorithm, and reads with a base number exceeding 85% aligned to the bacterial contaminant genome are removed, and finally clean reads of high quality are obtained.
6. The method of claim 1 or 5, wherein the method is based on a next-generation sequencing technology. The splicing assembly in step (3) is specifically as follows: a de Bruijn graph assembly algorithm based on K-mer iteration is used to obtain contigs, and short sequences with a length below 1.5 kb are filtered out.
Citation Information
Patent Citations
Metagenome data analysis method based on next-generation sequencing technology
CN112071366A
High-throughput macro fungus molecular identification method based on next-generation sequencing technology
CN115094129A