Method and device for quantifying enterovirome
By constructing a virus cluster-specific marker protein database and comparing the intestinal metagenome sequencing data, the problems of low accuracy and low efficiency of enterovimes in the prior art were solved, and rapid and accurate enterovimes quantification was achieved.
Patent Information
- Application Number
- CN202311813511.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has low accuracy in quantitative analysis of enterovirus groups, and due to the large database size and low quantitative efficiency, it is impossible to quantify enterovirus groups quickly and accurately.
By constructing a virus cluster-specific marker protein database, the intestinal metagenome sequencing data were compared to this database to achieve quantification of the enteroviral group. This method improves the accuracy of virus quantification by eliminating the near-source sequence interference of other microorganisms, and greatly compresses the database sequence scale, improving the efficiency of tool operation.
The rapid and accurate quantification of the enterovitrome is achieved, the quantitative efficiency is improved, and it can be used in the quantification of the viral group with a wide range of intestinal microbiota metagenomic sequencing data.
Smart Images

Figure CN120220810A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of gut microbiota, and more particularly, to a method and device for quantifying the gut virome. Background Art
[0002] The viral community is an important component of the human gut microbiota. It is estimated that there are approximately 10 9-10 viral particles per gram of feces. Researchers refer to the collection of all viruses in the gut as the gut virome. The gut virome is highly diverse and exhibits high inter-individual heterogeneity. With the rapid development of metagenomic sequencing technology, researchers can directly detect the nucleic acid information of all microorganisms in a community without relying on isolation and cultivation. Based on metagenomic sequencing technology and the development of bioinformatics tools for gut bacteria, our understanding of gut bacteria and other microorganisms has been greatly improved. However, due to the lack of gut virome analysis tools, our current understanding of the gut virome is very limited.
[0003] Existing studies have shown that the characteristics of gut virome changes are related to various diseases such as inflammatory bowel disease, diabetes, hypertension, and colon cancer. The gut virome has great potential to regulate the gut microbiota and thus affect human health. Therefore, developing relevant tools for the gut virome and continuously enhancing our understanding of the gut virome are of great significance for comprehensively analyzing the mechanisms of disease occurrence and development.
[0004] In the study of the human gut virome, it is necessary to characterize the gut virome characteristics of individuals based on the differences in the types and relative abundances of viruses in the guts of different individuals. Based on metagenomic sequencing data of the gut microbiota, researchers can quickly complete this process - virome quantitative analysis. Through virome quantitative analysis, researchers can carefully understand the composition and content of viruses in an individual's gut, and then be able to identify the differential characteristics of the gut virome reflected by factors such as time, intervention methods, or disease status at the population level. Therefore, accurately and efficiently implementing virome quantitative analysis is crucial for further understanding the functional characteristics of the gut virome and revealing its relationship with human disease and health.
[0005] Currently, the existing technical solutions for realizing gut virome quantitative analysis are mainly embodied in the viral whole-genome sequence alignment strategy. The specific steps are as follows: 1) Construct a reference database of viral genome sequences; 2) Use a read alignment tool to align the metagenomic sequencing data of the sample to the viral genome reference database, and obtain reliable alignment results by setting alignment filtering conditions such as similarity and coverage; 3) Based on the alignment results, calculate the relative abundances of various viruses in the reference database. Specifically, according to different ideas for constructing the reference database, it can be roughly divided into two categories: based on public libraries and based on self-built libraries.
[0006] Quantitative tools based on public libraries refer to using publicly available viral genome sequences such as viral genomes in NCBI Refseq and viral genomes in ICTV (International Committee on Taxonomy of Viruses) as reference databases. Quantitative tools based on self-built libraries refer to researchers de novo assembling metagenomic sequencing data from their research outputs and predicting viral sequences to de novo construct a reference genome database for alignment. Both strategies are based on the complete genome sequences of viruses as reference databases for quantifying the gut virome.
[0007] Existing quantitative techniques all use the complete genome sequences of viruses as reference databases for nucleotide-level sequence alignment. Since metagenomic sequencing data of gut microbiota simultaneously contains reads from both viral and bacterial genomes, and there are similar gene sequence fragments between the gut virome and its host bacteria. Inevitably, existing techniques will align reads from host bacteria to the viral genome, unable to exclude the interference of bacterial sequences on viral quantification, resulting in low accuracy.
[0008] In addition, due to the rapid growth of the scale of gut virus genomes, the database is constantly increasing. The quantitative strategy based on alignment of complete viral genome sequences is inefficient. Since 2020, multiple teams have successively published multiple large gut virome databases such as Gut Phage Database (GPD), Metagenomic Gut Virus (MGV), Cenote HumanVirome Database (CHVD), and Gut Virome Database (GVD), etc. Among them, the largest GPD database has identified 140,000 viral species, and the total base number of the genome database has reached 5.3 Gb11. The scale of the gut virus genome database is growing rapidly. The existing quantitative strategy for the gut virome based on alignment of complete genome sequences has a large database scale and low quantitative efficiency.
[0009] Therefore, there is currently no effective solution to how to quickly and accurately quantify the gut virome. Summary of the Invention
[0010] The main objective of the present invention is to provide a method and device for quantifying the gut virome to solve the problem of low accuracy in quantifying the gut virome in the prior art.
[0011] To achieve the above objective, according to one aspect of the present invention, there is provided a method for quantifying the gut virome, the method comprising: constructing a database of marker proteins specific to viral clusters; aligning the metagenomic sequencing data of the gut with the database of marker proteins to achieve quantification of the gut virome.
[0012] Further, constructing a virus cluster-specific marker protein database includes: constructing virus clusters in the intestine; screening marker proteins of the virus clusters, and constructing a virus cluster-specific marker protein database.
[0013] Further, constructing virus clusters in the intestine includes: constructing protein clusters of intestinal virus genomes; de novo constructing virus clusters in the intestine based on the characteristics of the common protein clusters among intestinal virus genomes.
[0014] Further, constructing protein clusters of intestinal virus genomes includes: collecting sequences of known intestinal virus genomes and performing quality control to obtain a set of intestinal virus genomes; identifying and translating coding genes on the intestinal virus genomes to obtain protein sequences; clustering all virus proteins according to the similarity among the protein sequences, thereby constructing protein clusters of intestinal virus genomes.
[0015] Further, screening marker proteins of intestinal virus clusters and constructing a virus cluster-specific marker protein database using the marker proteins includes: screening protein clusters related to each virus cluster according to the information of the common genomes between the protein clusters and the virus clusters, thereby establishing an association relationship between the virus clusters and the protein clusters; calculating the Jaccard coefficient between the protein clusters and the virus clusters according to the occurrence of the protein clusters related to the virus clusters in each genome of the virus clusters, and screening protein clusters that meet the following two conditions simultaneously, denoted as candidate marker protein clusters: 1) the Jaccard coefficient between the protein cluster and the virus cluster ≥ 0.5; 2) the Jaccard coefficient between the protein cluster and other virus clusters < 0.2; selecting representative protein sequences for the candidate marker protein clusters, and performing homology quality control and determining the number of marker proteins for the representative protein sequences; selecting the determined number of marker proteins in descending order of the Jaccard coefficient between the protein clusters and the virus clusters, thereby constructing a virus cluster-specific marker protein database.
[0016] Further, selecting representative protein sequences for the candidate marker protein clusters and performing homology quality control and determining the number of marker proteins for the representative protein sequences includes: pairwise comparing the internal protein sequences of the candidate marker protein clusters, and screening the protein sequence that is closest to other protein sequences as the representative sequence of a certain candidate marker protein cluster, constructing a set of representative protein sequences of all candidate marker protein clusters; removing the representative protein sequences with homologous sequences to other microorganisms outside the virus clusters from the set of representative protein sequences of the candidate marker protein clusters; taking the minimum value of the number of the remaining representative protein sequences and 20% of the number of protein sequences on the virus genomes included in the virus cluster as the target number of marker proteins; preferably, controlling the number of marker proteins for each virus cluster ≤ the number threshold, preferably 50.
[0017] Further, comparing the gut metagenomic sequencing data with the marker protein database to achieve quantification of the gut virome includes: aligning the reads in the gut microbiota metagenomic sequencing data to the marker protein database; taking the viral clusters as units, counting the number of reads aligned to each marker protein and dividing by the length of the marker protein to calculate the read coverage depth of each marker protein, and further obtaining the read coverage depth of the viral clusters; excluding false-positive viral clusters in the sequencing data, and denoting the viral clusters in which the number of marker proteins with read coverage among all marker proteins is more than 50% as identified viral clusters; normalizing the read coverage depths of the viral clusters in all identified viral clusters to obtain the final relative abundances of the viral clusters contained in the sample to be tested.
[0018] Further, for the marker proteins of each viral cluster, removing a predetermined number of marker proteins with the highest and lowest coverage depths in the ranking, and taking the mean value of the coverage depths of the remaining marker proteins in the middle as the read coverage depth of the viral cluster.
[0019] According to the second aspect of the present application, there is provided an apparatus for quantifying the gut virome, the apparatus including: a construction unit configured to construct a marker protein database specific to viral clusters; an alignment and quantification unit configured to compare gut metagenomic sequencing data with the marker protein database to achieve quantification of the gut virome.
[0020] Further, the construction unit includes: a viral cluster construction module configured to construct viral clusters in the gut; a marker protein database construction module configured to screen marker proteins of viral clusters and construct a marker protein database specific to viral clusters.
[0021] Further, the viral cluster construction module includes: a protein cluster construction sub-module configured to construct protein clusters of gut virus genomes; a viral cluster construction sub-module configured to de novo construct viral clusters in the gut based on the characteristics of common protein clusters among gut virus genomes.
[0022] Further, the protein cluster construction sub-module includes: a collection module configured to collect sequences of known gut virus genomes and perform quality control to obtain a set of gut virus genomes; an identification and translation module configured to identify and translate coding genes on gut virus genomes to obtain protein sequences; a protein clustering module configured to cluster all viral proteins according to the similarity between protein sequences, thereby constructing protein clusters of gut virus genomes.
[0023] Furthermore, the labeled protein database construction module includes: a virus cluster-protein cluster association module, which is configured to screen the protein clusters related to each virus cluster according to the information of the common genomes between the protein clusters and the virus clusters, so as to establish the association relationship between the virus clusters and the protein clusters; a candidate labeled protein cluster screening module, which is configured to calculate the Jaccard coefficient between the protein cluster and the virus cluster according to the occurrence of the protein clusters related to the virus cluster in each genome of the virus cluster, and screen the protein clusters that meet the following two conditions simultaneously, which are recorded as candidate labeled protein clusters: 1) the Jaccard coefficient between the protein cluster and the virus cluster ≥ 0.5; 2) the Jaccard coefficient between the protein cluster and other virus clusters < 0.2; a quality control and quantity determination module, which is configured to select representative protein sequences for the candidate labeled protein clusters, and perform homology quality control and labeled protein quantity determination on the representative protein sequences; a labeled protein database construction sub-module, which is configured to select the determined number of labeled proteins according to the Jaccard coefficient between the protein cluster and the virus cluster from high to low, so as to construct a virus cluster-specific labeled protein database.
[0024] Furthermore, the quality control and quantity determination module includes: a representative protein sequence screening module, which is configured to perform pairwise alignment between the internal protein sequences of the candidate labeled protein clusters, and screen the protein sequence that is closest to other protein sequences as the representative protein sequence of a certain candidate labeled protein cluster, and construct a set of representative protein sequences of all candidate labeled protein clusters; a quality control removal module, which is configured to remove the representative protein sequences with homologous sequences to other microorganisms outside the virus cluster from the set of representative protein sequences of the candidate labeled protein clusters; a quantity determination module, which is configured to take the minimum value of the number of remaining representative protein sequences and 20% of the number of protein sequences on the viral genomes included in the virus cluster as the target quantity of the labeled proteins; preferably, in the quantity determination module, the number of labeled proteins for each virus cluster is controlled ≤ the quantity threshold, preferably 50.
[0025] Furthermore, the alignment and quantification unit includes: an alignment module, which is configured to align the reads in the intestinal microbiota metagenomic sequencing data to the labeled protein database; a read coverage depth calculation module, which is configured to take the virus cluster as a unit, count the number of reads aligned to each labeled protein and divide it by the length of the labeled protein to calculate the read coverage depth of each labeled protein, and further obtain the read coverage depth of the virus cluster; a false positive exclusion module, which is configured to exclude the false positive virus clusters in the sequencing data, and record the virus clusters in which the number of labeled proteins with read coverage accounts for more than 50% of all labeled proteins as the identified virus clusters; a quantification module, which is configured to normalize the read coverage depths of the virus clusters in all identified virus clusters to obtain the final relative abundances of the virus clusters contained in the test sample.
[0026] According to a third aspect of the present application, there is provided a computer-readable storage medium, which includes a stored program. When the program runs, it controls the device where the storage medium is located to execute any of the above methods for quantifying the gut virome.
[0027] According to a fourth aspect of the present application, there is provided a processor for running a program. When the program runs, it executes any of the above methods for quantifying the gut virome.
[0028] By applying the technical solution of the present invention, by constructing a marker protein database specific to virus clusters, the so-called "specific" means that this marker protein database excludes the interference of closely related sequences of other microorganisms, thus improving the accuracy of virus quantification. Aligning the gut metagenomic sequencing data to this marker protein database enables rapid quantification of the virome in gut metagenomic data. Compared with whole-genome sequence alignment, the marker protein database of the present application greatly compresses the scale of database sequences and significantly improves the running efficiency of the tool. Therefore, the quantification process of the present application is suitable for being extended to the virome quantification of a wide range of gut microbiota metagenomic sequencing data. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0030] Figure 1 It shows a schematic flowchart of a method for quantifying the gut virome according to Embodiment 1 of the present invention.
[0031] Figure 2 It shows a detailed flowchart of a method for quantifying the gut virome according to Embodiment 2 of the present invention.
[0032] Figure 3A and Figure 3B It shows an evaluation of the consistency between the quantitative abundance and the theoretical abundance obtained by the method for quantifying the gut virome of the present application through the constructed test dataset in Embodiment 2 of the present invention. Among them, Figure 3A It shows that the virome abundance quantified in this embodiment is higher than the existing quantitative results based on whole-genome alignment; Figure 3B It shows that the quantification accuracy of this embodiment is higher than that of the existing quantification methods.
[0033] Figure 4 It shows that the method for quantifying the gut virome of the present application is verified for its quantification accuracy in public data according to Embodiment 3 of the present invention.
[0034] Figure 5The structural schematic diagram of the device for quantifying the gut virome according to Embodiment 4 of the present invention is shown. Detailed implementation manners
[0035] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the embodiments.
[0036] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0037] It should be noted that the terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to describe the embodiments of the present application herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0038] For the convenience of description, some nouns or terms related to the embodiments of the present application are described below:
[0039] Gut microbiota: A large number of microorganisms existing in the animal intestine. This group of microorganisms rely on the animal intestine for life and at the same time help the host complete a variety of physiological and biochemical functions. In the present application, it specifically refers to human gut microbiota.
[0040] Gut virome: Among the human gut microbiota population, in addition to the most widely studied bacteria and archaea (the two account for ~94% of the total DNA of gut microbiota biomass), viruses (~5%), protozoa (~0.2%) and fungi (~0.1%) constitute a smaller but also very important part of the gut microbial community. Among them, the viruses living in the human intestine are divided into three categories: bacteriophages, DNA viruses and RNA viruses. Among them, bacteriophages are viruses that infect prokaryotic cells, and the latter two are viruses that infect eukaryotic cells.
[0041] Metagenome: Refers to the sum total of the genetic material of all microorganisms in a specific environment, including culturable and non-culturable microorganisms.
[0042] Metagenomic sequencing: Generally refers to extracting genomic DNA from environmental samples and performing high-throughput sequencing. It breaks away from the technical limitations of microbial isolation and culture in traditional research, interprets the diversity and abundance of microbial populations at the genomic level, and explores the relationships between microorganisms, the environment, and the host.
[0043] As mentioned in the background art, there are problems such as low accuracy and low quantification efficiency in existing methods for quantifying the human gut virome. To improve this situation, based on an in-depth understanding of the sequence characteristics of gut viruses and making full use of the current rich genomic data resources of gut viruses, the present application creatively proposes an improved scheme for quantitative analysis of the human gut virome based on the "viral taxon - marker protein" strategy. The improvement principle and idea of the improved scheme of the present application are as follows:
[0044] 1) Based on the publicly available gut virus genomes, de novo clustering is performed according to the characteristics of common proteins to construct viral taxons. Furthermore, marker proteins are screened for the viral taxons to construct a reference database of marker proteins. By excluding the interference of closely related sequences of other microorganisms from the marker protein database, the accuracy of virus quantification is improved.
[0045] 2) A reference database is constructed based on the marker protein sequences, and the metagenomic sequencing reads are aligned to the marker protein database to achieve virus classification and quantification. The scheme of the present application uses the marker protein as the reference database, which significantly compresses the scale of the database sequences compared with the whole genome, improves the running efficiency of the tool, and enhances the quantification efficiency.
[0046] Therefore, the scheme of the present application can achieve: By using the metagenomic sequencing data of the human gut microbiota, the types and relative abundances of gut taxons in the sample can be quickly output within a short time, realizing accurate and efficient quantitative analysis of the gut virome.
[0047] Based on the above inventive concept, the applicant has proposed a series of schemes of the present application.
[0048] Example 1
[0049] This example provides a method for quantifying the gut virome. Among them, Figure 1 The flow schematic diagram of the method for quantifying the gut virome in this example is shown. The method includes:
[0050] S101, constructing a marker protein database specific to the virus cluster;
[0051] S102. Align the gut metagenomic sequencing data with the marker protein database to achieve quantification of the gut virome.
[0052] In this embodiment, by constructing a marker protein database specific to virus clusters, the so-called "specific" means that this marker protein database excludes the interference of homologous sequences of other microorganisms, thus improving the accuracy of virus quantification. Aligning the gut metagenomic sequencing data with this marker protein database enables rapid quantification of the virome in gut metagenomic data. Compared with whole-genome sequence alignment, the marker protein database of this application significantly compresses the scale of database sequences and greatly improves the operation efficiency of the tool. Therefore, the quantification process is suitable for being extended to the quantification of the virome in a wide range of gut microbiota metagenomic sequencing data.
[0053] In some preferred embodiments, the construction of the marker protein database specific to virus clusters includes: constructing virus clusters in the gut; screening the marker proteins of the virus clusters, and constructing a marker protein database specific to virus clusters using the marker proteins.
[0054] In order to construct a comprehensive virus taxonomic unit at the genus level of the human gut virome, in this application, de novo clustering is performed based on the common protein characteristics among the collected virus genomes to obtain virus taxonomic units - virus clusters. In some preferred embodiments, constructing virus clusters in the gut includes: constructing protein clusters of gut virus genomes; de novo constructing virus clusters in the gut based on the characteristics of the common protein clusters among gut virus genomes.
[0055] In some preferred embodiments, constructing protein clusters of gut virus genomes includes: collecting the sequences of known gut virus genomes and performing quality control to obtain a set of gut virus genomes; identifying and translating the coding genes on the gut virus genomes to obtain protein sequences; clustering all virus proteins according to the similarity of the protein sequences, thereby constructing protein clusters of gut virus genomes.
[0056] Specifically, the steps for constructing the viral clusters of the gut include: 1) Collect the latest published human gut virus genome sequences, including GPD, MGV, CHVD, and GVD, etc. According to indicators such as the virus genome length and genome integrity, use the Checkv tool (see Nayfach, S. et al. CheckV assesses the quality and completeness of metagenome-assembled viral genomes. Nat. Biotechnol. 2020 395 39,578-585 (2020).) for quality control to obtain a set of high-quality human gut virus genomes. 2) Based on the Prodigal gene prediction tool (see Bin Jang, H. et al. Taxonomic assignment of uncultivated prokaryotic virus genomes is enabled by gene-sharing networks. Nat. Biotechnol. 37, 632-639 (2019), identify the coding genes on the human gut virus genomes and translate them into protein sequences. 3) Based on the vContact2 tool (see Bin Jang, H. et al. Taxonomic assignment of uncultivated prokaryotic virus genomes is enabled by gene-sharing networks. Nat. Biotechnol. 37, 632-639 (2019), according to the similarity of protein sequences (obtained by protein sequence alignment), cluster all viral proteins to construct protein clusters (PCs) of human gut virus genomes. 3) Based on the characteristics of the common protein clusters between virus genomes, use network clustering methods such as Markov clustering to de novo construct the gut taxonomic unit at the genus level - viral clusters (VCs).
[0057] The following is an exemplary description of the steps of "screening the marker proteins of viral clusters and constructing a marker protein database specific to viral clusters using the marker proteins":
[0058] In some preferred embodiments, screening for marker proteins of enterovirus clusters and constructing a virus cluster-specific marker protein database using the marker proteins includes: screening for protein clusters associated with each virus cluster based on the information of the shared genomes between the protein clusters and the virus clusters, thereby establishing an association relationship between the virus clusters and the protein clusters; calculating the Jaccard coefficient between the protein clusters associated with the virus clusters and the virus clusters based on the occurrence of the protein clusters associated with the virus clusters in each genome of the virus clusters, and screening for protein clusters that simultaneously meet the following two conditions, denoted as candidate marker protein clusters: 1) the Jaccard coefficient between the protein cluster and the virus cluster (i.e., Jaccard index, also known as Jaccard similarity coefficient - Jaccard similarity coefficient, used to compare the similarity and difference between finite sample sets. The larger the Jaccard coefficient value, the higher the sample similarity) ≥ a first threshold (preferably 0.5); 2) the Jaccard coefficient between the protein cluster and other virus clusters < a second threshold (preferably 0.2); selecting a representative protein sequence for the candidate marker protein cluster, and performing homology quality control and determination of the number of marker proteins on the representative protein sequence; selecting the determined number of marker proteins in descending order of the Jaccard coefficient between the protein cluster and the virus cluster, thereby constructing a virus cluster-specific marker protein database.
[0059] The following is an exemplary description of "selecting a representative protein sequence for the candidate marker protein cluster and performing homology quality control and determination of the number of marker proteins":
[0060] (For example, the Diamond tool can be used, see Buchfink, B., Xie, C. & Huson, D. H. Fast and sensitive protein alignment using DIAMOND. Nat. Methods 12, 59 - 60 (2015)) to perform pairwise alignment between the internal protein sequences of the candidate marker protein clusters, and screen for the protein sequence with the most shared genomes and the highest homology with other protein sequences (i.e., protein sequences outside the candidate marker protein clusters) as the representative sequence of a certain candidate marker protein cluster, constructing a set of representative protein sequences of all candidate marker protein clusters; removing the representative protein sequences with homologous sequences to other microorganisms outside the virus cluster from the set of representative protein sequences of the candidate marker protein clusters; taking the minimum value of the number of remaining representative protein sequences and 20% of the number of protein sequences on the virus genomes contained in the virus cluster as the target number of marker proteins; to reduce the marker protein database, preferably, control the number of marker proteins for each virus cluster ≤ a number threshold, preferably 50.
[0061] The steps of "removing representative protein sequences with homologous sequences to other microorganisms outside the virus cluster" are exemplified as follows for operability:
[0062] Using the Diamond tool, pairwise alignment between representative protein sequences is completed. Meanwhile, Blastx (see Gish, W. & States, D. J. Identification of protein coding regions by database similarity search. Nat. Genet. 1993 33 3, 266 - 272 (1993)) is used to achieve pairwise alignment between representative protein sequences and the virus genome. Further, from the above two groups of alignment results, homologous regions (for example, ≥30 amino acids) with representative protein sequences or whole genome sequences of other virus clusters are screened, and then the Bedtools maskfasta tool (see Quinlan, A. R. & Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841 - 842 (2010)) is used to mask this homologous region. In addition, using the hmmsearch tool (see Li, H. et al. The Sequence Alignment / Map format and SAMtools. Bioinformatics 25, 2078 - 2079 (2009)), according to the hidden Markov models of bacteria and archaea, representative protein sequences with homologous sequences to prokaryotes in the remaining representative protein sequences are searched and removed.
[0063] In some preferred embodiments, aligning the intestinal metagenomic sequencing data with the marker protein database to achieve quantification of the intestinal virome includes: aligning the reads in the intestinal microbiota metagenomic sequencing data to the marker protein database; taking the virus cluster as a unit (for example, the alignment tool Samtools can be used), counting the number of reads aligned to each marker protein and dividing by the length of the marker protein to calculate the read coverage depth of each marker protein, and then obtaining the read coverage depth of the virus cluster; excluding false positive virus clusters in the sequencing data, and marking the virus clusters with the proportion of the number of marker proteins with read coverage in all marker proteins greater than or equal to the third threshold (for example, more than 50%) as identified virus clusters; normalizing the read coverage depths of each virus cluster in all identified virus clusters to obtain the final relative abundances of the virus clusters contained in the test sample.
[0064] In some preferred embodiments, for the marker proteins of each viral cluster, a predetermined number (e.g., 10% each) of the marker proteins with the highest and lowest coverage depths are removed, and the mean of the coverage depths of the remaining intermediate number (e.g., 80%) of the marker proteins is used as the read coverage depth of the viral cluster.
[0065] The above alignment with the marker protein database enables rapid quantification of the virome. By excluding false-positive viral clusters, the quantification accuracy is further improved. Compared with the whole-genome sequence, using the viral cluster-specific marker protein sequences as the database significantly compresses the scale of the database sequences, improving the operation rate and the quantification efficiency of the virome.
[0066] Example 2
[0067] This example provides a more preferred method for quantifying the human gut virome. Among them, Figure 2 The detailed flowchart of the method of this example is shown:
[0068] 1. Construct viral taxonomic units at the genus level based on common protein features
[0069] To construct a comprehensive viral taxonomic unit at the genus level of the human gut virome, this example performs de novo clustering based on the common protein features among the collected viral genomes. The specific steps include: 1) Collect the latest published human gut virus genome sequences, including GPD, MGV, CHVD, and GVD, etc. According to indicators such as the viral genome length and genome integrity, use the Checkv14 tool for quality control to obtain a high-quality set of human gut virus genomes. 2) Based on the Prodigal15 gene prediction tool, identify the coding genes on the human gut virus genomes and translate them into protein sequences. 3) Based on the vContact2 tool15, according to the similarity of the protein sequences, cluster all viral proteins to construct protein clusters (PCs) of the human gut virus genomes. Further, based on the common protein cluster features among the viral genomes, use network clustering methods such as Markov clustering to de novo construct the gut taxonomic unit at the genus level - viral clusters (VCs).
[0070] 2. Screen marker proteins of the gut viral clusters and construct a marker protein database
[0071] To screen marker proteins for the gut viral clusters constructed above, this example adopts the following steps:
[0072] 1) Establish virus cluster - protein cluster associations: Based on the information of the common genomes between the protein clusters and virus clusters constructed above, screen the protein clusters related to each virus cluster; 2) Screen candidate protein clusters: Calculate the Jaccard Index between the protein clusters related to the virus clusters and the virus clusters according to the occurrence of the protein clusters in each genome of the virus clusters, and screen the candidate marker protein clusters. The candidate marker protein clusters need to meet two conditions simultaneously: one is that the Jaccard Index between the protein cluster and the virus cluster is ≥ 0.5; the other is that the Jaccard Index between this protein cluster and other virus clusters is < 0.2. 3) Select representative protein sequences for the candidate protein clusters: Use Diamond 16 tool to complete pairwise alignments between the internal protein sequences of the candidate marker protein clusters, and screen the protein with the closest distance to other protein sequences as the representative sequence of this candidate protein cluster, and construct a set of representative protein sequences for the candidate protein clusters; 4) Quality control of candidate representative protein sequences, remove representative proteins with homologous sequences to other microorganisms: Use Diamond tool to complete pairwise alignments between the representative protein sequences. At the same time, use Blastx 17 to achieve pairwise alignments between the representative protein sequences and the virus genomes. Further, from the above two sets of alignment results, screen homologous regions (≥ 30 amino acids) with representative proteins or whole genome sequences of other virus clusters, and further use Bedtools maskfasta 18 tool to mask this homologous region. In addition, use hmmsearch 19 tool to search and remove representative protein sequences with homologous sequences to prokaryotes in the remaining representative protein sequences according to the hidden Markov models of bacteria and archaea. 5) Reduce the representative protein database: Take the minimum value of the number of remaining representative protein sequences and 20% of the number of proteins on the virus genomes included in the virus cluster as the target number, and reduce the marker protein database. The maximum number of marker proteins for each virus cluster does not exceed 50. After determining the number of marker proteins, select marker proteins from high to low according to the Jaccard Index between the virus cluster and the protein cluster, and construct a marker protein database for virus clusters at the genus level.
[0073] 3. Calculate the relative abundances of viromes at the genus level based on alignments with the marker protein database
[0074] To achieve the quantification of enterovirus clusters based on the marker protein database, the following steps were adopted in this embodiment: 1) Based on the Diamond protein-level alignment tool (Diamond, a software with similar functions to BLAST, which can align protein sequences or their corresponding nucleotides with a protein database and has the following characteristics: 500 to 20,000 times faster than BLAST; frameshift alignment for long sequences; low resource consumption and can run on ordinary desktops and laptops; diverse output formats), the human gut microbiota metagenomic sequencing reads were aligned to the above-constructed marker protein database. 2) Tools such as Samtools were used to count the number of reads aligned to each marker protein and divide it by the length of the marker protein to calculate the read coverage depth of each marker protein. To further reduce the difference in read coverage depth on different marker proteins of the same virus cluster, for the marker proteins of each virus cluster, the top 10% and bottom 10% with the highest coverage depth ranking were removed, and the average value of the coverage depth of the middle 80% of the marker proteins was used as the read coverage depth of the virus cluster. 3) Controlling false-positive virus clusters: To further reduce the false positives caused by random alignment, only when the number of marker proteins with read coverage among all the marker proteins of a virus cluster accounts for more than 50%, the virus cluster is considered to be identified. 4) Among all the identified virus clusters, the read coverage depth of the virus clusters was normalized as the final relative abundance of the virus clusters contained in the sample.
[0075] 4. Evaluate the efficacy of the quantification method
[0076] In this embodiment, the efficacy of the above-mentioned quantification method for the enterovirus group was evaluated from dimensions such as the accuracy of quantification and the efficiency of operation. The specific steps are as follows:
[0077] 1) Refer to the quantification scheme: Select the strategy based on the whole-genome database alignment as the reference for the horizontal evaluation of efficacy. The whole-genome database uses the 13XXX genomes of the constructed marker virus clusters to build a reference database. The bowtie2 tool is used to align the reads to the reference database, and tools such as Samtools 19 are used to count the number of reads aligned to each genome, and further calculate the relative abundance of each virus group using Reads Per Kilobase Million (RPKM):
[0078]
[0079] 2) Simulated datasets: Ten metagenomic datasets were simulated. Each dataset contained 100 randomly selected viral genomes, and the genomic sequences of their host bacteria were added as background sequence interference, and simulated into 150bp reads. The total number of reads in the simulated datasets was 50M, of which the number of reads of viral genomes accounted for 10%, and the number of reads of bacteria accounted for 90%.
[0080] 3) Evaluation metrics: When evaluating the accuracy of viral abundance quantification, recall, precision, F1 score, and the Bray-Curtis distance between the simulated abundance and the measured abundance were used as measurement metrics to evaluate the gap between the quantification results and the true values, so as to further measure the accuracy of the quantification tool. When evaluating the quantification efficiency, the efficiency of the quantification scheme in terms of memory occupancy and running time was evaluated.
[0081] The viral quantification method provided by this embodiment of the present invention: By constructing a marker protein database specific to viral clusters and aligning the reads to this marker protein database, viral metagenomic quantification of intestinal metagenomic data is achieved, and the following beneficial effects can be obtained:
[0082] 1) The marker protein database excludes the interference of closely related sequences of other microorganisms and improves the accuracy of viral quantification.
[0083] According to the steps in step 4 above for evaluating the efficacy of the quantification method, the consistency between the quantification abundance obtained by the quantification method of the present application and the theoretical abundance was evaluated.
[0084] The results are as Figure 3A shown. In the constructed test dataset, the viral metagenomic abundance quantified in this embodiment can explain 85.7% of the theoretical abundance, which is higher than the quantification result (85.1%) of the reference whole-genome alignment. Further statistics on the alignment of reads from bacteria and reads from viruses showed that the results are as Figure 3B shown. The protein sequence alignment-based scheme provided by this embodiment can significantly reduce the proportion of reads from bacteria misaligned to viruses, confirming the quantification accuracy of this embodiment.
[0085] 2) Using the marker protein sequences of viral clusters as the database, the scale of the database sequences is greatly compressed, and the running efficiency of the tool is greatly improved.
[0086] As shown in Table 1, the quantification method provided by this embodiment can generally complete the viral metagenomic quantification of 50M reads (150bp) in about 100 minutes. The time required for quantification is only 1 / 8 of the whole-genome quantification scheme and occupies less memory, enabling efficient viral metagenomic quantification of large-scale metagenomic samples.
[0087] Table 1 Computational resources for virome quantification
[0088]
[0089] Example 3: Using the publicly available simulated dataset to verify the quantification accuracy of the quantification method of the present application in the publicly available data.
[0090] Situation of the publicly available simulated dataset (see Kleiner, M., Hooper, L.V. & Duerkop, B.A. Evaluation of methods to purify virus-like particles for metagenomic sequencing of intestinal viromes. BMC Genomics 16, 1-15 (2015)): Six viruses (P22( 19585-B1 TM ), T7( BAA-1025-B2 TM ), T3( BAA-1025-B1 TM , ΦVPE25, M13 (New England Biolabs, Ipswich, MA) and Φ6), and two bacteria (Listeria monocytogenes EGD-e and Bacteroides thetaiotaomicron VPI5482) were added to the feces of germ-free mice. At the same time, a sequencing method for enriching virus particles (by enriching intestinal virus particles, extracting viral nucleic acids, amplifying and purifying them, and finally performing on-machine sequencing) and a metagenomic sequencing method were used to obtain sequencing data.
[0091] Operation process: Input the above sequencing data into the software, and the software outputs the corresponding taxonomic virus clusters at the genus level of the virome and the abundances of each virus cluster in the sample (i.e., quantitative analysis is performed using the quantification method of the present application). At the same time, a quantification method based on whole-genome alignment (i.e., the existing quantification method) is used as a reference.
[0092] The results are as Figure 4 shown. Both schemes can identify 5 / 6 of the virus species in the simulated microbial community. However, whether it is the sequencing data of virus particle enrichment or the metagenomic sequencing data: for the virus species (dark blue) additionally identified by the whole-genome alignment-based strategy, their abundances account for a very high proportion (>50%), while the abundance proportions of other viruses identified by the strategy of the present invention based on marker protein alignment are lower (<25%), reflecting the quantification accuracy of the present invention.
[0093] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0094] Example 4
[0095] Figure 5 It is a device for quantifying the gut virome according to an embodiment of the present application. The device includes: a construction unit 10 and a comparison and quantification unit 20.
[0096] The construction unit 10 is configured to construct a marker protein database specific to virus clusters; the comparison and quantification unit 20 is configured to compare the gut metagenomic sequencing data with the marker protein database to achieve the quantification of the gut virome.
[0097] Optionally, the construction unit includes: a virus cluster construction module configured to construct virus clusters in the gut; a marker protein database construction module configured to screen the marker proteins of the virus clusters and construct a marker protein database specific to virus clusters.
[0098] Optionally, the virus cluster construction module includes: a protein cluster construction sub-module configured to construct protein clusters of gut virus genomes; a virus cluster construction sub-module configured to de novo construct virus clusters in the gut based on the characteristics of the common protein clusters among gut virus genomes.
[0099] Optionally, the protein cluster construction sub-module includes: a collection module configured to collect the sequences of known gut virus genomes and perform quality control to obtain a set of gut virus genomes; an identification and translation module configured to identify and translate the coding genes on the gut virus genomes to obtain protein sequences; a protein clustering module configured to cluster all virus proteins according to the similarity between protein sequences, thereby constructing protein clusters of gut virus genomes.
[0100] Optionally, the labeled protein database construction module includes: a virus cluster - protein cluster association module, which is configured to screen protein clusters related to each virus cluster according to the information of the shared genomes between the protein clusters and the virus clusters, so as to establish an association relationship between the virus clusters and the protein clusters; a candidate labeled protein cluster screening module, which is configured to calculate the Jaccard coefficient between the protein cluster and the virus cluster according to the occurrence of the protein clusters related to the virus cluster in each genome of the virus cluster, and screen protein clusters that simultaneously meet the following two conditions, denoted as candidate labeled protein clusters: 1) the Jaccard coefficient between the protein cluster and the virus cluster ≥ 0.5; 2) the Jaccard coefficient between the protein cluster and other virus clusters < 0.2; a quality control and quantity determination module, which is configured to select a representative protein sequence for the candidate labeled protein clusters, and perform homology quality control and labeled protein quantity determination on the representative protein sequence; a labeled protein database construction sub - module, which is configured to select the determined number of labeled proteins according to the Jaccard coefficient between the protein cluster and the virus cluster from high to low, so as to construct a virus - cluster - specific labeled protein database.
[0101] Optionally, the quality control and quantity determination module includes: a representative protein sequence screening module, which is configured to perform pairwise alignment between the internal protein sequences of the candidate labeled protein clusters, and screen the protein sequence with the closest distance to other protein sequences as the representative protein sequence of a certain candidate labeled protein cluster, and construct a set of representative protein sequences of all candidate labeled protein clusters; a quality control removal module, which is configured to remove the representative protein sequences with homologous sequences to other microorganisms outside the virus cluster from the set of representative protein sequences of the candidate labeled protein clusters; a quantity determination module, which is configured to take the minimum value of the number of remaining representative protein sequences and 20% of the number of protein sequences on the viral genomes contained in the virus cluster as the target quantity of the labeled proteins; preferably, in the quantity determination module, the number of labeled proteins for each virus cluster is controlled ≤ quantity threshold, preferably 50.
[0102] Optionally, the alignment and quantification unit includes: an alignment module, which is configured to align the reads in the intestinal microbiota metagenomic sequencing data to the labeled protein database; a read coverage depth calculation module, which is configured to take the virus cluster as a unit, count the number of reads aligned to each labeled protein and divide it by the length of the labeled protein, calculate the read coverage depth of each labeled protein, and further obtain the read coverage depth of the virus cluster; a false positive exclusion module, which is configured to exclude false positive virus clusters in the sequencing data, and denote the virus clusters in which the number of labeled proteins with read coverage accounts for more than 50% of all labeled proteins as identified virus clusters; a quantification module, which is configured to normalize the read coverage depths of each virus cluster in all identified virus clusters to obtain the final relative abundances of the virus clusters contained in the test sample.
[0103] The device for determining the formula of the above-mentioned intestinal probiotic supplement includes a processor and a memory. The above-mentioned construction unit 10, comparison and quantification unit 20, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions.
[0104] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, effective analysis of the intestinal microbial flora and the state of the flora can be carried out.
[0105] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0106] An embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the method for quantifying the above-mentioned intestinal virus is implemented.
[0107] An embodiment of the present application provides a processor, and the processor is used to run a program, wherein when the program runs, the method for quantifying the above-mentioned intestinal virus is executed.
[0108] An embodiment of the present application provides a device, the device includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, the following steps are implemented: constructing a marker protein database specific to virus clusters; comparing the intestinal metagenomic sequencing data with the marker protein database to achieve quantification of the intestinal virome.
[0109] The device herein can be a server, PC, PAD, mobile phone, etc.
[0110] The present application also provides a computer program product, which is suitable for executing a program initialized with the following method steps when executed on a data processing device: constructing a marker protein database specific to virus clusters; comparing the intestinal metagenomic sequencing data with the marker protein database to achieve quantification of the intestinal virome.
[0111] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0112] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows and / or blocks. Figure 1 in one or more flows and / or blocks Figure 1 in one or more blocks.
[0113] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more flows and / or blocks. Figure 1 in one or more flows and / or blocks Figure 1 in one or more blocks.
[0114] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or blocks. Figure 1 in one or more flows and / or blocks Figure 1 in one or more blocks.
[0115] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0116] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0117] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0118] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0119] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system or computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0120] From the above description, it can be seen that the above embodiments of the present invention achieve the following technical effects:
[0121] The quantitative method and device of the present invention construct a marker protein database specific to virus clusters. This marker protein database excludes the interference of closely related sequences of other microorganisms and improves the accuracy of virus quantification. Further, by aligning reads to this marker protein database, virus metagenome quantification of intestinal metagenomic data is achieved. The solution of the present application greatly compresses the scale of the database sequences and significantly improves the running efficiency of the tool.
[0122] At present, the solutions of the method and device of the present application achieve virus quantification at the genus level based on the "virus cluster - protein" strategy. Subsequently, the database construction strategy can be extended to the species level to achieve rapid quantification of virus species. In addition, the solutions of the method and device of the present application can accurately and efficiently generate the abundance spectrum of the gut virome, and can be applied to large-scale gut virome scientific research services; in the future, it may be applied to fields such as virus-related health detection or clinical tests.
[0123] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for quantifying the gut virome, characterized in that, The method includes: Constructing a marker protein database specific to virus clusters; Comparing the gut metagenomic sequencing data with the marker protein database to achieve quantification of the gut virome.
2. The method according to claim 1, wherein Constructing a marker protein database specific to virus clusters includes: Constructing virus clusters in the gut; Screening marker proteins of the virus clusters to construct the marker protein database specific to the virus clusters.
3. The method according to claim 2, wherein Constructing virus clusters in the gut includes: Constructing protein clusters of gut virus genomes; De novo constructing the virus clusters in the gut based on the characteristics of the common protein clusters among the gut virus genomes.
4. The method according to claim 3, wherein Constructing protein clusters of gut virus genomes includes: Collecting sequences of known gut virus genomes and performing quality control to obtain a set of gut virus genomes; Identifying and translating the coding genes on the gut virus genomes to obtain protein sequences; Clustering all virus proteins according to the similarity among the protein sequences to construct the protein clusters of the gut virus genomes.
5. The method according to claim 3, wherein Screening marker proteins of the virus clusters and constructing the marker protein database specific to the virus clusters using the marker proteins includes: According to the information of the common genomes between the protein clusters and the virus clusters, screening the protein clusters related to each virus cluster to establish the association relationship between the virus clusters and the protein clusters; Calculating the Jaccard coefficient between the protein cluster and the virus cluster according to the occurrence of the protein cluster related to the virus cluster in each genome of the virus cluster, and screening the protein clusters that meet the following two conditions simultaneously, denoted as candidate marker protein clusters: 1) The Jaccard coefficient between the protein cluster and the virus cluster ≥ 0.5; 2) The Jaccard coefficient between the protein cluster and other virus clusters < 0.2; Selecting representative protein sequences for the candidate marker protein clusters and performing homology quality control and determination of the number of marker proteins on the representative protein sequences; Selecting the determined number of marker proteins in descending order of the Jaccard coefficient between the protein cluster and the virus cluster to construct the marker protein database specific to the virus clusters.
6. The method according to claim 5, wherein Selecting representative protein sequences for the candidate marker protein clusters and performing homology quality control and determination of the number of marker proteins on the representative protein sequences includes: Performing pairwise comparison between the internal protein sequences of the candidate marker protein clusters and screening the protein sequence with the closest distance to other protein sequences as the representative sequence of a certain candidate marker protein cluster to construct a set of representative protein sequences of all candidate marker protein clusters; Removing the representative protein sequences with homologous sequences to other microorganisms outside the virus cluster from the set of representative protein sequences of the candidate marker protein clusters; Taking the minimum value of the number of the remaining representative protein sequences and 20% of the number of protein sequences on the virus genomes contained in the virus cluster as the target number of marker proteins; Preferably, controlling the number of marker proteins of each virus cluster ≤ the number threshold, preferably 50.
7. The method according to any one of claims 1 to 6, characterized in that Comparing the gut metagenomic sequencing data with the marker protein database to achieve quantification of the gut virome includes: Align the reads in the intestinal microbiota metagenomic sequencing data to the marker protein database; Taking virus clusters as units, count the number of reads aligned to each marker protein and divide it by the length of the marker protein to calculate the read coverage depth of each marker protein, and then obtain the read coverage depth of the virus cluster; Exclude the false-positive virus clusters in the sequencing data, and mark the virus clusters in which the number of marker proteins with read coverage accounts for more than 50% of all marker proteins as identified virus clusters; Normalize the read coverage depths of the virus clusters in all the identified virus clusters to obtain the final relative abundances of the virus clusters contained in the test sample.
8. The method according to claim 7, wherein For the marker proteins of each virus cluster, remove a predetermined number of the marker proteins with the highest and lowest coverage depths, and use the mean value of the coverage depths of the remaining marker proteins in the middle as the read coverage depth of the virus cluster.
9. An apparatus for quantifying the gut virome, characterized in that, The device includes: A construction unit configured to construct a marker protein database specific to virus clusters; An alignment and quantification unit configured to align the intestinal metagenomic sequencing data with the marker protein database to achieve quantification of the intestinal virome.
10. The device according to claim 9, characterized in that, The construction unit includes: A virus cluster construction module configured to construct virus clusters in the intestine; A marker protein database construction module configured to screen the marker proteins of the virus clusters and construct the marker protein database specific to the virus clusters.
11. The device according to claim 10, wherein The virus cluster construction module includes: A protein cluster construction sub-module configured to construct protein clusters of intestinal virus genomes; A virus cluster construction sub-module configured to de novo construct the virus clusters in the intestine based on the characteristics of the common protein clusters among the intestinal virus genomes.
12. The device according to claim 11, wherein The protein cluster construction sub-module includes: A collection module configured to collect the sequences of known intestinal virus genomes and perform quality control to obtain a set of intestinal virus genomes; An identification and translation module configured to identify and translate the coding genes on the intestinal virus genomes to obtain protein sequences; A protein clustering module configured to cluster all virus proteins according to the similarity between the protein sequences, thereby constructing the protein clusters of the intestinal virus genomes.
13. The device according to claim 11, characterized in that, The marker protein database construction module includes: A virus cluster-protein cluster association module configured to screen the protein clusters related to each virus cluster according to the information of the common genomes between the protein clusters and the virus clusters, thereby establishing the association relationship between the virus clusters and the protein clusters; A candidate marker protein cluster screening module configured to calculate the Jaccard coefficient between the protein cluster and the virus cluster according to the occurrence of the protein clusters related to the virus cluster in each genome of the virus cluster, and screen the protein clusters that simultaneously meet the following two conditions and mark them as candidate marker protein clusters: 1) The Jaccard coefficient between the protein cluster and the virus cluster is ≥ 0.5; 2) The Jaccard coefficient between the protein cluster and other virus clusters is < 0.2; A quality control and quantity determination module, which is configured to select representative protein sequences for the candidate marker protein clusters, and perform homology quality control and marker protein quantity determination on the representative protein sequences; A marker protein database construction sub-module, which is configured to select the determined number of marker proteins from high to low according to the Jaccard coefficient between the protein clusters and the virus clusters, so as to construct the marker protein database specific to the virus clusters.
14. The device according to claim 13, wherein The quality control and quantity determination module includes: A representative protein sequence screening module, which is configured to perform pairwise comparison between the internal protein sequences of the candidate marker protein clusters, and screen the protein sequence with the closest distance to other protein sequences as the representative protein sequence of a certain candidate marker protein cluster, and construct a set of representative protein sequences of all the candidate marker protein clusters; A quality control removal module, which is configured to remove the representative protein sequences with homologous sequences to other microorganisms outside the virus clusters from the set of representative protein sequences of the candidate marker protein clusters; A quantity determination module, which is configured to use the minimum value of the number of the remaining representative protein sequences and 20% of the number of protein sequences on the viral genomes contained in the virus clusters as the target quantity of the marker proteins; Preferably, in the quantity determination module, the number of marker proteins for each virus cluster is controlled ≤ a quantity threshold, preferably 50.
15. The device according to any one of claims 9 to 14, characterized in that The alignment and quantification unit includes: An alignment module, which is configured to align the reads in the gut microbiota metagenomic sequencing data to the marker protein database; A read coverage depth calculation module, which is configured to take the virus clusters as units, count the number of reads aligned to each marker protein and divide it by the length of the marker protein to calculate the read coverage depth of each marker protein, and further obtain the read coverage depth of the virus clusters; A false positive exclusion module, which is configured to exclude false positive virus clusters in the sequencing data, and mark the virus clusters in which the number of marker proteins with read coverage accounts for more than 50% of all marker proteins as identified virus clusters; A quantification module, which is configured to normalize the read coverage depths of the virus clusters in all the identified virus clusters to obtain the final relative abundances of the virus clusters contained in the test sample.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the method for quantifying the gut virome according to any one of claims 1 to 8.
17. A processor, characterized in that, The processor is used to run the program, wherein when the program runs, it executes the method for quantifying the gut virome according to any one of claims 1 to 8.