A high-throughput identification and quantification method for environmental pathogens based on virulence genes
Through high-throughput identification and quantitative methods based on virulence genes, combined with metagenomic sequencing and virulence gene comparison recognition tools, the accuracy and efficiency of pathogen identification and quantification in environmental samples in the prior art are solved, and more accurate pathogen identification and quantification are achieved, with high application value.
Patent Information
- Application Number
- CN202410208743.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-26
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-02-26
AI Technical Summary
The prior art has not yet provided a method that can accurately and efficiently identify and quantify pathogens in environmental samples, especially in environments such as soil and water, where identification bias and inefficiency are problems.
High-throughput recognition and quantitative methods based on virulence genes were used to extract the total DNA of environmental samples for second-generation or third-generation metagenomic sequencing. Combined with species recognition tools and virulence gene alignment recognition tools, species carrying core virulence genes were screened out, and the ratio R was calculated to determine whether they were pathogenic bacteria.
It achieves more accurate identification and quantification of pathogens in environmental samples, improves identification efficiency, and has high application value in environmental pathogen risk assessment and transmission control.
Smart Images

Figure CN118086518B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of metagenomic analysis methods, and in particular relates to a high-throughput identification and quantification method of environmental pathogens based on virulence genes. Background Art
[0002] With the development of sequencing technology, metagenomic methods have been greatly developed in microbial detection. The general process of metagenomic analysis includes: experimental design, sample information organization, metagenomic sequencing, host removal, reads-based analysis or assemble-based analysis. Among the reads-based species composition analysis tools, the Kraken tool generally shows more favorable speed, memory usage and accuracy compared with other similar tools. The MetaPhlAn2 tool uses parts that can distinguish taxa to reduce the memory footprint of the index, but it prevents the tool from classifying many reads. Therefore, the Kraken tool is a more suitable pathogen identification tool.
[0003] At present, methods for pathogen analysis have been reported, but no pathogen identification and analysis methods specifically for environmental samples have been found. The Chinese invention patent with publication number CN110964854A discloses a kit for the simultaneous detection of seven respiratory pathogen nucleic acids and its application; the Chinese invention patent with publication number CN111074001A discloses 9 respiratory pathogen nucleic acids coordinated multiplex PCR detection. These kits or methods can only be used for specific detection of pathogens. The Chinese invention patent with publication number CN115852001A discloses a wheat pathogen detection method and application, which can comprehensively detect all pathogens that may be contained in the wheat sample to be tested. It adopts the form of comparing the pathogen database to determine the pathogen, and there is a certain deviation in the analysis of pathogens. Therefore, it is urgent to provide a more accurate and efficient high-throughput identification and quantification method for pathogens. Summary of the invention
[0004] The purpose of the present invention is to solve the deficiencies in the prior art and provide a high-throughput identification and quantification method for environmental pathogens based on virulence genes.
[0005] The specific technical solutions adopted by the present invention are as follows:
[0006] The present invention provides a high-throughput identification and quantification method for pathogens based on virulence genes, and the specific steps are as follows:
[0007] S1: Extract total DNA from target environmental samples and perform second-generation or third-generation metagenomic sequencing, perform quality control and filtering on sequencing data, and obtain clean reads;
[0008] S2: Use a species identification tool to perform species identification on the clean reads in step S1 to obtain a first species set; each microbial species information in the first species set includes a species classification number, a species name, and a number of reads; use a metagenomic assembly software to splice the clean reads in step S1 to form an overlapping group file; then re-input the overlapping group file into the species identification tool for species identification to obtain a second species set; each microbial species information in the second species set includes an overlapping group number, a species classification number, a species name, and a number of reads;
[0009] S3: Based on the species classification number and the species name, the first species set and the second species set are matched, and the microbial species that exist in both the first species set and the second species set form a third species set; each microbial species information in the third species set includes an overlapping group number, a species classification number, a species name, and a read number;
[0010] S4: using a virulence gene comparison and identification tool to identify virulence genes of all the contigs in the contig file obtained in step S2, thereby obtaining a virulence gene set; each virulence gene information in the virulence gene set includes a virulence gene name and its corresponding contig number;
[0011] S5: Based on the contig number, the third species set obtained in step S3 and the virulence gene set obtained in step S4 are matched to screen out species with virulence genes in the third species set, count the total number of reads and the number of virulence genes contained therein, and calculate the ratio R; if the species contains the core virulence genes in the core virulence gene set, and the ratio R is greater than or equal to the threshold, it is judged to be a pathogen;
[0012]
[0013] S6: The relative abundance of pathogens is obtained by the ratio of the total number of reads of the pathogen to the total number of reads of the sample.
[0014] Preferably, the species identification tool is Kraken2 software or centrifuge software.
[0015] Preferably, the virulence gene identification uses blast software and VFDB database.
[0016] Preferably, the metagenomic assembly software adopts megahit software or Flye software.
[0017] Preferably, the matching of the first species set and the second species set is implemented by writing a script in R language.
[0018] Preferably, the calculation of R is implemented by writing a script in R language.
[0019] Preferably, the relative abundance of the pathogen is obtained according to the ratio of the number of reads of the pathogen to the total number of reads of the environmental sample.
[0020] Preferably, the number of virulence genes in the pathogen is obtained by matching the overlapping group number and the species name in the third species set with the virulence gene set.
[0021] Preferably, the threshold of R is 3.64E-06, which is the minimum R value of 113 species and 4140 strains of typical environmental pathogens; if R is greater than or equal to 3.64E-06 and the species contains core virulence genes, it is judged to be a pathogen.
[0022] Preferably, the method for determining the core virulence gene set is as follows:
[0023] The virulence factors in the pathogen virulence factor database VFDB and the proteins in the non-pathogenic bacteria protein database were compared for amino acid sequence using the BLAST RBH method; if a virulence factor has one or more orthologous proteins in non-pathogenic bacteria, the virulence factor is identified as a non-specific virulence factor that exists in both pathogens and non-pathogenic bacteria; if a virulence factor only exists in pathogens, it is a pathogen-specific virulence factor. If the proportion of specific virulence factors in a certain type of virulence factor is higher than that of non-specific virulence factors, all virulence genes of this type of virulence factor are judged as core virulence genes, and all core virulence genes constitute the core virulence gene set.
[0024] Compared with the prior art, the present invention has the following beneficial effects:
[0025] Based on the association analysis between virulence genes and pathogens, the present invention provides a more accurate method for identifying and quantifying pathogens in environmental samples, targeting zoonotic pathogens that are widely distributed but have low abundance in environmental samples such as soil and water. It has high application value in environmental pathogen risk assessment and transmission control. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is the implementation process of the high-throughput identification and quantification method of pathogens based on virulence genes provided by the present invention;
[0027] Figure 2 It is the identification and quantification results of pathogens in a soil sample based on three analytical methods: (a) the types of pathogens detected by the three methods; (b) the relative abundance of pathogens detected by the three methods, among which: "MBPD" is a pathogen identification method based on 16S rRNA gene; "species comparison" is a method for determining pathogens based on metagenomic data and species comparison results; "this method" is a high-throughput identification and quantification method for environmental pathogens based on virulence genes provided in this embodiment. DETAILED DESCRIPTION
[0028] The present invention is further described and illustrated below in conjunction with the accompanying drawings and specific embodiments. The technical features of each embodiment of the present invention can be combined accordingly without conflicting with each other.
[0029] Example 1
[0030] (1) Weigh 0.25 g of exposed topsoil (0-20 cm) around a domestic waste collection point, extract DNA from the environmental sample using the PowerSoil Prokit kit, and perform second-generation metagenomic sequencing (sequencing volume 10 GB); perform quality control and filtering on the second-generation metagenomic sequencing data to obtain clean reads.
[0031] (2) Using a species identification tool to perform species identification on the clean reads in step (1) to obtain a first species set. Each microbial species information of the first species set includes a species classification number, a species name, and a number of reads. In this embodiment, Kraken2 software is used as a species identification tool, and each sample obtains a *.report and a *.output file.
[0032] After the clean reads in step (1) are spliced using metagenomic assembly software, an overlapping group file is obtained; the overlapping group file is then input into the Kraken2 software, and species identification is performed again to obtain a second species set. In this embodiment, the megahit tool is used as the metagenomic assembly software to splice the clean reads to obtain several overlapping groups (contigs), and each sample obtains a *.fa result file. Each contig has a unique identifier, which is called the overlapping group number (contig_id).
[0033] (3) Based on the species classification number and species name, the first species set and the second species set are matched, and the third species set is composed of microbial species that exist in both the first species set and the second species set. In this embodiment, the integration and output are obtained by obtaining files A and B, the content of file A includes the species name (accurate to species), the number of reads, and the species classification number (taxid), and the content of file B is the species classification number (taxid) and the contig number (contig_id). The purpose of this method using reads and contigs data to run the kraken2 tool separately is to improve the accuracy of species identification. Only species that appear in both results are correctly identified, while ensuring that the virulence gene can correspond to the species.
[0034] In this embodiment, the matching of the first species set and the second species set in this step is implemented by writing a script in R language.
[0035] (4) Using the contig result file *.fa obtained in step (2) as input, the VFDB database and blast software are used to identify virulence genes, and a file C containing a virulence gene set is obtained. Each virulence gene information in the virulence gene set mainly includes the virulence gene name and its corresponding contig number.
[0036] (5) Summarize file A, file B, and file C, and match the third species set obtained in step (3) with the virulence gene set obtained in step (4) based on the overlapping group number, screen out species with virulence genes in the third species set, count the total number of reads and the number of virulence genes contained, and calculate the ratio R; if the species contains the core virulence genes in the core virulence gene set, and the ratio R is greater than or equal to the threshold, it is judged to be a pathogen.
[0037]
[0038] The threshold of R is 3.64E-06, which is the minimum R value of 113 species and 4140 strains of typical environmental pathogens. If R is greater than or equal to 3.64E-06 and the species contains core virulence genes, it is judged to be a pathogen.
[0039] Among them, the method for determining the core virulence gene set is as follows:
[0040] The virulence factors in the pathogen virulence factor database VFDB and the proteins in the non-pathogenic bacteria protein database were compared for amino acid sequence using the BLAST RBH method; if a virulence factor has one or more orthologous proteins in non-pathogenic bacteria, the virulence factor is identified as a non-specific virulence factor that exists in both pathogens and non-pathogenic bacteria; if a virulence factor only exists in pathogens, it is a pathogen-specific virulence factor. If the proportion of specific virulence factors in a certain type of virulence factor is higher than that of non-specific virulence factors, all virulence genes of this type of virulence factor are judged as core virulence genes, and all core virulence genes constitute the core virulence gene set.
[0041] In this embodiment, the non-pathogenic bacteria protein database uses the public data provided by https: / / gold.jgi.doe.gov / .
[0042] In this embodiment, the matching of the microbial set and the virulence gene set, and the calculation of the ratio R of the number of virulence genes to (the total number of read segments of the species×the average sequencing fragment length) are both implemented by writing a script in R language.
[0043] from Figure 2(a) It can be seen that this method can identify 23 pathogens. Although the other two methods (pathogen identification method based on 16S rRNA gene and method for determining pathogens based on metagenomic data and species comparison results) can identify more than 60 species, most of the pathogens in their results do not carry core virulence genes and may be false positives. Figure 2 (b) shows that while ensuring the accuracy of identification, the average relative abundance of pathogens obtained by this method is higher than that of the other two methods, and the distribution range is relatively wide.
[0044] The above-described embodiment is only an application case of the present invention, but it is not intended to limit the present invention. A person skilled in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution and effect obtained by equivalent replacement or equivalent transformation of the method described in the present invention shall fall within the protection scope of the present invention.
Claims
1. A high-throughput identification and quantification method for environmental pathogens based on virulence genes, characterized in that: The specific steps are as follows: S1: Extract total DNA from target environmental samples and perform second-generation or third-generation metagenomic sequencing, perform quality control and filtering on sequencing data, and obtain clean reads; S2: Use a species identification tool to perform species identification on the clean reads in step S1 to obtain a first species set; each microbial species information in the first species set includes a species classification number, a species name, and a number of reads; use a metagenomic assembly software to splice the clean reads in step S1 to form an overlapping group file; then re-input the overlapping group file into the species identification tool for species identification to obtain a second species set; each microbial species information in the second species set includes an overlapping group number, a species classification number, a species name, and a number of reads; S3: Based on the species classification number and the species name, the first species set and the second species set are matched, and the microbial species that exist in both the first species set and the second species set form a third species set; each microbial species information in the third species set includes an overlapping group number, a species classification number, a species name, and a read number; S4: using a virulence gene comparison and identification tool to identify virulence genes of all the contigs in the contig file obtained in step S2, thereby obtaining a virulence gene set; each virulence gene information in the virulence gene set includes a virulence gene name and its corresponding contig number; S5: Based on the contig number, the third species set obtained in step S3 and the virulence gene set obtained in step S4 are matched to screen out species with virulence genes in the third species set, count the total number of reads and the number of virulence genes contained therein, and calculate the ratio R; if the species contains the core virulence genes in the core virulence gene set, and the ratio R is greater than or equal to the threshold, it is judged to be a pathogen; S6: The relative abundance of pathogens is obtained by the ratio of the total number of reads of the pathogen to the total number of reads of the sample; The threshold of R is 3.64E-06, which is the minimum value of R value of 113 species and 4140 strains of typical environmental pathogens; if R is greater than or equal to 3.64E-06 and the species contains core virulence genes, it is judged as a pathogen; The method for determining the core virulence gene set is as follows: The virulence factors in the pathogen virulence factor database VFDB and the proteins in the non-pathogenic bacteria protein database were compared for amino acid sequences using the BLAST RBH method; if a virulence factor has one or more orthologous proteins in non-pathogenic bacteria, the virulence factor is identified as a non-specific virulence factor that exists in both pathogens and non-pathogenic bacteria; if a virulence factor only exists in pathogens, it is a pathogen-specific virulence factor; if the proportion of specific virulence factors in a certain type of virulence factors is higher than that of non-specific virulence factors, all virulence genes of this type of virulence factors are judged to be core virulence genes, and all core virulence genes constitute the core virulence gene set.
2. The high-throughput identification and quantification method of environmental pathogens based on virulence genes according to claim 1, characterized in that: The species identification tool is Kraken2 software or centrifuge software.
3. The high-throughput identification and quantification method of environmental pathogens based on virulence genes according to claim 1, characterized in that: The virulence gene identification was performed using blast software and VFDB database.
4. The high-throughput identification and quantification method of environmental pathogens based on virulence genes according to claim 1, characterized in that: The metagenomic assembly software adopts megahit software or Flye software.
5. The high-throughput identification and quantification method of environmental pathogens based on virulence genes according to claim 1, characterized in that: The matching of the first species set and the second species set is implemented by writing a script in R language.
6. The high-throughput identification and quantification method of environmental pathogens based on virulence genes according to claim 1, characterized in that: The R calculation is implemented by writing scripts in R language.
7. The high-throughput identification and quantification method of environmental pathogens based on virulence genes according to claim 1, characterized in that: The number of virulence genes in the pathogen is obtained by matching the overlapping group number and the species name in the third species set with the virulence gene set.
Citation Information
Patent Citations
Kit for simultaneously detecting nucleic acids of seven respiratory pathogens and application thereof
CN110964854A
Collaborative multiplex PCR detection method for nucleic acids of nine respiratory tract pathogens
CN111074001A
Wheat pathogenic bacterium detection method and application thereof
CN115852001A
Pathogenic microorganism virulence gene association model as well as establishment method and application thereof
CN112837745A
Metagenome-based clinical important pathogenic bacterium virulence gene detection method and system
CN113223618A