A pathogenic metagenomic data analysis method
By removing host sequences, filtering internal references and contaminating sequences, and combining rapid classification and accurate alignment methods, the problems of complex procedures and low accuracy in pathogen metagenomic analysis have been solved, achieving rapid and accurate pathogen species identification and efficient computation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU MATRIDX BIOTECH CO LTD
- Filing Date
- 2026-02-10
- Publication Date
- 2026-06-09
AI Technical Summary
Existing pathogen metagenomic analysis methods are complex, require extensive manual intervention, are slow, suffer from severe interference from host background DNA sequences, and have their accuracy affected by contaminating sequences.
By employing methods such as removing host sequences, filtering internal references and contaminating sequences, rapid classification, and precise alignment, combined with a pre-constructed species database and a reference genome database, candidate pathogen species are rapidly classified and screened, and high-precision alignment is performed to generate pathogen metagenomic analysis results.
It achieves rapid and accurate identification of pathogen species, significantly improves analysis speed and accuracy, reduces false positive rate, and has high computational performance and repeatability.
Smart Images

Figure CN122177226A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics, specifically the field of metagenomic analysis, and more specifically relates to a method for analyzing pathogen metagenomic data. Background Technology
[0002] With the development of high-throughput sequencing technology, metagenomic detection has become an important supplementary method for clinical pathogen detection. By sequencing all genetic material in a sample, this technology can unbiasedly identify potential pathogens such as bacteria, viruses, and fungi. However, existing metagenomic analysis methods for pathogens typically suffer from the following problems: complex analysis procedures, requiring significant manual intervention, and slow analysis speed; the large amount of host background DNA sequences in the sample severely interferes with the detection of pathogen sequences; and contamination sequences in experimental reagents or the environment may lead to false positive results, affecting the accuracy of the results. Therefore, existing technologies struggle to obtain high-precision pathogen classification results while ensuring analysis speed, necessitating an automated analysis method that can rapidly and accurately identify pathogen species from metagenomic sequencing data. Summary of the Invention
[0003] To address the problems of complex analysis procedures, significant host background interference, and the impact of contaminated sequences on accuracy in existing pathogen metagenomic data analysis methods, this invention provides a method for rapidly and accurately identifying pathogen species from sequencing data containing a large number of host sequences and complex background sequences.
[0004] The technical solution adopted in this invention is: a method for analyzing pathogen metagenomic data, comprising the following steps:
[0005] S1. Remove the host sequence from the original sequencing data to obtain the host-free sequence;
[0006] S2. Filter out the intrinsic parameter sequences and contaminating sequences in the host sequence to obtain the effective sequence;
[0007] S3. Based on a pre-constructed pathogen species database containing a list of species, the effective sequences are quickly classified to obtain species classification information; species in the species list with a sequence number less than the screening threshold are selected as candidate pathogen species, and their corresponding reads are extracted as candidate pathogen species sequences.
[0008] S4. Compare each candidate pathogen species obtained in step S3 with its corresponding species reference genome database to obtain species comparison information; the species comparison information includes at least the presence and abundance information of the candidate pathogen species in the sample.
[0009] S5. Output the metagenomic analysis results of the pathogen; the metagenomic analysis results of the pathogen shall include at least the species classification information obtained in step S3 and the species comparison information obtained in step S4.
[0010] Preferably, step S1 includes: performing preliminary classification on the raw sequencing data based on a pre-constructed host reference genome database and obtaining unclassified sequences according to a preset classification threshold; aligning the unclassified sequences to the host reference genome database and removing any possible host sequences to obtain host-free sequences.
[0011] Preferably, step S2 includes: aligning the host-de-host sequence to a pre-constructed database of internal references and contaminated sequences, filtering out reads that match the alignment, and obtaining a valid sequence.
[0012] Preferably, the parameters for comparison in step S2 include any one or a combination of seed length, maximum number of mismatches, optimal matching mode of the output comparison result, and sequences whose comparison quality is lower than a preset threshold.
[0013] Preferably, the species reference genome database in step S4 includes all sequences of the corresponding family and / or genus of the species.
[0014] As a preferred option, the comparison method in step S4 includes: for each candidate species, calling the corresponding family or genus sequence in the species reference genome database for parallel comparison.
[0015] Preferably, the pathogen metagenomic analysis results in step S5 include at least species ID, species name, species level, number of species sequences, cumulative number of species sequences, and species abundance information.
[0016] This invention also provides a pathogen metagenomic data analysis system, comprising:
[0017] The host sequence removal module is used to remove the host sequence from the raw sequencing data to obtain the host-free sequence;
[0018] The module for filtering internal parameters and contaminated sequences is used to filter out internal parameter sequences and contaminated sequences in the host sequence to obtain a valid sequence.
[0019] The rapid classification and analysis module is used to quickly classify effective sequences based on a pre-built pathogen species database containing a list of species to obtain species classification information; species in the species list with a number of sequences less than the screening threshold are selected as candidate pathogen species, and their corresponding reads are extracted as candidate pathogen species sequences.
[0020] The precise alignment analysis module is used to align each candidate pathogen species obtained from the rapid classification analysis module with its corresponding species reference genome database to obtain species alignment information; the species alignment information includes at least the presence and abundance information of the candidate pathogen species in the sample.
[0021] The report generation module is used to output the metagenomic analysis results of pathogens; the metagenomic analysis results of pathogens include at least the species classification information obtained from the rapid classification analysis module and the species comparison information obtained from the precise comparison analysis module.
[0022] Preferably, the host sequence removal module is used to perform preliminary classification of the raw sequencing data based on a pre-constructed host reference genome database and obtain unclassified sequences according to a preset classification threshold. After aligning the unclassified sequences to the host reference genome database, the module removes any possible host sequences to obtain host-removed sequences.
[0023] Preferably, the filtering intrinsic parameter and contaminated sequence module is used to align the hostless sequence to a pre-built intrinsic parameter and contaminated sequence database, filter out the read segments that match the alignment, and obtain the effective sequence.
[0024] The present invention also provides a pathogen metagenomic data analysis device, including a processor and a memory; the processor and the memory are connected via a communication bus; wherein, the processor is used to call and execute a program stored in the memory; the memory is used to store the program, the program being used at least to execute the pathogen metagenomic data analysis method.
[0025] The beneficial effects of this invention are as follows: By introducing a candidate species screening mechanism, this invention combines the two-step analysis of "rapid classification" (S3) and "precise alignment" (S4), significantly improving the speed and accuracy of pathogen identification and achieving a balance between rapid analysis and high-precision verification of metagenomic data. This method first uses a pre-constructed pathogen species database containing a species list to perform large-scale screening of effective sequences using a rapid classification algorithm to quickly determine the pathogen range. Then, it precisely verifies these candidate targets using a high-precision alignment algorithm, significantly improving the accuracy and computational efficiency of pathogen identification. Compared with existing technologies, this invention features fast analysis speed, low false positive rate, and strong reproducibility. Attached Figure Description
[0026] Figure 1 This is a simplified flowchart of a pathogen metagenomic data analysis method according to an embodiment of the present invention.
[0027] Figure 2This is a schematic diagram of the candidate species screening mechanism between step S3 (rapid classification analysis) and step S4 (precise alignment analysis) in a pathogen metagenomic data analysis method according to an embodiment of the present invention. Detailed Implementation
[0028] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features can be combined with each other. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or a list formation.
[0029] The methods described in this specification are executed by a computer device, which may be a terminal device or a server with computing capabilities. The terminal device may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, and desktop computers.
[0030] The computer device may be configured with sequencing data analysis software. The method described in this specification can be configured within a pathogen metagenomic data analysis calculation tool or module, which is integrated into the sequencing data analysis software. After acquiring sample sequencing data, the computer device is used to analyze the aforementioned files and perform relevant pathogen metagenomic data analysis calculations. For details, please refer to the pathogen metagenomic data analysis method provided in the embodiments section of this invention (…). Figure 1 ).
[0031] This invention provides a method for pathogen metagenomic data analysis, mainly comprising five steps: host sequence removal, filtering of internal references and contaminating sequences, rapid classification analysis, precise alignment analysis, and report generation. Specifically, step S1 removes host sequences from the raw sequencing data using a combination of rapid initial screening and precise alignment; step S2 filters out internal references and contaminating sequences from the removed host sequences using alignment analysis; step S3 rapidly classifies the remaining valid sequences based on a pre-constructed pathogen species database containing a species list, and selects a candidate pathogen species sequence set according to a preset threshold; step S4 performs precise analysis of the candidate sequences using a high-precision alignment method based on a species reference genome database to confirm species information and abundance; and step S5 synthesizes the analysis results to generate an analysis result containing species classification information, species alignment information, etc. The method specifically includes the following steps.
[0032] Step S1: Removal of Host Sequences. For the acquired raw sequencing reads, the Kraken2 software and a database constructed based on the host reference genome are used to initially remove host sequences using a confidence level as a classification threshold. In one or more embodiments, the confidence level can be a decimal value between 0 and 1. Subsequently, the unclassified sequences are aligned to the host reference genome database using Bowtie2 software to further confirm and remove any possible host sequences, ultimately obtaining the host-removed sequences. In one or more embodiments, the alignment parameters can be a preset set of parameters in Bowtie2 software, such as sensitive-local, very-sensitive-local, fast-local, and very-fast-local. Alternatively, parameters can be manually specified to precisely adjust the alignment parameters.
[0033] Step S2: Filtering Internal References and Contaminating Sequences. Using the host-decontaminated sequence as input, a database constructed based on internal reference sequences and common contaminating sequences in the laboratory is selected, and bowtie2 is used for alignment. The matched reads are filtered to obtain valid sequence data after removing internal references and contamination. The alignment parameters include any one or more of the following: seed length, maximum number of mismatches, optimal matching mode for output alignment results, and filtering sequences with alignment quality below a preset threshold. In one or more embodiments, the alignment parameters can be a set of parameters preset by bowtie2 software, such as sensitive-local, very-sensitive-local, fast-local, and very-fast-local. In addition, the alignment parameters can be manually specified and precisely adjusted.
[0034] Step S3: Rapid Classification Analysis. Input the valid sequences obtained in Step 2 into Kraken2, and use a pre-built pathogen species database to rapidly classify them based on confidence levels. For example... Figure 2 As shown, this step sets filtering conditions, selecting species from the species list whose sequence count is less than the filtering threshold as candidate pathogen species, and extracting their corresponding reads as a candidate sequence set. In one or more embodiments, the confidence level can be a decimal value between 0 and 1.
[0035] Step S4: Precise Alignment Analysis. For each candidate pathogen species screened in Step 3, the system automatically calls the species reference genome database of its corresponding genus or family and performs parallel precise alignment using bowtie2. Based on the alignment results, the presence and abundance information of the species in the sample are finally determined. The species reference genome database includes all sequences of the corresponding family and / or genus of the species. Correspondingly, for each candidate species, the corresponding family or genus sequences in its species reference genome database are simultaneously called for parallel alignment, thereby shortening the alignment time and improving processing efficiency. In one or more embodiments, the alignment parameters can be a set of parameters preset by bowtie2 software, such as sensitive-local, very-sensitive-local, fast-local, and very-fast-local. In addition, the alignment parameters can be manually specified and precisely adjusted.
[0036] Step S5: Report Generation. Based on the species classification and alignment information obtained in steps S3 and S4, and according to the species classification hierarchy, a complete pathogen metagenomic analysis result file is output, as shown in the example in Table 1. The result file includes species classification ID, species classification hierarchy, species name, number of species sequences, cumulative species sequences, species abundance, etc.
[0037] In one specific embodiment, the above-described method was used to perform pathogen metagenomic analysis on the sequencing data of a sample. The sample was peripheral blood, which underwent metagenomic DNA library construction and sequencing on a CELOS sequencer, yielding 1,609,433 reads. Based on a host reference genome database and using a confidence threshold of 0.5, a rapid initial host removal process was performed on the sequencing data, resulting in 1,134,605 unclassified sequences. After alignment with the host database using Bowtie2 with a sensitive-local parameter, 129,345 host sequences were removed, ultimately resulting in 1,005,260 host-removed sequences. Alignment of the host-removed sequences with an internal reference and contamination sequence database, using a sensitive-local alignment parameter, yielded 935,738 valid sequence data. Then, based on a pathogen species database, the valid sequence data were rapidly classified using a confidence threshold of 0.5. Candidate pathogen species sequences were obtained by selecting from the species list and using a selection threshold of 10,000,000. For each candidate pathogen species selected, the system automatically invoked the reference genome database of its corresponding genus or family for precise comparison. Using comparison coverage and comparison score as comparison indicators (thresholds for each indicator), the system identified the presence of pathogen species such as *Corynebacterium propinquum*, *Rhizomucor pusillus*, *Streptococcus parasanguinis*, and *Human betaherpesvirus 5* in the sample. Partial analysis results are shown in Table 1.
[0038] Table 1. Analysis Results
[0039]
[0040] In one specific embodiment, a pathogen metagenomic data analysis system is also provided, including a host sequence removal module, an internal reference and contaminated sequence filtering module, a precise alignment analysis module, and a report generation module. The host sequence removal module is used to perform preliminary classification of raw sequencing data based on a pre-constructed host reference genome database and obtain unclassified sequences according to a preset classification threshold. After aligning the unclassified sequences to the host reference genome database, any possible host sequences are removed to obtain host-free sequences. The internal reference and contaminated sequence filtering module is used to align the host-free sequences to a pre-constructed internal reference and contaminated sequence database, filtering out reads that match the alignment to obtain valid sequences. The rapid classification analysis module is used to rapidly classify valid sequences based on a pre-constructed pathogen species database containing a species list to obtain species classification information; species in the species list with a sequence count less than a screening threshold are selected as candidate pathogen species, and their corresponding reads are extracted as candidate pathogen species sequences. The precise alignment analysis module is used to align each candidate pathogen species obtained from the rapid classification analysis module with its corresponding species reference genome database to obtain species alignment information; the species alignment information includes at least the presence and abundance information of the candidate pathogen species in the sample. The report generation module is used to output the pathogen metagenomic analysis results; the pathogen metagenomic analysis results include at least the species classification information obtained from the rapid classification analysis module and the species alignment information obtained from the precise alignment analysis module.
[0041] In one specific embodiment, a pathogen metagenomic data analysis device is also provided, including a processor and a memory; the processor and the memory are connected via a communication bus; wherein, the processor is used to call and execute a program stored in the memory; the memory is used to store the program, the program being used at least to execute the pathogen metagenomic data analysis method.
[0042] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope of the present invention.
Claims
1. A method for analyzing pathogen metagenomic data, characterized in that, Includes the following steps: S1. Remove the host sequence from the original sequencing data to obtain the host-free sequence; S2. Filter out the intrinsic parameter sequences and contaminating sequences in the host sequence to obtain the effective sequence; S3. Based on a pre-constructed pathogen species database containing a list of species, effective sequences are rapidly classified to obtain species classification information; Species in the species list with a sequence number less than the screening threshold will be selected as candidate pathogen species, and their corresponding reads will be extracted as candidate pathogen species sequences. S4. Compare each candidate pathogen species obtained in step S3 with its corresponding species reference genome database to obtain species comparison information; the species comparison information includes at least the presence and abundance information of the candidate pathogen species in the sample. S5. Output the metagenomic analysis results of the pathogen; the metagenomic analysis results of the pathogen shall include at least the species classification information obtained in step S3 and the species comparison information obtained in step S4.
2. The method as described in claim 1, characterized in that, Step S1 includes: based on a pre-constructed host reference genome database, performing preliminary classification on the raw sequencing data and obtaining unclassified sequences according to a preset classification threshold, aligning the unclassified sequences to the host reference genome database and removing any possible host sequences to obtain host-free sequences.
3. The method as described in claim 1, characterized in that, Step S2 includes: aligning the host-de-host sequence to a pre-built database of internal parameters and contaminated sequences, filtering out reads that match the alignment, and obtaining valid sequences.
4. The method as described in claim 3, characterized in that, The parameters for comparison in step S2 include any one or a combination of seed length, maximum number of mismatches, optimal matching mode of the output comparison result, and sequences whose comparison quality is lower than a preset threshold.
5. The method as described in claim 1, characterized in that, The species reference genome database mentioned in step S4 includes the complete sequences of the corresponding family and / or genus of the species.
6. The method as described in claim 5, characterized in that, The comparison method in step S4 includes: for each candidate species, calling the corresponding family or genus sequence in the species reference genome database for parallel comparison.
7. The method as described in claim 1, characterized in that, The pathogen metagenomic analysis results described in step S5 shall include at least the species ID, species name, species level, number of species sequences, cumulative number of species sequences, and species abundance information.
8. A pathogen metagenomic data analysis system, characterized in that, include: The host sequence removal module is used to remove the host sequence from the raw sequencing data to obtain the host-free sequence; The module for filtering internal parameters and contaminated sequences is used to filter out internal parameter sequences and contaminated sequences in the host sequence to obtain a valid sequence. The rapid classification and analysis module is used to quickly classify valid sequences based on a pre-built pathogen species database containing a list of species, and obtain species classification information; Species in the species list with a sequence number less than the screening threshold will be selected as candidate pathogen species, and their corresponding reads will be extracted as candidate pathogen species sequences. The precise alignment analysis module is used to align each candidate pathogen species obtained from the rapid classification analysis module with its corresponding species reference genome database to obtain species alignment information; the species alignment information includes at least the presence and abundance information of the candidate pathogen species in the sample. The report generation module is used to output the metagenomic analysis results of pathogens; the metagenomic analysis results of pathogens include at least the species classification information obtained from the rapid classification analysis module and the species comparison information obtained from the precise comparison analysis module.
9. The system as described in claim 8, characterized in that, The host sequence removal module is used to perform preliminary classification of raw sequencing data based on a pre-constructed host reference genome database and obtain unclassified sequences according to a preset classification threshold. After aligning the unclassified sequences to the host reference genome database, any possible host sequences are removed to obtain host-removed sequences; and / or, The filtering internal parameter and contaminated sequence module is used to align the hostless sequence to a pre-built internal parameter and contaminated sequence database, filter out the read segments that match the alignment, and obtain the valid sequence.
10. A pathogen metagenomic data analysis device, characterized in that, It includes a processor and a memory; the processor and the memory are connected via a communication bus; wherein the processor is used to call and execute a program stored in the memory; the memory is used to store the program, which is at least used to execute the pathogen metagenomic data analysis method according to any one of claims 1 to 7.