Pathogenic microorganism genome database, construction method, computer system and application

By combining the ANI and AF indicators with contig and k-mer contamination identification methods, a high-quality pathogenic microorganism genome database was constructed, which solved the problems of database contamination and classification errors in existing technologies and achieved high accuracy and efficient analysis of pathogenic microorganism detection.

CN120600129APending Publication Date: 2025-09-05AUTOBIO DIAGNOSTICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510601805.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing pathogen metagenomic testing has problems such as database contamination, classification errors, large data volume and slow analysis speed, resulting in low detection rate and accuracy of pathogenic microorganisms, making it difficult to apply in clinical practice.

Method used

ANI combined with AF dual indicators were used to identify classification errors. A high-quality pathogenic microorganism genome database was constructed by combining the contig contamination sequence and k-mer short sequence contamination identification methods. ANI and AF values ​​were calculated using FastANI software. Contamination sequences were identified and replaced using the BLAST database to construct a self-built database.

Benefits of technology

It improves the accuracy and analysis efficiency of pathogen identification, ensures species genome diversity, shortens analysis time, reduces false positives and false negatives, and improves the accuracy and sensitivity of pathogen infection species identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600129A_ABST
    Figure CN120600129A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bioinformatics, in particular to a pathogenic microorganism genome, a construction method, a computer system and application. According to the method, a genome data quality control strategy in the pathogenic microorganism genome database creation process is innovated, ANI and AF double indexes are creatively adopted for classification error recognition, contig pollution sequence recognition and k-mer short sequence pollution recognition are combined for sequence pollution recognition and processing on genome data, the number of genomes is reserved to the maximum, and the number of the genomes is reduced to the maximum. According to the method, contig level sequence errors are accurately recognized, strain specific sequences are prevented from being mistakenly deleted, the diversity of species genomes is guaranteed, and the constructed database can improve the identification accuracy / sensitivity of pathogen infected species.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and in particular to a pathogenic microorganism genome, a construction method, a computer system and applications. Background Art

[0002] Pathogen metagenomic high-throughput sequencing (mNGS) technology has been widely used in the auxiliary diagnosis of difficult clinical infections due to its advantages such as short detection time, high resolution, and ability to identify infections caused by rare and emerging pathogens or mixed infections. A key component of pathogen metagenomic technology is the pathogen genome database, and the quality of the database directly affects the detection rate, accuracy, and other analytical performance of pathogen metagenomic testing.

[0003] However, due to the lack of standardized bioinformatics analysis processes and high-quality clinical microbial genome reference databases, the further development of the clinical application of this technology has been restricted to a certain extent. In the analysis process currently used in the industry for pathogen metagenomic detection, there are generally two ways to construct a database. One is to select a genome or sequence for each species as a reference. The advantage of this method is that the database constructed is that it takes up little storage resources and has a fast analysis speed, but it is easy to miss detections, and the pathogen detection rate and accuracy are very low, which basically cannot be used in actual clinical practice. The other is to include all genomes of the same species in the reference database. Although this method solves the problem of missed detection that is prone to the previous method, its data volume is too large, the analysis speed is slow, and the data quality in public databases is uneven. There are often classification errors, contaminated sequences and / or redundant sequences, which are prone to pathogen identification errors, false positives and other defects, and it is basically impossible to actually apply it in clinical practice.

[0004] To address these existing issues, the industry has gradually conducted scientific research on methods for constructing pathogen genome databases and developed practical tools. Current research has found that reference sequence databases suffer from various common problems, with database contamination being the most recognized issue. Systematic evaluations have identified 2,161,746 contaminated sequences in NCBI GenBank and 114,035 in its higher-quality subset, RefSeq. However, database issues extend far beyond contamination. The default databases used in most popular tools are subject to classification errors, inappropriate inclusion and exclusion criteria, and errors in the sequences themselves. This phenomenon is the result of most metagenomic tools simply mirroring NCBI resources (including NCBI GenBank, RefSeq, Taxonomy, and the BLAST nucleotide database) into their databases. Taxonomic misannotation is prevalent in NCBI GenBank and RefSeq, affecting an estimated 3.6% of prokaryotic genomes in GenBank and approximately 1% of its curated subset, RefSeq.

[0005] Currently, there are also some commercially available tools for screening and identifying sequences with genomic classification errors and contamination in public databases. For example, checkM software is currently a good software for assessing genome integrity and contamination. It has been verified that the software provides accurate estimates of genome integrity and contamination based on lineage-specific marker genes. The software requires assembled genome files and raw sequencing data for accurate assessment. This makes calculation impossible for genome files uploaded earlier to the NCBI database due to the lack of raw data. For recently uploaded genome files, the need to download raw data also increases the download volume and complexity. Raw data is often more than 1,000 times the size of the genome file, and there are no standardized genome file management methods. Furthermore, checkM software currently only calculates genome integrity and contamination rate and does not provide output for contamination sequences.

[0006] FastANI is a software that calculates the nucleotide sequence consistency of genome files. FDA-ARGOS and the Chinese Academy of Sciences' microbial database gcPathogen both use the ANI (average nucleotide identity) value calculated by the software to determine whether the genome classification is incorrect. The threshold is generally 95%, and a value below 95% is considered a classification error. The problem with this type of method is that the ANI calculation is only based on homologous sequences, and non-homologous sequences are not involved in the calculation. Genomes with a low proportion of homologous sequences but a high homology sequence consistency rate cannot be accurately identified. Moreover, the software can only determine whether the genome is misclassified based on the ANI value, and cannot output contaminated sequences for further removal of contaminated sequences while retaining the entire genome file. For species with fewer genome files themselves, the genome is usually directly filtered because the ANI value does not meet the quality requirements, resulting in a small number of species genomes and insufficient diversity. Summary of the Invention

[0007] In order to overcome the shortcomings of the prior art, one of the objectives of the present invention is to provide a method for constructing a pathogenic microorganism genome database. The constructed database covers a rich diversity of pathogen genome data, has the advantages of high pathogen identification accuracy, short analysis time, and cost savings.

[0008] A second object of the present invention is to provide a computer system for constructing a pathogenic microorganism genome database.

[0009] The third object of the present invention is to provide a pathogenic microorganism genome database that ensures the diversity of species genomes and can improve the accuracy / sensitivity of metagenomic infectious species identification compared to traditional databases.

[0010] A fourth object of the present invention is to provide an application of a pathogenic microorganism genome database in metagenomic sequencing analysis and detection of pathogenic microorganisms.

[0011] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0012] A method for constructing a pathogenic microorganism genome database, comprising:

[0013] 1) Obtain genomic data;

[0014] 2) Confirm reference genome and non-reference genome;

[0015] 3) Perform classification error identification, and / or contig sequence contamination identification, and / or short sequence contamination identification on the non-reference genome data;

[0016] 4) Step 3) identifying and filtering the genome data retained after the removal process and merging it with the reference genome data;

[0017] The genome data of each species were processed according to the above method and merged into a complete pathogenic microorganism genome database;

[0018] The specific method for identifying misclassifications involves comparing a species' non-reference genome with the reference genome and calculating the ANI and AF values. The AF value is the number of homologous sequence fragments divided by the total number of sequence fragments. ANI and AF thresholds are set, and a genome is identified as misclassified if either the ANI or AF value does not meet the threshold. It should be noted that the ANI value represents the average nucleotide identity rate.

[0019] Optionally, genomes with ANI>threshold 1 and AF>threshold 2 are identified as qualified genomes, and then contig sequence contamination identification and short sequence contamination identification are performed in sequence; threshold 1 is 85% to 90%; threshold 2 is 75% to 80%.

[0020] Preferably, the thresholds for bacteria, archaea, and parasites are: threshold 1 is 85% to 90%; threshold 2 is 75 to 80%; fungi threshold 1 is 80% to 85%; threshold 2 is 70 to 75%.

[0021] In a specific embodiment of the present invention, the sources of genomic data in step 1) include the NCBI refseq database and the GenBank database; wherein in step 2), the genome in the NCBI refseq database is determined as the reference genome, and the genomic data from other data sources are defined as non-reference genomes.

[0022] Optionally, the pathogenic microorganisms are bacteria and archaea, and step 3) only performs classification error identification on the non-reference genome data; before classification error identification, the integrity of the non-reference genome is also screened. As an example, the screening conditions can be to retain the genome with the best ANI matching species as the target species, the genome status is normal, the genome assembly level is at the chromosome level, the assembly coverage is at least 20X or more, the genome completeness is more than 70%, and the genome predicted contamination rate is less than 10%; and remove genomes whose genome size error with the reference genome is more than 50%.

[0023] Optionally, the pathogenic microorganism is a fungus or a parasite, and step 3) performs classification error identification, and / or contig sequence contamination identification, and short sequence contamination identification on the non-reference genome data in sequence.

[0024] Optionally, the specific method for identifying contig sequence contamination is: aligning the correctly classified non-reference genome contig-level sequences to the reference genomes of each species, and counting the alignment consistency rate (ide) and the alignment base percentage of the contigs to the reference genomes of each species. If the alignment consistency rate (ide) and the alignment base percentage of the non-target species alignment results are higher than those of the target species, then the contig sequence is judged to be a contamination sequence derived from the non-target species, and the contig sequence is further deleted from the genome file.

[0025] Optionally, a specific method for identifying short sequence contamination is as follows: breaking the correctly classified non-reference genome into kmers of length L and step size k; aligning all kmers to the reference genome, and calculating the alignment consistency rate (ide) of the kmer to the reference genome of each species. If the alignment consistency rate (ide) of the non-target species alignment result is higher than that of the target species by more than 10%, the kmer is judged to be a contaminating sequence, and the bases of the kmer at the genome position are further replaced by consecutive Ns;

[0026] It should be understood that the continuous N replacement of bases means that all the contaminating base sequences are replaced with N.

[0027] Optionally, the length L is consistent with the sequencing length; the step length k is 50 bp. In a specific embodiment of the present invention, the length L = 150 bp.

[0028] Furthermore, in a specific embodiment of the present invention, kmers that are not aligned to the reference genomes of various species are determined to be strain-specific sequences and are not processed to avoid accidental deletion of strain-specific sequences.

[0029] A computer system for constructing a pathogenic microorganism genome database comprises a processor and a memory; the processor and the memory are communicatively connected, wherein the memory is used to store a computer program, and the processor is used to call the computer program, wherein the computer program comprises program instructions, and when the program instructions are executed by the processor, the method described above is performed;

[0030] Optionally, the program instructions for executing classification error identification include using FastANI software to output the ANI value, the number of homologous sequence fragments, and the total number of sequence fragments, and inputting the next program instruction to run the calculation of AF value = the number of homologous sequence fragments divided by the total number of sequence fragments to calculate the homologous sequence.

[0031] The application of the high-quality genome database of pathogenic microorganisms constructed by the above construction method in the detection of pathogenic microorganisms by metagenomic sequencing analysis.

[0032] Beneficial effects of the present invention:

[0033] 1) The present invention innovates the genome data quality control strategy during the creation of the pathogenic microorganism genome database, especially the creative use of ANI combined with AF dual indicators for classification error identification, contig contamination sequence identification combined with k-mer short sequence contamination identification to identify and process sequence contamination in genome data, maximize the number of retained genomes, accurately identify contig-level sequence errors, avoid the accidental deletion of strain-specific sequences, ensure the diversity of species genomes, and the constructed database can improve the accuracy and sensitivity of pathogen infection species identification.

[0034] 2) Optimize ANI and AF thresholds for different species categories to further improve the quality control of genomic data, accurately identify misclassified genomes while maximizing the number of retained genomes to ensure species genomic diversity (especially for species with fewer genomes, such as fungal parasites);

[0035] 3) Contig contamination sequence identification is combined with k-mer short sequence contamination identification to identify and process sequence contamination in genomic data, thereby improving the accuracy of species genomic data in the database, simplifying the database, and improving analysis efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a schematic diagram of the statistical analysis results of the distribution of ANI values ​​and AF values ​​after alignment of the Stenotrophomonas maltophilia strain genome with the reference genome;

[0037] Figure 2 Schematic diagram comparing the differences in analytical accuracy of the Stenotrophomonas maltophilia database constructed using different methods in Example 1;

[0038] Figure 3 Schematic diagram of the statistical analysis results of the distribution of ANI values ​​and AF values ​​after alignment of the Blastomyces dermatitidis strain genome with the reference genome;

[0039] Figure 4 Schematic diagram of the identification results of contig fragments of strain genome GCA_000151595.1;

[0040] Figure 5 This is a schematic diagram comparing the differences in analysis accuracy of the Blastomyces dermatitidis database constructed using different methods in Example 2. DETAILED DESCRIPTION

[0041] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0042] Unless otherwise specified, the techniques used and covered herein are standard methods known to those skilled in the art. The materials, methods and examples are for illustrative purposes only and are not intended to limit the scope of the present invention in any way.

[0043] The present invention relates to computer media or computer system products, specifically including computer-readable storage media on which computer-readable program instructions for executing the present invention are loaded; wherein "computer-readable storage media" refers to a tangible device that can maintain and store instructions used by an instruction execution device.

[0044] It should be understood that during the construction of the pathogenic microorganism genome database of the present invention, genomes with plasmid sequences may be processed to remove the plasmid sequences, and methods known to those skilled in the art may be used for such processing.

[0045] Example 1: Constructing a bacterial database using the method of the present invention

[0046] 1. Data preparation: Taking Stenotrophomonas maltophilia (taxid: 40324) as an example, first download the reference genome GCF_900186865.1 (genome tagged with refseq) and all other strain genome files from the NCBI database. Relevant information can be searched in the genome module and jumped to the relevant download link. The downloaded data includes the result labels of the strain genome comparison with the reference genome (below_threshold_mismatch, mismatch, and species_match). The strain genome with the mismatch label is removed, and the other strain genomes are run in the next step to identify and remove classification errors;

[0047] 2. Genomic classification error identification: FastANI software was used to calculate the average nucleotide identity (ANI) value of the genome of a single strain and the reference genome of the species, and the number of homologous sequence fragments and the total number of sequence fragments were output. The homologous sequence fraction AF was calculated as the number of homologous sequence fragments / the total number of sequence fragments.

[0048] The accuracy of genome classification is determined by combining the ANI value and the fraction of homologous sequences (AF). Thresholds are set, with ANI > 90% and AF > 75%. Genomes that do not meet these criteria are considered misclassified and removed. The remaining strain genomes are merged with the reference genome to form a self-built database for the species.

[0049] This example uses Stenotrophomonas maltophilia as an example to evaluate the beneficial effects of constructing a bacterial species database using the method of the present invention:

[0050] (1) The statistical analysis results of the distribution of ANI and AF values ​​after comparing the genome of Stenotrophomonas maltophilia strain with the reference genome are as follows: Figure 1 As shown in the figure above, it can be clearly seen that these data are divided into five clusters, indicating that the species diversity of Stenotrophomonas maltophilia is relatively rich. If it is simply divided according to ANI>95%, it will lead to a lack of species diversity and cause false negatives; but if all are retained, it will cause false positives.

[0051] (2) Accuracy Verification: To verify the accuracy of the self-built database constructed by the method of the present invention, i.e., to preserve species diversity and eliminate misclassified species, species with complete assembly level were randomly selected, namely GCF_000072485.1, GCF_002951115.1, GCF_003030985.1, GCF_008693985.1, GCF_009676405.1, GCF_009676425.1, GCF_009676525.1, GCF_009676605.1, GCF_011386925.1, and GCF_012647025.1. Using the simulation software dwgsim, a simulated dataset with a length of 75 bp, an error rate of 0.08%-0.2%, and a coverage of 100X was generated.

[0052] Comparison of analytical accuracy differences between different databases, such as Figure 2 As shown, the genome database is the reference genome of the species and all other strain genomes downloaded from the NCBI database; the database of ANI≥95% genomes only selects genomes with ANI≥95%. Figure 2 The results show that the accuracy of the simulated dataset for the database constructed from all genomes was 100%; the accuracy of the database constructed from genomes with an ANI ≥ 95% was an average of 94.6%; and the accuracy of the self-constructed database was 97.4%. This indicates that the accuracy of the database constructed using the present method is comparable to that of all genomes, but the analysis speed is improved by 8.9%. Therefore, the database constructed using the present method can maintain accuracy while taking into account species diversity and reducing analysis time.

[0053] Example 2:

[0054] 1. Data preparation: Taking Blastomyces dermatitidis (taxid: 5039) as an example, first download the reference genome GCA_000003525.2 and all other strain genome files from the NCBI database. Relevant information can be searched in the genome module and jumped to the relevant download link. These data will serve as the basis for subsequent analysis;

[0055] 2. Identification of genome classification errors: Use FastANI software to calculate the average nucleotide identity (ANI) value of a single strain genome and the reference genome or representative genome of the species, and output the number of homologous sequence fragments and the total number of sequence fragments, and calculate the homologous sequence ratio AF = number of homologous sequence fragments / total number of sequence fragments; combine the ANI value and the homologous sequence ratio (AF) to judge the classification accuracy of the genome. Set the threshold, ANI>80% and AF>70%. If the conditions are not met, it is judged as a genome classification error and removed. The classification results are distributed as follows Figure 3 As shown;

[0056] 3. Identification of contig sequence contamination in the strain genome: Construct a local BLAST database based on the RefSeq reference genome database of all fungal species. Align the contig sequences of the genome to the local BLAST database, and count the coverage ratio and coverage similarity of the contigs aligned to the reference genome of each species. If the coverage ratio and coverage similarity of the non-target species are higher than those of the target species (Blastomyces dermatitidis), the contig sequence is judged to be a contamination sequence from the non-target species; after marking it as a contamination sequence from the non-target species, the contig sequence is further deleted from the genome file and replaced with continuous N bases;

[0057] Taking the GCA_000151595.1 genome as an example, the number of contigs removed by contig sequence identification was counted. Figure 4 As shown, approximately 19% of the sequences were identified as the closely related species Blastomyces gilchristii;

[0058] 4. Identification of contamination in long fragment sequences of the strain genome: Continue to break the remaining strain genome after step 3 into kmers with a length of 150bp and a step length of 50bp, and align the generated kmers to the local BLAST database using blastn. Count the consistency rate of each kmer aligned to the reference genome of each species. Identify kmers whose consistency rate with non-target species is more than 10% higher than that of the target species (Blastomyces dermatitidis) and judge them as potential contaminating kmers. Mark the positions of the kmers identified as contamination in the genome. Use genome editing tools to replace these positions with consecutive N bases to remove contamination.

[0059] The strain genome data after the above processing is integrated with the reference genome to serve as a self-built database of the species constructed according to the method of the present invention.

[0060] This example takes Blastomyces dermatitidis as an example to illustrate the beneficial effects of constructing a database using the method of the present invention:

[0061] (1) Taking the strain genome GCA_000151595.1 as an example, the comparison results before and after long fragment sequence contamination identification treatment are statistically analyzed, as shown in Table 1 below:

[0062] Table 1

[0063] GCA_000151595.1 N quantity Total bases N content ratio (%) Before treatment 5451528 72293421 7.54 After processing 25452111 72293421 35.21

[0064] As can be seen from the above table, the N content increased from 7.54% to 35.21%, indicating that 27.67% of the regions were aligned to other species, further demonstrating that the database construction method of the present invention can more fully identify and remove assembly errors.

[0065] (2) In order to further verify the beneficial effects of the method of the present invention, the analysis accuracy of the genome databases constructed with different quality control levels was compared. The genome data in three different databases (original genome database: the reference genome downloaded from the NCBI database and the database constructed by the genomes of all strains; "moderate" streamlined genome (the database constructed by removing contig contamination fragments as shown in step 3); "heavy" streamlined genome (the self-built database constructed by the method of the present invention)) were regenerated from scratch using the simulation software dwgsim. The simulated data sets were 75 bp in length, with an error rate of 0.08%-0.2% and a coverage of 100X. The simulated data generated by the three databases were compared with the NCBI refseq database, and the correctly classified reads were calculated. The results are shown in Figure 2. Figure 5As shown, the results show that the accuracy of the simulated data set in the original genome analysis results is less than 95%, mainly due to the contamination sequences in the original genome, which cannot be correctly aligned to the species; and with the removal of different levels of contamination fragments, the read recognition accuracy gradually increases, indicating that the database constructed by the method of the present invention for fungal species can fully identify contaminated data, while improving the accuracy of the analysis results, streamlining the amount of database data and increasing the analysis speed.

[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for constructing a pathogenic microorganism genome database, characterized in that: include: 1) Obtain genomic data; 2) Confirm reference genome and non-reference genome; 3) Perform classification error identification, and / or contig sequence contamination identification, and / or short sequence contamination identification on the non-reference genome data; 4) Step 3) identifying and filtering the genome data retained after the removal process and merging it with the reference genome data; The genome data of each species were processed according to the above method and merged into a complete pathogenic microorganism genome database; The specific method for identifying classification errors is to compare the non-reference genome of the species with the reference genome and calculate the ANI value and AF value, where the AF value is the number of homologous sequence fragments divided by the total number of sequence fragments; set the ANI and AF thresholds, and comprehensively judge that if either the ANI value or the AF value does not meet the threshold setting conditions, it will be identified as a misclassified genome.

2. The method for constructing a pathogenic microorganism genome database according to claim 1, wherein: Genomes with ANI>threshold 1 and AF>threshold 2 were identified as qualified genomes, and then contig sequence contamination identification and short sequence contamination identification were performed in sequence; the threshold 1 was 85%-90%; the threshold 2 was 75%-80%.

3. The method for constructing a pathogenic microorganism genome database according to claim 2, wherein: The thresholds for bacteria, archaea, and parasites are: threshold 1 is 85% to 90%; threshold 2 is 75% to 80%; threshold 1 for fungi is 80% to 85%; threshold 2 is 70% to 75%.

4. The method for constructing a pathogenic microorganism genome database according to any one of claims 1 to 3, wherein: In step 1), the genomic data are obtained from the NCBI refseq database and the GenBank database; in step 2), the genome in the NCBI refseq database is determined as the reference genome, and the genomic data from other data sources are defined as non-reference genomes; Optionally, the pathogenic microorganism is bacteria or archaea, and step 3) only performs classification error identification on the non-reference genome data; before the classification error identification, the integrity of the non-reference genome is also screened; Optionally, the pathogenic microorganism is a fungus or a parasite, and step 3) performs classification error identification, and / or contig sequence contamination identification, and short sequence contamination identification on the non-reference genome data in sequence.

5. The method for constructing a pathogenic microorganism genome database according to claim 4, wherein: The specific method for identifying contig sequence contamination is as follows: align the correctly classified non-reference genome contig-level sequences to the reference genomes of each species, and count the alignment consistency rate (ide) and alignment base percentage of the contigs to the reference genomes of each species. If the alignment consistency rate (ide) and alignment base percentage of the non-target species alignment results are higher than those of the target species, then the contig sequence is judged to be a contamination sequence derived from the non-target species, and the contig sequence is further deleted from the genome file.

6. The method for constructing a pathogenic microorganism genome database according to claim 4, wherein: The specific method for identifying short sequence contamination is as follows: the correctly classified non-reference genome is broken into kmers of length L and step size k; all kmers are aligned to the reference genome, and the alignment consistency rate (ide) of the kmer to the reference genome of each species is calculated. If the alignment consistency rate (ide) of the non-target species alignment result is higher than that of the target species by more than 10%, the kmer is judged as a contaminating sequence, and the bases of the kmer are further replaced with consecutive Ns at the genome position. Optionally, the length L is consistent with the sequencing length; the step length k is 50 bp.

7. The method for constructing a pathogenic microorganism genome database according to claim 6, wherein: Kmers that were not aligned to the reference genome of each species were considered strain-specific sequences and were not processed.

8. A pathogenic microorganism genome database constructed by the method according to any one of claims 1 to 7.

9. A computer system for constructing a pathogenic microorganism genome database, characterized in that: comprising a processor and a memory; the processor and the memory are communicatively connected, wherein the memory is used to store a computer program, the processor is used to call the computer program, the computer program includes program instructions, and when the program instructions are executed by the processor, the method according to any one of claims 1 to 8 is executed; Optionally, the program instructions for executing classification error identification include using FastANI software to output the ANI value, the number of homologous sequence fragments, and the total number of sequence fragments, and inputting the next program instruction to run the calculation of AF value = the number of homologous sequence fragments divided by the total number of sequence fragments to calculate the homologous sequence.

10. Use of the high-quality pathogenic microorganism genome database according to claim 8 in the detection of pathogenic microorganisms by metagenomic sequencing analysis.