Pathogenic microorganism genome reference database and construction method thereof
By employing multi-source data fusion and hierarchical database construction, the lack of uniformity and scalability in existing pathogen databases has been addressed, enabling efficient pathogen detection and risk monitoring, which is suitable for clinical and research applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG HONG KONG MACAO GREATER BAY AREA PRECISION MEDICINE RESEARCH INSTITUTE (GUANGZHOU)
- Filing Date
- 2025-12-22
- Publication Date
- 2026-05-01
AI Technical Summary
Existing pathogen reference databases lack a unified, high-quality, and scalable large-scale system, resulting in low pathogen comparison efficiency and poor identification accuracy, failing to meet the needs of broad-spectrum clinical detection and rapid analysis.
A pathogenic microorganism genome reference database is constructed by employing multi-source data fusion, hierarchical database construction, redundancy removal and standardization processing, and linkage with clinical knowledge system. This includes multi-source data integration, standardized mapping, sequence clustering, hierarchical database construction, and automated updates, supporting high-precision pathogen detection and risk monitoring.
It has achieved a comprehensive, hierarchical, and quality-controllable pathogen database, with excellent pathogen detection, classification, comparison, and risk monitoring capabilities, and is suitable for clinical and research scenarios.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics, and more specifically, to a pathogenic microorganism genome reference database and its construction method. Background Technology
[0002] With the rapid development of high-throughput sequencing (HTS) technology, nucleic acid-based pathogen detection has been widely applied in clinical microbiology and public health. Next-generation sequencing (NGS) can acquire large-scale data from clinical infection specimens within 24-72 hours, and it has significant advantages over traditional bacterial culture, immunological detection, and PCR amplification methods in detecting difficult-to-culture, rare, or emerging pathogens. Therefore, constructing a comprehensive and accurately annotated pathogen genome reference sequence database has become a key infrastructure in modern pathogen identification systems.
[0003] In traditional detection methods, pathogen culture relies on specific growth conditions, significantly limiting its sensitivity and detection range. While PCR methods offer high specificity, prior knowledge of the target sequence is required, making them unsuitable for unknown or highly variable pathogens. With the widespread adoption of whole-genome sequencing (WGS) and metagenomic next-generation sequencing (mNGS), clinicians can obtain the complete genomic information of all microorganisms in a sample in a single sequencing run, enabling unbiased and broad-spectrum pathogen detection.
[0004] However, such sequencing technologies are highly dependent on the quality of the reference database. The lack of a large-scale, standardized, and accurately annotated pathogen reference genome library will severely impact pathogen alignment efficiency, identification accuracy, and even lead to false positives or false negatives. Current technologies typically focus on constructing reference data for specific diseases or types of pathogens, and a unified, high-quality, scalable, large-scale pathogen genome reference database system has not yet been established. Therefore, current technologies still cannot meet the comprehensive needs of broad-spectrum clinical detection, rapid analysis, and continuous database updates.
[0005] Therefore, constructing a large-scale pathogenic microorganism reference database (PGDQ) has become an important prerequisite for the clinicalization and standardization of mNGS. Summary of the Invention
[0006] This application is based on the inventor's discoveries and understanding of the following facts and problems: Rapid identification of pathogenic microorganisms using high-throughput sequencing technology has become an important means of infectious disease diagnosis and public health monitoring. This invention proposes a method and application platform for constructing a pathogenic microorganism genome reference database based on multi-source data fusion, hierarchical database construction, redundancy removal and standardization processing, and linkage with clinical knowledge system, which has good data integration and retrieval capabilities.
[0007] Therefore, in a first aspect of the present invention, the present invention proposes a method for constructing a pathogenic microorganism genome reference database. According to an embodiment of the present invention, the method includes the following steps: (1) acquiring multi-source pathogen sequence data from multiple public databases and autonomous clinical data, integrating them according to preset data source priority rules, and achieving semantic uniformity and format compatibility of multi-source data through a cross-database data standardization mapping system; (2) cleaning and quality control the integrated data to obtain cleaned data; (3) based on the cleaned data, using a sequence clustering algorithm combined with functional gene integrity verification to perform function-oriented redundancy removal and standardized naming to obtain a standardized sequence dataset; (4) based on the standardized sequence dataset, constructing a pan-genome model, and building a hierarchical database based on the pan-genome model, the hierarchical database including at least a core database and an extended database; (5) based on the hierarchical database, constructing a clinical priority-oriented query system; and (6) automatically and dynamically updating the hierarchical database to achieve the emergency inclusion of emerging pathogens. The pathogenic microorganism genome reference database constructed according to the method described in this invention is comprehensive, hierarchically structured, and of controllable quality, and can be used in clinical and research scenarios. At the same time, the accompanying query system and knowledge service module can enhance the ability to detect, classify, compare, track, and monitor pathogens.
[0008] According to an embodiment of the present invention, the method further includes the following steps in step (1): establishing a unified mapping table for metadata fields, mapping the core fields of different databases to a unified name, data type and value range; based on the pathogenic microorganism ontology terminology database, using natural language processing (NLP) algorithms combined with homologous sequence alignment, unifying the annotation rules of different databases and correcting annotation contradictions; when there is information conflict in multi-source data, automatically adjudicating according to the rule of "data source weight + clinical evidence priority", and retaining high-credibility data.
[0009] According to an embodiment of the present invention, in step (3) of the method, the redundancy removal threshold of the sequence clustering algorithm is a sequence similarity ≥ 95%.
[0010] According to an embodiment of the present invention, the function-oriented deredundancy processing further includes: performing core functional gene integrity verification on each cluster and retaining representative sequences covering all functional gene variation types; assigning clinical priority weights to clinically frequently detected strains, locally prevalent strains, and highly pathogenic strains, and retaining their strains separately and marking their clinical attributes; establishing an association index between redundant sequences and representative sequences to support reversible traceability of redundant sequences.
[0011] According to an embodiment of the present invention, the core library consists of the RefSeq complete genome and pan-genome, used for clinical examination and low false positive scenarios; the extended library consists of supplementary data from GenBank, NT library and professional databases, used for scientific research and pathogen monitoring; the hierarchical database also includes an application library, which is a customized database for specific disease scenarios; the pan-genome model includes a core gene set and a variable gene set, wherein the core gene set consists of gene sequences common to all strains of the same species, and the variable gene set consists of gene sequences unique to some strains within the species.
[0012] According to an embodiment of the present invention, the clinical priority-oriented query system includes: a multi-dimensional retrieval architecture that supports retrieval by combination of at least one of the following conditions: species name, Tax ID, and sequence ID; a dynamic index structure that adopts a differentiated construction strategy for different levels of databases; and a genome-clinical phenotype bidirectional linkage function, including: generating association tags by positively associating pathogen genome features with clinical diagnosis and treatment data, and establishing a clinical phenotype index library to support reverse retrieval of genome features based on clinical phenotypes.
[0013] According to an embodiment of the present invention, the dynamic index structure includes: a Minimizer+BWT hybrid index for the core library; a compressed hierarchical index for the extended library, which performs genomic differential encoding on low-access-frequency sequences to construct a primary index-secondary index; and a "disease-gene panel association index" for the application library; the dynamic index structure updates only the index corresponding to newly added / changed sequences during incremental database updates.
[0014] According to an embodiment of the present invention, in the bidirectional linkage function of genome-clinical phenotype: the generation of the association tag includes: associating pathogen genome features with clinical diagnosis and treatment data to generate a "genome-clinical association tag" containing treatment recommendations; the association tag is graded according to the strength of evidence and dynamically corrected as the database is updated; the clinical phenotype index library supports inputting clinical phenotypes to retrieve corresponding genome features and strain sequences.
[0015] According to embodiments of the present invention, the contaminated sequence filtering includes masterless sequence removal, host background sequence removal, and low-abundance noise sequence filtering. Furthermore, the autonomous clinical data acquisition also adopts a multi-center data standardization access and privacy enhancement framework, including: providing a unified API interface and data upload tools to support automatic parsing and format conversion of various sequencing data formats; implementing at least two levels of privacy protection, with primary removal of direct identifiers and advanced processing of quasi-identifiers using differential privacy technology; and requiring access data to pass clinical relevance verification and sequence quality verification.
[0016] According to an embodiment of the present invention, the automated update mechanism is a periodically executed incremental update, and the emergency inclusion of emerging pathogens includes: performing quality prediction on emerging pathogen sequences and generating quality scores through a deep learning model; determining their taxonomic status through rapid k-mer alignment and preliminary phylogenetic tree construction; and for high-risk emerging pathogens that meet the quality score criteria, completing rapid entry into the database and index construction through an emergency update green channel.
[0017] In a second aspect, this invention proposes a pathogenic microorganism genome reference database. According to embodiments of the invention, the pathogenic microorganism genome reference database is constructed using the methods described above. The database proposed by this invention is comprehensive, hierarchically structured, of controllable quality, and suitable for clinical and research use, possessing excellent capabilities for pathogen detection, classification, comparison, tracking, and risk monitoring.
[0018] The beneficial effects of this invention are at least as follows: The pathogenic microorganism database proposed in this invention is comprehensive, hierarchically structured, and of controllable quality, possessing excellent capabilities for pathogen detection, classification, comparison, tracking, and risk monitoring. Attached Figure Description
[0019] Figure 1 This is a classification diagram of viral respiratory pathogens according to an embodiment of the present invention.
[0020] Figure 2 This is a classification diagram of fungal respiratory pathogens according to an embodiment of the present invention.
[0021] Figure 3 This is a classification diagram of bacterial respiratory pathogens according to an embodiment of the present invention.
[0022] Figure 4 This is a classification diagram of atypical respiratory pathogens according to an embodiment of the present invention.
[0023] Figure 5 This is a classification diagram of enterovirus pathogens according to an embodiment of the present invention.
[0024] Figure 6This is a diagram illustrating the classification of intestinal bacterial pathogens according to an embodiment of the present invention. Detailed Implementation
[0025] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0026] The endpoints and any values of the ranges disclosed herein are not limited to the precise ranges or values, and these ranges or values should be understood to include values close to these ranges or values. For numerical ranges, the endpoint values of the various ranges, the endpoint values of the various ranges and individual point values, and individual point values can be combined with each other to obtain one or more new numerical ranges, which should be considered as specifically disclosed herein.
[0027] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0028] To facilitate understanding of this invention, certain technical and scientific terms are specifically defined below. Unless explicitly defined elsewhere in this document, all other technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. In the description of this invention, the terms used herein have been explained and described; these explanations and descriptions are merely for the purpose of facilitating understanding and should not be construed as limiting the scope of protection of this invention.
[0029] In this article, the term "de-redundancy" refers to the process of identifying and removing duplicate or highly similar sequences in bioinformatics analysis.
[0030] In this article, the term "mapping" refers to the process of locating reads generated by sequencing onto a reference genome using sequence alignment algorithms, and determining their most likely origin.
[0031] In this article, the term "pangenome" refers to the sum of genomic information of all individuals within a species, including the core genome and the accessory genome. The core genome is the set of genes present in all tested strains / individuals and typically performs housekeeping functions (such as replication, transcription, translation, and basal metabolism). It is relatively conserved in number and represents the "essential genetic backbone" of the species. The accessory genome is the set of genes present only in some strains, including plasmids, phage fragments, transposons, resistance islands, virulence islands, and secondary metabolic synthesis clusters. Its content rapidly increases or decreases with changes in ecological niche or selective pressure and is the main genetic source of phenotypic differences (pathogenicity, drug resistance, environmental adaptation) among strains.
[0032] In this article, the terms "metagenomic sequencing" or "mNGS" refer to the use of high-throughput sequencing technology to obtain the sequence data of all microbial genomes in a sample, and then using bioinformatics analysis methods to perform sequence alignment to identify the types and abundance of microorganisms.
[0033] The technical solution of this application is described in detail below: Databases and their construction methods In some embodiments of the present invention, the present invention proposes a method for constructing a pathogenic microorganism genome reference database. According to an embodiment of the present invention, the method includes the following steps: (1) acquiring multi-source pathogen sequence data from multiple public databases and autonomous clinical data, integrating them according to preset data source priority rules, and achieving semantic uniformity and format compatibility of multi-source data through a cross-database data standardization mapping system; (2) cleaning and quality control the integrated data to obtain cleaned data; (3) based on the cleaned data, using a sequence clustering algorithm combined with functional gene integrity verification to perform function-oriented redundancy removal and standardized naming to obtain a standardized sequence dataset. Standardization can improve the consistency of the database; (4) based on the standardized sequence dataset, constructing a pan-genome model, and building a hierarchical database based on the pan-genome model. The hierarchical database includes at least a core database and an extended database; (5) based on the hierarchical database, constructing a clinical priority-oriented query system; and (6) automatically and dynamically updating the hierarchical database to achieve the emergency inclusion of emerging pathogens. The pathogenic microorganism genome reference database constructed according to the method described in this invention is comprehensive, hierarchically structured, and of controllable quality, and can be used in clinical and research scenarios. At the same time, the accompanying query system and knowledge service module can enhance the ability to detect, classify, compare, track, and monitor pathogens.
[0034] According to embodiments of the present invention, the pathogenic microorganisms include bacteria, fungi, viruses, protozoa, and parasites.
[0035] According to an embodiment of the present invention, the method further includes the following steps in step (1): establishing a unified mapping table for metadata fields, mapping the core fields of different databases to a unified name, data type and value range; based on the pathogenic microorganism ontology terminology database, using natural language processing (NLP) algorithms combined with homologous sequence alignment, unifying the annotation rules of different databases and correcting annotation contradictions; when there is information conflict in multi-source data, automatically adjudicating according to the rule of "data source weight + clinical evidence priority", and retaining high-credibility data.
[0036] According to embodiments of the present invention, the public database includes: RefSeq, GenBank, NT database, and specialized databases; the specialized databases include at least one of LPSN, ICTV, GISAID, MycoBank, etc. These databases contain high-quality, manually verified complete genomes; GenBank supplements species missing and low-coverage groups; the specialized databases provide typing and classification information. The LPSN (List of Prokaryotic Names with Standing in Nomenclature) provides hierarchical information on prokaryotes, including domain, kingdom, phylum, class, order, family, genus, and species. The ICTV (International Committee on Taxonomy of Viruses) is a viral classification system, providing a "Master Species List" (MSL), a 15-level taxonomic tree, exemplary viral strain information, viral metadata resources (VMR), and naming rules, covering DNA / RNA viruses, satellite nucleic acids, viroids, etc. The GISAID (Global Initiative on Sharing All Influenza Data) database primarily contains influenza viruses, now covering multiple viruses including COVID-19, monkeypox, and avian influenza. Its core libraries include EpiCoV™ and EpiFlu™, providing viral genome sequences, sampling time / location, patient clinical and epidemiological metadata, and online comparison and evolutionary analysis tools. MycoBank (Fungal Bank) includes information on new and old fungal names, namers, publications, and type specimens, and is linked to fungal culture collections, serving as an authoritative reference for fungal taxonomy research. The data sources are prioritized as follows: RefSeq, GenBank, NT database, specialized databases, and proprietary clinical sequencing data.
[0037] According to an embodiment of the present invention, in step (3) of the method, the redundancy removal threshold of the sequence clustering algorithm is a sequence similarity ≥ 95%. Extremely large-scale genomic data often contain a large number of highly similar repetitive sequences. Directly using these sequences for metagenomic sequencing data can lead to increased false positives, cross-alignment between species, and a significant decrease in processing speed. Therefore, the present invention uses CD-HIT (threshold ≥ 95 similarity) to remove redundant sequences and retain the longest representative sequence, achieving a balance between lightweight design and accuracy.
[0038] According to an embodiment of the present invention, the function-oriented deredundancy processing further includes: performing core functional gene integrity verification on each cluster and retaining representative sequences covering all functional gene variation types; assigning clinical priority weights to clinically frequently detected strains, locally prevalent strains, and highly pathogenic strains, and retaining their strains separately and marking their clinical attributes; establishing an association index between redundant sequences and representative sequences to support reversible traceability of redundant sequences.
[0039] According to an embodiment of the present invention, the core library consists of the RefSeq complete genome and pan-genome, used for clinical testing and low false positive scenarios; the extended library consists of supplementary data from GenBank, the NT library, and professional databases, used for scientific research and pathogen monitoring; the hierarchical database also includes an application library, which is a customized database for specific disease scenarios, capable of rapid detection, customized scenarios, and specific disease monitoring; the pan-genome model includes a core gene set and a variable gene set, wherein the core gene set consists of gene sequences common to all strains of the same species, and the variable gene set consists of gene sequences unique to some strains within the species.
[0040] According to an embodiment of the present invention, the clinical priority-oriented query system includes: a multi-dimensional retrieval architecture that supports retrieval by combination of at least one of the following conditions: species name, Tax ID, and sequence ID; a dynamic index structure that adopts a differentiated construction strategy for different levels of databases; and a genome-clinical phenotype bidirectional linkage function, including: generating association tags by positively associating pathogen genome features with clinical diagnosis and treatment data, and establishing a clinical phenotype index library to support reverse retrieval of genome features based on clinical phenotypes.
[0041] According to an embodiment of the present invention, the dynamic index structure includes: a Minimizer+BWT hybrid index for the core library; a compressed hierarchical index for the extended library, which performs genomic differential encoding on low-access-frequency sequences to construct a primary index-secondary index; and a "disease-gene panel association index" for the application library; the dynamic index structure updates only the index corresponding to newly added / changed sequences during incremental database updates.
[0042] According to an embodiment of the present invention, in the bidirectional linkage function between genome and clinical phenotype: the generation of the association tag includes: associating pathogen genome features with clinical diagnosis and treatment data to generate a "genome-clinical association tag" containing treatment recommendations; the association tag is graded according to the strength of evidence and dynamically corrected with database updates; the clinical phenotype index library supports reverse retrieval of corresponding genome features and strain sequences by inputting a clinical phenotype. The database construction method of the present invention also provides enhanced functions such as clinical knowledge linkage, subtype morphology display, asynchronous task scheduling, and security isolation mechanisms.
[0043] According to embodiments of the present invention, the contaminated sequence filtering includes masterless sequence removal, host background sequence removal, and low-abundance noise sequence filtering. Furthermore, the autonomous clinical data acquisition also adopts a multi-center data standardization access and privacy enhancement framework, including: providing a unified API interface and data upload tools to support automatic parsing and format conversion of various sequencing data formats; implementing at least two levels of privacy protection, with primary removal of direct identifiers and advanced processing of quasi-identifiers using differential privacy technology; and requiring access data to pass clinical relevance verification and sequence quality verification.
[0044] According to an embodiment of the present invention, the automated update mechanism is a periodically executed incremental update, and the emergency inclusion of emerging pathogens includes: performing quality prediction on emerging pathogen sequences and generating quality scores through a deep learning model; determining their taxonomic status through rapid k-mer alignment and preliminary phylogenetic tree construction; and for high-risk emerging pathogens that meet the quality score criteria, completing rapid entry into the database and index construction through an emergency update green channel.
[0045] In some embodiments of the present invention, a pathogenic microorganism genome reference database is proposed. According to embodiments of the present invention, the pathogenic microorganism genome reference database is constructed using the methods described above. The database proposed by the present invention is comprehensive, hierarchically structured, quality-controllable, and applicable to clinical and research settings, possessing excellent capabilities for pathogen detection, classification, comparison, tracking, and risk monitoring. The database proposed by the present invention is a high-precision integrated database system for multi-source microbial sequencing data. Through unified data inclusion standards, cross-platform data fusion strategies, hierarchical indexing structures, and dynamic update mechanisms, this database achieves comprehensive integration of microbial genome, pathogen characteristics, drug resistance, and virulence factor data. Furthermore, the database proposed by the present invention can be applied to various high-throughput sequencing data analysis software, providing a comprehensive and accurate data foundation for data analysis.
[0046] Embodiments of the present invention will now be described in more detail, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the invention, and should not be construed as limiting the invention.
[0047] Example 1: Method for constructing reference sequences for respiratory pathogens This study constructed a pathogen detection spectrum by systematically screening frequently detected common pathogens in the population using national regulations, international and domestic laboratory testing guidelines, previous literature evidence, and real-world clinical sample testing data. The spectrum primarily included the "Catalogue of Human Infectious Pathogens" (2023 edition), relevant CLSI (Clinical and Laboratory Standards Institute) guidelines, and domestic consensus / guidelines on clinical laboratory testing and infectious diseases (2020–2022). More than 20 common pathogens were selected, including: viruses such as influenza virus, respiratory syncytial virus, adenovirus, and human metapneumovirus; bacteria such as Streptococcus pneumoniae, Haemophilus influenzae, and Staphylococcus aureus; atypical pathogens such as Mycoplasma pneumoniae and Chlamydia trachomatis; and fungi such as Aspergillus and Candida. Respiratory pathogen information was extracted from the following databases: NCBI RefSeq (for high-quality genomes), NCBI Taxonomy (for classification and nomenclature information), GenBank (for supplementing variant sequences), and GISAID (for supplementing epidemic strain data for some viruses such as influenza virus and RSV). A total of 34 sequences of 7 viral respiratory pathogens, 6 sequences of 3 fungal pathogens, 346 sequences of 7 bacterial respiratory pathogens, and 10 sequences of 3 atypical cytokine pathogens were collected.
[0048] To improve database consistency, this embodiment performs the following standardized steps on the data during the data collection process: (I) Standardization process for species naming and classification information Step S1: Collection of original species information Obtain original species information of pathogenic microorganisms from public databases (including but not limited to NCBI Taxonomy, RefSeq, GenBank) and authoritative catalogs, including Latin names, synonyms and historical naming records.
[0049] Step S2: Verification of Latin scientific names Using the taxonomic ID (TaxID) in the NCBI Taxonomy database as a unique identifier, different naming forms of the same species are mapped and merged to ensure that each pathogen corresponds to only one standard Latin name.
[0050] Step S3: Standardized mapping of Chinese and English names Based on national standards, industry norms, and commonly used medical literature, a one-to-one correspondence between Latin scientific names and their Chinese and English names is established to eliminate ambiguity caused by synonyms, multiple translations, or abbreviations.
[0051] Step S4: Taxonomic Hierarchy Completion and Verification For each pathogen, complete taxonomic information should be provided, including at least the following levels: Kingdom, Phylum, Class, Order, Family, Genus, and Species. When there are inconsistencies in the classification of the database sources, the latest version of NCBI Taxonomy should be used for standardization.
[0052] (II) Structured Construction Process of Pathogen Data Units Step S5: Define the unified data unit template For each pathogen, a data object with a uniform format is constructed. The data object contains at least the following fields: Chinese name, English name, Latin name, taxonomic information, pathogen category (virus / bacteria / atypical pathogen / fungus), common detected sample types, including but not limited to: nasopharyngeal swabs, throat swabs, sputum, bronchoalveolar lavage fluid, blood, urine, body fluids, etc.
[0053] Step S6: Sample type information integration Based on clinical testing guidelines, literature statistics, and real-world testing data, the high-frequency detection sample types of different pathogens are summarized and stored using predefined enumeration values to ensure field comparability.
[0054] Step S7: Regulatory Attribute Labeling Based on the catalog of pathogenic microorganisms and related regulations issued by the National Health Commission, binary labeling is used to determine whether pathogens are included in the national monitoring system and whether they are highly pathogenic.
[0055] Step S8: Consistency and Integrity Verification Perform field integrity checks and logical consistency checks on the completed data units to ensure that all required fields are filled in and there are no conflicting information.
[0056] (III) Quality control standards and construction process of reference genome sets To ensure the accuracy, representativeness, and reliability of the reference genome set and subsequent alignment analysis, this embodiment implements a rigorous quality filtering and screening process for genome sequences from RefSeq and GenBank.
[0057] Step S1: Download the genome sequence Download the corresponding whole or near-whole genome sequences from the RefSeq and GenBank databases, categorized by pathogen species, prioritizing those labeled as... Complete Genome or Chromosome-level assembly sequence.
[0058] Step S2: Metadata association associates the source information of each genome sequence, including: species name and TaxID, sequencing platform, and assembly method.
[0059] Step S3: Genome quality control standards Integrity criteria, duplication and redundancy control, and contamination detection: CD-HIT (threshold ≥95% similarity) is used to remove redundant sequences and retain the longest representative sequence.
[0060] Step S4: Reference genome screening and deduplication Based on integrity and continuity criteria, low-quality assemblies are initially screened, and genomes that clearly do not meet the requirements are removed. Whole-genome similarity calculations are performed on candidate genomes within the same species, and they are clustered according to a similarity threshold. Only the genome with the highest quality score is retained as the representative genome in each cluster.
[0061] III. Pathogen Classification Display The database classifies pathogens based on multi-dimensional attributes. In this embodiment, pathogens are classified according to their source of infection into: viral respiratory pathogens, bacterial respiratory pathogens, atypical respiratory pathogens, and fungal respiratory pathogens, as detailed in the appendix. Figure 1 Appendix Figure 2 Appendix Figure 3 and attached Figure 4 As shown in Table 1, users can quickly locate pathogens through the classification list. The system also supports multi-condition searches, such as species name, sample type, and pathogen category.
[0062] Table 1 Overview of Respiratory Pathogens
[0063] Example 2: Method for constructing reference sequences for intestinal pathogens This study constructed a pathogen detection spectrum by systematically screening frequently detected common pathogens in the population using national regulations, international and domestic laboratory testing guidelines, previous literature evidence, and real-world clinical sample testing data. The spectrum primarily included the "Catalogue of Human Infectious Pathogens" (2023 edition), relevant CLSI (Clinical and Laboratory Standards Institute) guidelines, and domestic consensus / guidelines on clinical laboratory testing and infectious diseases (2020–2022). A total of 11 common pathogens were selected, including 4 viral enteric pathogens and 7 bacterial enteric pathogens. The viral enteric pathogens were rotavirus (11 sequences), norovirus (6 sequences), adenovirus (enteric type, 2 sequences), and astrovirus (1 sequence). Bacterial enteric pathogens include Escherichia coli (pathogenic, 225 sequences), Shigella (165 sequences), Salmonella (47 sequences), Vibrio cholerae (68 sequences), Campylobacter jejuni (22 sequences), Clostridium difficile (133 sequences), and Yersinia enterocolitica (74 sequences).
[0064] To improve database consistency, this embodiment performs the following standardized steps on the data during the data collection process: (II) Standardization Process for Species Naming and Classification Information Step S1: Collection of information on the original species Obtain original species information of pathogenic microorganisms from public databases (including but not limited to NCBI Taxonomy, RefSeq, GenBank) and authoritative catalogs, including Latin names, synonyms and historical naming records.
[0065] Step S2: Verification of Latin scientific names Using the taxonomic ID (TaxID) in the NCBI Taxonomy database as a unique identifier, different naming forms of the same species are mapped and merged to ensure that each pathogen corresponds to only one standard Latin name.
[0066] Step S3: Standardized mapping of Chinese and English names Based on national standards, industry norms, and commonly used medical literature, a one-to-one correspondence between Latin scientific names and their Chinese and English names is established to eliminate ambiguity caused by synonyms, multiple translations, or abbreviations.
[0067] Step S4: Taxonomic Hierarchy Completion and Verification For each pathogen, complete taxonomic information should be provided, including at least the following levels: Kingdom, Phylum, Class, Order, Family, Genus, and Species. When there are inconsistencies in the classification of the database sources, the latest version of NCBI Taxonomy should be used for standardization.
[0068] (II) Structured Construction Process of Pathogen Data Units Step S5: Define the unified data unit template For each pathogen, a data object with a uniform format is constructed. The data object contains at least the following fields: Chinese name, English name, Latin name, taxonomic information, pathogen category (virus / bacteria / atypical pathogen / fungus), common detected sample types, including but not limited to: nasopharyngeal swabs, throat swabs, sputum, bronchoalveolar lavage fluid, blood, urine, body fluids, etc.
[0069] Step S6: Sample type information integration Based on clinical testing guidelines, literature statistics, and real-world testing data, the high-frequency detection sample types of different pathogens are summarized and stored using predefined enumeration values to ensure field comparability.
[0070] Step S7: Regulatory Attribute Labeling Based on the catalog of pathogenic microorganisms and related regulations issued by the National Health Commission, binary labeling is used to determine whether pathogens are included in the national monitoring system and whether they are highly pathogenic.
[0071] Step S8: Consistency and Integrity Verification Perform field integrity checks and logical consistency checks on the completed data units to ensure that all required fields are filled in and there are no conflicting information.
[0072] (III) Quality control standards and construction process of reference genome sets To ensure the accuracy, representativeness, and reliability of the reference genome set and subsequent alignment analysis, this embodiment implements a rigorous quality filtering and screening process for genome sequences from RefSeq and GenBank.
[0073] Step S1: Download the genome sequence Download the corresponding whole or near-whole genome sequences from the RefSeq and GenBank databases according to the pathogen species, prioritizing sequences labeled as Complete Genome or Chromosome-level assembly.
[0074] Step S2: Metadata association associates the source information of each genome sequence, including: species name and TaxID, sequencing platform, and assembly method.
[0075] Step S3: Genome quality control standards Integrity criteria, duplication and redundancy control and contamination detection: Integrity criteria, duplication and redundancy control and contamination detection: CD-HIT (threshold ≥95% similarity) is used to remove redundant sequences and retain the longest representative sequence.
[0076] Step S4: Reference genome screening and deduplication Based on integrity and continuity criteria, low-quality assemblies are initially screened, and genomes that clearly do not meet the requirements are removed. Whole-genome similarity calculations are performed on candidate genomes within the same species, and they are clustered according to a similarity threshold. Only the genome with the highest quality score is retained as the representative genome in each cluster.
[0077] III. Pathogen Classification Display The database classifies pathogens based on multi-dimensional attributes. In this embodiment, pathogens are divided into viral enteric pathogens and bacterial enteric pathogens according to their source of infection. Specific results are shown in the appendix. Figure 5 and attached Figure 6 As shown. Users can quickly locate pathogens through the classification list. The system also supports multi-condition searches, such as species name, sample type, pathogen category, etc.
[0078] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0079] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for constructing a pathogenic microorganism genome reference database, characterized in that, Includes the following steps: (1) Acquire multi-source pathogen sequence data from multiple public databases and independent clinical data, integrate them according to the preset data source priority rules, and achieve semantic uniformity and format compatibility of multi-source data through a cross-database data standardization mapping system; (2) Clean and quality control the integrated data to obtain cleaned data; (3) Based on the cleaned data, a sequence clustering algorithm is used in combination with functional gene integrity verification to perform function-oriented redundancy removal and standardized naming to obtain a standardized sequence dataset; (4) Based on the standardized sequence dataset, construct a pan-genome model and build a hierarchical database based on the pan-genome model. The hierarchical database includes at least a core library and an extended library. (5) Based on the hierarchical database, construct a clinical priority-oriented query system; and (6) The hierarchical database is automatically and dynamically updated to enable the emergency inclusion of emerging pathogens.
2. The method according to claim 1, characterized in that, In step (1), The public databases include at least one of NCBI RefSeq, GenBank, NT database, PATRIC, National Pathogenic Microorganism Resource Center, ICTV, GISAID, LPSN, and MycoBank; The autonomous clinical sequencing data comes from the whole genome sequences of clinical pathogens, raw mNGS sequencing data, and de-identified patient clinical information from medical institutions. Use the rsync or datasets tools to automate the downloading and incremental updates of data from public databases.
3. The method according to claim 1, characterized in that, Step (1) further includes: Establish a unified mapping table for metadata fields, mapping the core fields of different databases to a unified name, data type, and value range; Based on a pathogenic microorganism ontology terminology database, annotation rules from different databases are unified and annotation inconsistencies are corrected by combining natural language processing (NLP) algorithms with homology sequence alignment. When there are conflicting information from multiple data sources, the system automatically decides based on the rule of "data source weight + clinical evidence priority", retaining the data with high credibility.
4. The method according to claim 1, characterized in that, In step (3), the redundancy removal threshold of the sequence clustering algorithm is a sequence similarity ≥ 95%; The function-oriented redundancy removal process further includes: For each cluster, the integrity of the core functional genes is verified, and representative sequences covering all functional gene variation types are retained; Clinically frequently detected strains, locally prevalent strains, and highly pathogenic strains are assigned clinical priority weights, and their strains are retained separately and labeled with clinical attributes. Establish an association index between redundant sequences and representative sequences to support reversible tracing of redundant sequences.
5. The method according to claim 1, characterized in that, The core library consists of the RefSeq complete genome and pangenome, and is used for clinical examination and low false positive scenarios. The extended library consists of supplementary data from GenBank, the NT library, and specialized databases, and is used for scientific research and pathogen monitoring. The hierarchical database also includes an application library, which is a customized database for specific disease scenarios; The pan-genome model includes a core gene set and a variable gene set. The core gene set consists of gene sequences common to all strains of the same species, while the variable gene set consists of gene sequences unique to some strains within the species.
6. The method according to claim 1, characterized in that, The clinical priority-oriented query system includes: A multi-dimensional search architecture that supports searches based on at least one of the following conditions: species name, Tax ID, and sequence ID; A dynamic index structure, wherein the dynamic index structure adopts a differentiated construction strategy for different levels of database; The bidirectional linkage function between genome and clinical phenotype includes: generating association tags by positively linking pathogen genome features with clinical diagnosis and treatment data, and establishing a clinical phenotype index library to support reverse retrieval of genome features from clinical phenotypes.
7. The method according to claim 6, characterized in that, The dynamic index structure includes: Minimizer+BWT hybrid index for the core library; For the compressed hierarchical index of the extended library, genomic differential encoding is performed on low-access-frequency sequences to construct a primary index-secondary index; For the application library's "disease-gene panel association index"; The dynamic index structure only updates the index corresponding to the newly added / changed sequence during incremental database updates.
8. The method according to claim 6, characterized in that, In the bidirectional linkage function between genome and clinical phenotype: The generation of the association tags includes: associating pathogen genomic characteristics with clinical diagnosis and treatment data to generate "genome-clinical association tags" containing treatment recommendations; the association tags are graded according to the strength of evidence and are dynamically corrected as the database is updated; The clinical phenotype index library supports reverse retrieval of corresponding genomic features and strain sequences by inputting a clinical phenotype.
9. The method according to claim 1, characterized in that, Contaminated sequence filtering includes hostless sequence removal, host background sequence removal, and low-abundance noise sequence filtering. The autonomous clinical data collection also employs a multi-center data standardization access and privacy enhancement framework, including: It provides a unified API interface and data upload tools, supporting automatic parsing and format conversion of various sequencing data formats; At least two levels of privacy protection are implemented: primary protection removes direct identifiers, and advanced protection uses differential privacy technology to process quasi-identifiers; accessed data must pass clinical relevance verification and sequence quality verification. Optionally, the automated update mechanism is a periodically executed incremental update; Optionally, the emergency inclusion of the emerging pathogens includes: Quality prediction and quality score generation for emerging pathogen sequences are performed using deep learning models. Its taxonomic position was determined by rapid k-mer alignment and preliminary phylogenetic tree construction; and For high-risk emerging pathogens that meet the quality score criteria, rapid data entry and index construction are completed through an emergency update green channel.
10. A pathogenic microorganism genome reference database, characterized in that, The pathogenic microorganism genome reference database is constructed by the method described in claims 1 to 9.