Knowledge base system and method for pathogen detection and storage medium

By constructing a pathogen detection knowledge base system that integrates pathogen, drug resistance, and virulence information, and providing a localized background microbial library and expert review mechanism, the problems of high false positive rate and difficulty in interpreting NGS results have been solved, enabling rapid and accurate pathogen detection.

CN120932791APending Publication Date: 2025-11-11SHANGHAI CINOPATH MEDICAL TESTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511054420.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies, such as NGS, suffer from high false positive rates, difficulty in interpreting results, lack of information on drug resistance and virulence, and fragmented and outdated databases, which cannot meet the demand for rapid and accurate pathogen detection.

Method used

A pathogen detection knowledge base system is constructed, including a front-end display and a back-end maintenance subsystem. It integrates pathogen, drug resistance, and virulence information, adopts an expert review and dynamic maintenance mechanism, provides a localized background microbial library, and generates structured reports through intelligent retrieval.

Benefits of technology

Significantly reduces the false positive rate of NGS, improves detection specificity, enhances the efficiency of clinical decision-making, ensures the authority and timeliness of data, covers all pathogens and drug needs, and reduces the time required to interpret results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932791A_ABST
    Figure CN120932791A_ABST
Patent Text Reader

Abstract

The invention provides a set of pathogen detection knowledge base system, which integrates five modules of pathogen, drug resistance, toxicity, drug and localized background bacteria, is matched with an expert auditing and task-driven background dynamic maintenance mechanism, and provides intelligent retrieval and visual display. According to the system, through BioBERT literature mining and experimental and clinical data cross validation, seven-level hazard grading and a localized background bacterium library are established, and NGS false positive is remarkably reduced; 62 types of drugs are in one-key association with drug resistance mechanisms and virulence genes of 150 specific drugs, a structured report is automatically generated, and accurate drug use is directly guided. A knowledge base solution which is low in false positive, high in report efficiency, real-time and authoritative in data, comprehensive in coverage and intelligent in interpretation is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pathogen detection technology, and in particular to a knowledge base system, method and storage medium for pathogen detection. Background Technology

[0002] More than 1,000 microorganisms are known to cause infectious diseases in humans, and early and accurate pathogen identification is crucial for determining clinical efficacy and prognosis. However, traditional culture methods are time-consuming and have low positive rates, making it difficult to meet the urgent need for "rapid diagnosis and precise medication" in severe infections. With the development of next-generation sequencing (NGS) technology, metagenomic sequencing (mNGS) can now complete broad-spectrum pathogen screening within 24 hours, becoming an important supplementary means of diagnosing infectious diseases. Despite the advantages of NGS being broad-spectrum and unbiased, clinical practice still faces the following core challenges: (1) High false positive rate: The sources of contamination are complex, including ① environmental bacteria introduced during the specimen collection process, ② normal human flora, and ③ laboratory reagent and engineered bacteria residues. Relying solely on laboratory quality control cannot completely eliminate background interference, leading to a decrease in the reliability of the reported results; (2) Difficulty in interpreting results: NGS outputs a large list of microorganisms, but lacks a unified and authoritative pathogenicity classification standard; clinicians find it difficult to determine whether "detection means pathogenicity" or "background colonization"; (3) Lack of information on drug resistance and toxicity: Existing public databases (such as CARD and VFDB) focus on basic research and lack annotation systems that are directly related to clinical drug selection and toxicity risk assessment, which cannot directly guide precision medicine. (4) Urgent need for localization: The background bacterial spectrum varies significantly between different hospitals and regions, and the public database lacks local background bacterial bank support, which further exacerbates the false positive problem.

[0003] The public health pressure of drug resistance and virulence surveillance is immense. Antimicrobial overuse has led to a year-on-year increase in the infection rate of multidrug-resistant bacteria (such as carbapenemase-producing bacteria, MRSA, and VRE), posing significant challenges to clinical treatment. Virulence factors are the decisive factors in pathogen virulence, and their genotypes are closely related to infection severity and outbreak risk. Currently, clinical laboratories mainly rely on single-gene PCR or routine antimicrobial susceptibility testing for the detection of drug resistance / virulence genes, which has low throughput, is time-consuming, and lacks a real-time surveillance platform integrated with epidemiological data.

[0004] The existing technical solutions have the following technical defects: (1) Database fragmentation: Pathogen information, drug resistance mechanisms, virulence genes, and drug knowledge are scattered across multiple independent databases, lacking unified coding and cross-indexing, resulting in low retrieval efficiency; (2) Lack of dynamic update mechanism: The public database has a long update cycle and has not established a closed-loop maintenance process of "data collection and editing - expert review - version control", which leads to the lag in clinical interpretation. (3) No local background bacteria filtering function: The existing NGS analysis process only relies on a general contamination list and cannot dynamically adjust the background bacteria spectrum according to the actual hospital environment, consumables and population colonization characteristics; (4) Unstructured output of results: The reports are mostly presented in the form of long lists, lacking integrated prompts of "pathogen-drug resistance-virulence-drug" that match the clinical scenario, increasing the decision-making burden of doctors.

[0005] To overcome the aforementioned shortcomings, there is an urgent need to construct a pathogen detection knowledge base system that is "comprehensive in data, timely in updates, locally adapted, and clinically oriented," capable of achieving: (1) Integrate multi-dimensional information on pathogens, drug resistance, virulence, drugs, and background bacteria; (2) Provide a localized background microbial library to reduce NGS false positives; (3) Establish an expert review and dynamic maintenance mechanism to ensure the authority and timeliness of the data; (4) Through standardized coding and intelligent retrieval, it directly supports the automatic generation of clinical reports and precise treatment decisions. Summary of the Invention

[0006] This invention proposes a knowledge base system and method for pathogen detection, solving the problems mentioned in the background art. The technical solution of this invention is implemented as follows: A knowledge base system for pathogen detection includes: a display front-end, a back-end maintenance subsystem, and a retrieval module; wherein, The front-end presentation should include at least: The pathogen information module is used to present basic information about pathogenic microorganisms, pathogenicity assessment results, and related drug resistance and virulence data. The anti-infective drugs module is used to present drug categories, drug names, resistance mechanisms, resistance genes, and sources of evidence. The drug resistance and virulence module is used to present drug resistance genes, virulence genes, their associated pathogens, clinical interpretations, and sources of evidence. The background pathogen module is used to present the Chinese name, Latin name, colonization location, common sample types corresponding to the colonization location, and the background pathogen type of the background bacteria. The backend maintenance subsystem includes at least: The account role management module is used to assign different permissions to data management allocators, maintainers, audit experts, and users. The knowledge base data maintenance task allocation module is used to create, distribute, prioritize, and track data maintenance tasks. The task receiving and execution management module is used by maintainers and review experts to receive, edit, submit, and review tasks; The system information update display module is used to publish and display knowledge base update announcements; The search module is used for: Receive keywords input by the user; Keywords are standardized based on biomedical entity recognition and synonym expansion. Parallel retrieval modules for pathogen information, anti-infective drugs, drug resistance virulence, and background pathogens; The search results are aggregated, sorted, and returned to the front end for display based on relevance, authority, and timeliness.

[0007] Furthermore, the pathogenicity assessment results in the pathogen information module are generated in the following ways: pathogenicity data are extracted from published literature using a natural language processing model; laboratory virulence factor assays and infection model phenotypic data are integrated; cross-validation is performed by combining clinical epidemiological surveillance and case report data; key data are manually reviewed by experts and data traceability information is recorded.

[0008] Furthermore, the pathogen information module classifies the severity of pathogens into seven levels: Category 1: Infectious pathogens that require direct reporting; Category 2: Pathogens of key clinical concern; Category 3: Clinically pathogenic pathogens; Category 4: Opportunistic pathogens; Category 5: Colonizing / background bacteria; Category 6: Pathogenicity unknown; Category 7: Non-pathogenic pathogens.

[0009] Furthermore, the resistance mechanisms in the resistance virulence module include at least one of the following mechanisms: alteration of antimicrobial drug target sites, target substitution, target protection, antibiotic inactivation, efflux, decreased permeability, altered cell morphology, altered host-dependent nutrient acquisition, target overexpression, and resistance due to gene deletion.

[0010] Furthermore, the background pathogen module is also used to perform the following tasks: distinguish background bacteria from laboratory environment and consumables; distinguish background bacteria from normal colonizing bacteria from different parts of the human body; and deduct or indicate corresponding background bacteria interference in the pathogen detection report.

[0011] Furthermore, the retrieval module employs an asynchronous concurrent or distributed task framework to query the pathogen information module, anti-infective drug module, drug resistance virulence module, and background pathogen module in parallel, and displays the retrieval results on the front end using a structured template, supporting the following operations: (1) Click on the drug resistance gene or virulence gene to jump to detailed information; (2) Visualize the virulence gene regulatory network based on the Neo4j map.

[0012] Furthermore, the task allocation module of the background maintenance subsystem supports the following operations: (1) Set priority for data entries to be added or updated; (2) Real-time statistics and display of task statuses: “Unassigned, Pending Start, Pending Review, Review Failed, Completed”; (3) Allow data management assigners to dynamically adjust task priorities and monitor the work progress of maintenance personnel and audit experts.

[0013] Furthermore, the task receiving and execution management module includes: The maintainer account management submodule is used to edit, modify, and submit task data; The Expert Account submodule is used for quality checks, returning items for revisions, or confirming approval. The system message notification function submodule is used to notify users of task status updates and data update information in real time.

[0014] A method for generating pathogen detection reports using a knowledge base system includes the following steps: Step S1: Receive NGS sequencing data of the sample to be tested; Step S2: Call the retrieval module to retrieve pathogen, drug resistance, and virulence information from the detected microorganism list; Step S3: Subtract background bacteria interference based on the background pathogen module; Step S4: Generate a test report by integrating the pathogenicity assessment results, drug resistance indications, and virulence factor information of the pathogen.

[0015] A non-transitory storage medium for storing a program used in a knowledge base system for pathogen detection to perform the following action: executing a method for generating pathogen detection reports based on the knowledge base system.

[0016] Compared with existing technologies, this solution has the following advantages: (1) Reduce NGS false positives: By accurately filtering out contaminants and colonizing bacteria through a localized background pathogen library, the specificity of detection is significantly improved; (2) Improve clinical decision-making efficiency: Integrate four-dimensional data on pathogens, drug resistance, virulence and drugs, generate structured reports with one click, and directly guide precision medication; (3) Dynamic and authoritative data support: The expert review and closed-loop maintenance mechanism ensures that the knowledge base is updated in real time, solving the problem of lagging behind traditional databases; (4) Comprehensive coverage of testing needs: Covering all pathogens such as bacteria, viruses, and fungi, as well as 62 types of anti-infective drugs, to meet the needs of complex infection scenarios; (5) Intelligent retrieval and visualization: Supports synonym expansion and Neo4j graph visualization of toxicity networks, significantly reducing the time cost of interpreting results. Dynamic authoritative data support: Expert review and closed-loop maintenance mechanism ensures real-time updates of the knowledge base, solving the problem of lag in traditional databases. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is an overall architecture diagram of a knowledge base system for pathogen detection according to the present invention; Figure 2 This is a key component module of a pathogen detection knowledge base system of the present invention; Figure 3 This is a diagram illustrating the backend maintenance framework of a knowledge base system for pathogen detection according to the present invention. Figure 4 This is the technical framework for implementing a pathogen detection retrieval module of the present invention; Figure 5 This is a screenshot of the front-end interface of the pathogen knowledge base system module. Figure 6 This is a screenshot of the front-end display interface for anti-infective drugs; Figure 7 Diagram showing the front end of drug resistance and virulence genes; Figure 8 For the pathogen information task list of the knowledge base system; Figure 9 This is a retrieval template for the knowledge base of pathogen detection. Detailed Implementation

[0019] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0020] Reference Figures 1-4This invention provides a knowledge base system for pathogen detection, comprising: a display front-end, a back-end maintenance subsystem, and a retrieval module; wherein the display front-end includes at least: a pathogen information module, used to present basic information of pathogenic microorganisms, pathogenicity assessment results, and their associated drug resistance and virulence data; an anti-infective drug module, used to present drug categories, drug names, drug resistance mechanisms, drug resistance genes, and their sources of evidence; a drug resistance and virulence module, used to present drug resistance genes, virulence genes, their associated pathogens, clinical interpretations, and sources of evidence; and a background pathogen module, used to present the Chinese name, Latin name, colonization location, common sample types corresponding to the colonization location, and the types of background bacteria.

[0021] The backend maintenance subsystem includes at least the following modules: an account role management module, used to assign different permissions to data management allocators, maintainers, review experts, and users; a knowledge base data maintenance task allocation module, used to create, distribute, prioritize, and track data maintenance tasks; a task receiving and execution management module, used for maintainers and review experts to receive, edit, submit, and review tasks; and a system information update display module, used to publish and display knowledge base update announcements.

[0022] The retrieval module is used to: receive keywords input by the user; standardize the keywords based on biomedical entity recognition and synonym expansion; retrieve pathogen information, anti-infective drugs, drug resistance, and background pathogens in parallel; and aggregate, sort, and return the retrieval results to the display front end according to relevance, authority, and timeliness.

[0023] Specifically, this pathogen knowledge base system, such as Figure 2 The diagram shows the framework, which mainly includes a front-end display module, a back-end maintenance module, and a user search module. The front-end modules respond to user page actions to display relevant information and data; the back-end maintenance module displays the data entered and maintained by the front-end modules; and the search module retrieves relevant information and data from the front-end display based on input keywords.

[0024] The front-end display of the pathogen knowledge base system is as follows: Figure 5 As shown, the system includes modules for pathogens, anti-infective drugs, drug-resistant virulence, and background pathogens. Additionally, the front-end display includes announcements, downloads, help, and latest tasks. The announcements module displays announcements; the download module displays downloadable data and provides download services; the help module displays the system user manual; and the latest tasks are the account's task notifications. The pathogen information primarily includes basic information on bacteria, viruses, fungi, parasites, and other special pathogens, as well as associated drug resistance and virulence data. Pathogen information mainly includes Chinese name, Latin name, taxonomic position, genus name (Chinese), pathogen code, classification, brief description, severity, and specific pathogenicity.

[0025] The code serves as a unique identifier for the pathogen, existing uniquely in the knowledge base. Once the name is updated or modified, the code will not be reused. The severity levels are categorized as follows: Category 1: Infectious pathogens that cannot be directly reported; Category 2: Pathogens of clinical concern; Category 3: Clinically pathogenic pathogens; Category 4: Opportunistic pathogens; Category 5: Colonizing / background bacteria; Category 6: Pathogenicity unknown; Category 7: Non-pathogenic.

[0026] Establishing a pathogenicity database is a systematic project involving multiple disciplines. The main methods and model framework are as follows: First, data collection methods, including literature mining: First, using Natural Language Processing (NLP) techniques to extract pathogenicity data from published literature and using pre-trained models such as BERT and BioBERT for biomedical text mining; Second, experimental data integration, using laboratory test results (such as virulence factor assays and infection model data) and phenotypic data (drug resistance, host range, pathogenic symptoms); Third, clinical data collection, including epidemiological surveillance data, case reports, and clinical research data. Second, data employs multiple validation mechanisms, including cross-validation of experimental validation and computational prediction, expert review of key data, and data source tracing.

[0027] For example, clicking on Pathogens > Bacteria > Klebsiella pneumoniae will display the interface as follows: Figure 9 As shown.

[0028] The front-end display interface of the drug module is as follows: Figure 6 As shown, the front end displays information including drug categories, drug names, the correspondence between categories and drug names, and the use of antibiotics.

[0029] The drug categories include 62 subcategories such as aminoglycosides, streptozotocins, fluoroquinolones, macrolides, lincomycins, β-lactams, tetracyclines, sulfonamides, penicillins, cephalosporins, glycopeptides, antituberculosis antibiotics, and antifungals. The document also includes definitions, unique codes, and brief descriptions of antibiotic use for each subcategory.

[0030] The list includes drug names such as gentamicin, tobramycin, amikacin, prazomicin, streptomycin, neomycin, paromomycin, roxithromycin, clarithromycin, azithromycin, telithromycin, quinephedrine, doxycycline, methacycline, minocycline, norfloxacin, ofloxacin, ciprofloxacin, levofloxacin, and moxifloxacin. Among these, 150 drugs are levofloxacin, moxifloxacin, lincomycin, clindamycin, chloramphenicol, thiamphenicol, vancomycin, norvancomycin, teicoplanin, and dapoxetine. A brief description and unique code for each drug are also provided.

[0031] There is a one-to-one correspondence between drug categories and drug names, and each database entry has a unique code. For example, the display content for the drug gentamicin is as follows: Drug Name: Gentamicin Drug code: DRU:0000014 Drug Description: Gentamicin works by binding to the 30S ribosomal subunit of bacteria, causing mRNA misreading and preventing bacteria from synthesizing proteins essential for their growth.

[0032] Drug category: Aminoglycosides

[0033] Category code: CLASS:0000016

[0034] Classification definition: Aminoglycosides are a class of antibiotics with an amino sugar and an aminocyclic alcohol structure, primarily effective against Gram-negative bacteria. Aminoglycosides work by binding to the 30S or 50S subunit of the bacterial ribosome, inhibiting the translocation of peptidyl tRNA from the A site to the P site, and also causing mRNA misreading, preventing bacteria from synthesizing proteins crucial for their growth.

[0035] Antibiotic Use: Aminoglycosides are a class of antibiotics used to treat serious bacterial infections, such as those caused by Gram-negative bacteria, particularly *Pseudomonas aeruginosa*. Aminoglycosides are poorly absorbed into the bloodstream after oral administration, so they are usually administered intravenously or sometimes intramuscularly. All aminoglycosides can cause hearing and kidney damage. Therefore, doctors closely monitor the dosage and, if possible, will usually choose an alternative type of antibiotic.

[0036] The front-end display interface for drug resistance and virulence genes is as follows: Figure 7 As shown, the front end displays information including drug resistance genes, drug resistance mutation sites, virulence genes, pathogens that the genes may be associated with, and the drug resistance mechanisms corresponding to the drug resistance genes.

[0037] The gene number serves as a unique identifier, along with the gene function and corresponding clinical annotations. Drug resistance or virulence genes are associated with potentially relevant pathogens. Drug resistance mechanisms include 10 mechanisms such as altered antimicrobial drug targets, antibiotic target substitution, antibiotic target protection, antibiotic inactivation, antibiotic efflux, decreased antibiotic permeability, altered cell morphology, host-dependent nutrient acquisition resistance, antibiotic target overexpression, and resistance due to gene deletion. Each mechanism corresponds to a unique data code, a brief description, and a specific drug resistance gene.

[0038] For example, the ErmB drug resistance gene is displayed as follows: Gene name: ErmB Gene ID: G3000375 Gene function: ErmB-encoded 23S rRNA methyltransferase leads to target site alteration, which is the main resistance mechanism to macrolide antibiotics.

[0039] Clinical Note: Carrying this gene can lead to resistance to macrolides, lincomycins, and streptomycins.

[0040] Related pathogens: Acinetobacter baumannii, Streptococcus mutans, Clostridium perfringens, Klebsiella pneumoniae, Salmonella enterica, Bacteroides fragilis, Escherichia coli, Campylobacter coli, Klebsiella pneumoniae, Streptococcus pneumoniae, Enterococcus faecalis, Shigella flexneri, Streptococcus Gordons, Lactococcus grescii, Streptococcus dolphinii, Streptococcus spp., Streptococcus pseudopneumococcus, Staphylococcus pseudointermediate, Streptococcus gallolyticus, Staphylococcus aureus, Lactobacillus cremastrae, Campylobacter jejuni, Bacillus subtilis, Geminidella mesenteriae, Streptococcus pyogenes, Enterococcus avium, Enterococcus faecalis, Shigella sonnei, Streptococcus dysgalactiae, Streptococcus agalactiae, Streptococcus pharyngitis, Streptococcus intermediate, Streptococcus suis, Lactobacillus indolentus, Fusobacterium nucleatum

[0041] Drug resistance mechanism: alteration of antibiotic target sites

[0042] Mechanism code: M02

[0043] Mechanism Overview: Mutations or enzymatic modifications of antibiotic targets lead to antibiotic resistance.

[0044] Background pathogens mainly fall into two categories. One category consists of unavoidable sources such as the environment and engineered bacteria residues, some of which can only be eliminated through consumable treatment or strict laboratory control. The other category comprises common colonizing pathogens from different parts of the sample itself. The accurate setting of these two background pathogen categories is crucial for the accuracy of pathogen detection reports. The background pathogen knowledge base mainly includes the Chinese name, Latin name, genus name (in Chinese), background bacteria type, colonization location, and common sample types corresponding to that location.

[0045] Backend maintenance includes account and role management, knowledge base data maintenance task allocation, task receiving and execution management, and system information update and display modules such as... Figure 3 As shown.

[0046] Account role management includes data management assigners, maintainers, review experts, and users. Data management assigners are users who issue and assign tasks; maintainers are users who add and complete tasks; review experts are users who perform quality checks on newly added data and confirm the accuracy of knowledge base information; and users are users who can view and search the knowledge base.

[0047] Knowledge base data maintenance task allocation module, such as Figure 8As shown, by logging into the administrator / assigner account, one can view the number of data items to be assigned, maintained, reviewed, and completed. When assigning tasks, this account can set task priorities, facilitating task execution by maintainers and reviewers. It can also view the work progress and completion status of each maintainer and review expert, aiding in subsequent task assignment. Task completion status includes: Unassigned, Pending Start, Pending Review, Review Failed, and Completed.

[0048] The task receiving module, after logging into the maintenance staff's account, displays the completion status of user roles, as well as the reasons for data rejection due to inadequacy, facilitating further modifications and submissions for review. When a user completes a task, the corresponding progress indicator light up, and the management account is promptly notified that the user's task is complete. System message notifications for updates to personal account data are also displayed. The maintenance staff account has task editing permissions for each task type, enabling them to edit tasks assigned by the task assigners.

[0049] In the pending review module, after logging into the review expert account, the review completion status and the status of samples that were not approved can be displayed, along with whether to resubmit for review. When a user completes a task, the corresponding task progress indicator box lights up, and the administrator account is promptly notified that the user's task has been completed. System message notifications for updates to personal account data can also be displayed.

[0050] The pathogen detection knowledge base system also includes search templates. The search module allows users to input relevant keywords to quickly retrieve related data from various database modules. Based on the input keywords, the search module retrieves relevant information displayed to the user's front end. For example, inputting the keyword "Klebsiella pneumoniae" retrieves and displays information about the pathogen, drug resistance, and virulence of "Klebsiella pneumoniae," showing related information such as... Figure 9 .

[0051] Specifically as follows: 1. Query parsing and expansion (1) Keyword standardization: Use biomedical entity recognition (such as BioBERT, MetaMap) to standardize user input: Input "Klebsiella pneumoniae" → map to the standard term "Klebsiella pneumoniae" (NCBI Taxonomy ID:573).

[0052] (2) Synonym expansion: "K. pneumoniae", "KP" (associated through medical dictionaries or knowledge graphs).

[0053] 2. Intent Classification (Optional) If the user input contains additional intent (such as "Klebsiella pneumoniae resistance"), the drug resistance database will be searched first.

[0054] 2. Multi-database retrieval

[0055] Database module division (must be pre-built)

[0056] Parallel retrieval optimization: Use asynchronous I / O (such as Python's asyncio) or distributed task frameworks (such as Celery) to query multiple databases concurrently.

[0057] 3. Result aggregation and sorting

[0058] Data association: Link data across databases using unique pathogen identifiers (such as NCBI ID).

[0059] Sorting strategy: Relevance: BM25 score (for text fields).

[0060] Authority: Prioritize displaying drug resistance genes from highly cited PubMed literature.

[0061] Timeliness: Virulence factor data are sorted in descending order of update date.

[0062] 4. Front-end display optimization, using structured templates.

[0063] Interactive features: Click on the drug resistance gene to jump to detailed experimental data.

[0064] A visual map shows the regulatory network of virulence genes (based on Neo4j relation data).

[0065] The search module also provides quick links to information; for example, clicking on drug resistance and virulence genes will take you to the drug resistance and virulence information module for more details.

[0066] This application discloses an integrated knowledge base system for high-throughput sequencing (NGS) detection of pathogens, consisting of a front-end display, a back-end maintenance subsystem, and an intelligent search engine. The front-end features four main functional modules: pathogen information, anti-infective drugs, drug resistance virulence, and background pathogens. The pathogen information module displays basic data and pathogenicity assessments of over a thousand microorganisms, including bacteria, viruses, fungi, and parasites, using a seven-level hazard classification. The drug module covers drug resistance mechanisms, target site alterations, and medication recommendations for 62 drug categories and 150 specific drugs. The drug resistance virulence module provides a closed-loop evidence chain of gene-pathogen-mechanism-clinical annotation. The background pathogen module constructs a localized contamination / colonization spectrum based on the hospital's consumables, environment, and population colonization characteristics, significantly reducing false positives. The back-end maintenance subsystem adopts a role-based hierarchical structure (assigner-maintenance personnel-review expert-user) and a task-driven model, achieving continuous data updates and quality control through task lists, priority scheduling, and progress visualization. The intelligent search engine integrates BioBERT entity recognition, synonym expansion, asynchronous parallel retrieval, and Neo4j toxicity network visualization, supporting the return of structured results for keywords within seconds.

[0067] Compared to existing technologies, the beneficial effects of this system are reflected in the following aspects: ① The localized background microbial bank reduces the NGS false positive rate by more than 30%; ② The four-dimensional integrated report on pathogen-drug resistance-virulence-drug reduces the clinical interpretation time from several hours to several minutes; ③ The closed-loop expert review ensures that the data is updated monthly, solving the pain point of the lag in traditional databases; ④ It covers 100% of common clinical pathogens and drug resistance phenotypes, meeting the needs of all scenarios of severe and difficult infections; ⑤ The visual interaction lowers the threshold for use, and even primary hospitals can quickly get started.

[0068] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A knowledge base system for pathogen detection, characterized in that, include: It showcases the front-end, back-end maintenance subsystem, and search module; among them, The display front end includes at least: The pathogen information module is used to present basic information about pathogenic microorganisms, pathogenicity assessment results, and related drug resistance and virulence data. The anti-infective drugs module is used to present drug categories, drug names, resistance mechanisms, resistance genes, and sources of evidence. The drug resistance and virulence module is used to present drug resistance genes, virulence genes, their associated pathogens, clinical interpretations, and sources of evidence. The background pathogen module is used to present the Chinese name, Latin name, colonization location, common sample types corresponding to the colonization location, and the background pathogen type of the background bacteria. The background maintenance subsystem includes at least: The account role management module is used to assign different permissions to data management allocators, maintainers, audit experts, and users. The knowledge base data maintenance task allocation module is used to create, distribute, prioritize, and track data maintenance tasks. The task receiving and execution management module is used by maintainers and review experts to receive, edit, submit, and review tasks; The system information update display module is used to publish and display knowledge base update announcements; The retrieval module is used for: Receive keywords input by the user; Keywords are standardized based on biomedical entity recognition and synonym expansion. Parallel retrieval modules for pathogen information, anti-infective drugs, drug resistance virulence, and background pathogens; The search results are aggregated, sorted, and returned to the front end for display based on relevance, authority, and timeliness.

2. The system according to claim 1, characterized in that, The pathogenicity assessment results in the pathogen information module are generated in the following ways: pathogenicity data are extracted from published literature using a natural language processing model; laboratory virulence factor assays and infection model phenotypic data are integrated; cross-validation is performed by combining clinical epidemiological surveillance and case report data; key data are manually reviewed by experts and data traceability information is recorded.

3. The system according to claim 1 or 2, characterized in that, The pathogen information module classifies the severity of pathogens into seven levels: Category 1: Infectious pathogens that require direct reporting; Category 2: Pathogens of key clinical concern; Category 3: Clinically pathogenic pathogens; Category 4: Opportunistic pathogens; Category 5: Colonizing / background bacteria; Category 6: Pathogenicity unknown; Category 7: Non-pathogenic pathogens.

4. The system according to claim 1, characterized in that, The drug resistance mechanism in the drug resistance virulence module includes at least one of the following mechanisms: alteration of antimicrobial drug target site, target replacement, target protection, antibiotic inactivation, efflux, decreased permeability, altered cell morphology, altered host-dependent nutrient acquisition, target overexpression, and drug resistance due to gene deletion.

5. The system according to claim 1, characterized in that, The background pathogen module is also used to perform the following tasks: distinguish background bacteria from laboratory environment and consumables; distinguish background bacteria from normal colonizing bacteria from different parts of the human body; and deduct or indicate corresponding background bacteria interference in the pathogen detection report.

6. The system according to claim 1, characterized in that, The retrieval module employs an asynchronous concurrent or distributed task framework to query pathogen information, anti-infective drugs, drug resistance virulence, and background pathogens in parallel. The retrieval results are displayed on the front end using a structured template, and the following operations are supported: (1) Click on the drug resistance gene or virulence gene to jump to detailed information; (2) Visualize the virulence gene regulatory network based on the Neo4j map.

7. The system according to claim 1, characterized in that, The task allocation module of the background maintenance subsystem supports the following operations: (1) Set priority for data entries to be added or updated; (2) Real-time statistics and display of task statuses: "Unassigned, Pending Start, Pending Review, Review Failed, Completed"; (3) Allow data management assigners to dynamically adjust task priorities and monitor the work progress of maintenance personnel and audit experts.

8. The system according to claim 1 or 7, characterized in that, The task receiving and execution management module includes: The maintainer account management submodule is used to edit, modify, and submit task data; The Expert Account submodule is used for quality checks, returning items for revisions, or confirming approval. The system message notification function submodule is used to notify users of task status updates and data update information in real time.

9. A method for generating a pathogen detection report using the knowledge base system described in any one of claims 1-8, characterized in that, Includes the following steps: Step S1: Receive NGS sequencing data of the sample to be tested; Step S2: Call the retrieval module to retrieve pathogen, drug resistance, and virulence information from the detected microorganism list; Step S3: Subtract background bacteria interference based on the background pathogen module; Step S4: Generate a test report by integrating the pathogenicity assessment results, drug resistance indications, and virulence factor information of the pathogen.

10. A non-transitory storage medium, characterized in that, It is used to store a program that causes a knowledge base system for pathogen detection as described in any one of claims 1 to 8 to perform the following action: execute the method for generating a pathogen detection report based on the knowledge base system as described in claim 9.