Biohazard big data analysis and monitoring early warning system
Through the biohazard big data analysis and monitoring and early warning system, the problems of lack of integrated analysis tools and poor environmental compatibility in the existing technology are solved, and comprehensive, efficient and accurate analysis of biohazard data and risk warning are achieved.
Patent Information
- Application Number
- CN202510614858.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-08
AI Technical Summary
The existing technology lacks integrated full-process analysis tools in the field of biohazard monitoring and prevention and control. The tools are scattered and the environmental compatibility is poor, the analysis efficiency and scalability are insufficient, and intelligent analysis and decision-making support is lacking.
It provides a biohazard big data analysis and monitoring and early warning system, including a graphical interface module, an analysis subsystem and Docker containerized environment module, which supports multi-level permission management and dynamic expansion, and processes sequencing data through quality control, error correction, assembly, and binning steps, and conducts in-depth analysis in combination with pathogen analysis, resistance gene and virulence assessment units.
It realizes comprehensive, efficient and accurate analysis of biohazard data, provides an intuitive and easy-to-use operating interface, ensures the accuracy and reliability of the data, and can promptly detect potential biohazard risks and take countermeasures.
Smart Images

Figure CN120452553A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of public health and safety technology, and in particular to a biological hazard big data analysis and monitoring and early warning system. Background Art
[0002] Systematic analysis of biohazards based on metagenomic and other biological big data is a potentially powerful approach to biohazard prevention and control. However, biological systems, such as the ecological environment, are nonlinear and complex, containing massive amounts of multidimensional data. There is no mature technology for analyzing biological big data to conduct risk assessments, early warning, and prevention of potential biohazards in the environment. Second- and third-generation sequencing data analysis software currently available on the market has significant deficiencies in the following areas, particularly in applications related to biohazard monitoring and prevention:
[0003] 1) Lack of integrated full-process analysis tools; 2) Scattered tools and poor environmental compatibility; 3) Insufficient analysis efficiency and scalability; 4) Lack of intelligent analysis and decision support. Summary of the Invention
[0004] The purpose of this application is to provide a biohazard big data analysis and monitoring early warning system that can comprehensively, efficiently and accurately analyze biohazard data.
[0005] To achieve the above objectives, this application provides the following solutions:
[0006] In a first aspect, the present application provides a biohazard big data analysis and monitoring early warning system, comprising:
[0007] The graphical interface module includes the main interface, parameter configuration interface, result display interface, and status monitoring interface. The main interface displays the analysis process framework diagram, the parameter configuration interface is used to configure parameters for the analysis steps, the result display interface is used to display visual charts and log information, and the status monitoring interface is used to dynamically display the progress bar and task status.
[0008] The analysis subsystem includes a core logic module, a unit analysis module, and a microbial community integrated analysis module; the core logic module includes a quality control unit, an error correction unit, an assembly unit, and a binning unit, and is used to process and optimize sequencing data; the sequencing data is the raw data obtained from biological samples through high-throughput sequencing technology; the unit analysis module includes a pathogen analysis unit, a resistance gene unit, a virulence assessment unit, and an evolutionary development unit, and is used to analyze the sequencing data to obtain information on pathogens, resistance genes, and virulence factors in biological samples; the microbial community integrated analysis module includes a traceability analysis unit, a mutation feature unit, a transmission evolution unit, and a transformation control unit, and is used to integrate the analysis results of each unit in the unit analysis module to analyze the structure, functional characteristics, and interactions of the microbial community;
[0009] The Docker containerized environment module is used to build independent container images for each unit in the unit analysis module and deploy a server as a control center in each independent container image; the server receives instructions from the graphical interface through the application program interface and calls the analysis tools in the container image;
[0010] Dynamic extension module, used to dynamically load new analysis tool modules based on the reflection mechanism and adjust tool parameters through configuration files or graphical interfaces.
[0011] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0012] This application provides a biological hazard big data analysis and monitoring and early warning system, which provides an intuitive and easy-to-use operating interface through a graphical interface module. The analysis process framework diagram displayed on the main interface helps users quickly understand the entire analysis process; the parameter configuration interface allows users to configure parameters for the analysis steps according to actual needs, enhancing the flexibility of the analysis; the result display interface clearly displays the analysis results through visual charts and log information, making it easier for users to understand and interpret; the status monitoring interface dynamically displays the progress bar and task status to ensure that users can grasp the analysis progress in real time. Secondly, the core logic module in the analysis subsystem processes and optimizes the raw sequencing data obtained by high-throughput sequencing technology through quality control units, error correction units, assembly units and binning units to ensure the accuracy and reliability of the data. The unit analysis module further analyzes the optimized sequencing data and deeply mines the information of pathogens, resistance genes and virulence factors in biological samples through pathogen analysis units, resistance gene units, virulence assessment units and evolutionary development units. The microbial community integrated analysis module integrates the analysis results of each unit in the unit analysis module and conducts in-depth analysis of the microbial community structure, functional characteristics and interactions. Finally, the early warning subsystem conducts risk assessment and early warning of biohazards through multiple modules, including risk assessment and early warning system design, dynamic research, and transmission analysis. By building assessment models, studying pathogen evolution and function, and predicting toxicity and transmission risks, the system can promptly identify potential biohazard risks and implement appropriate countermeasures. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0014] Figure 1A schematic diagram of the structure of a biological hazard big data analysis and monitoring and early warning system provided in one embodiment of the present application;
[0015] Figure 2 This is a diagram of the analysis process of the biological hazard big data analysis and monitoring and early warning system provided in one embodiment of the present application;
[0016] Figure 3 A logic diagram of an evaluation system provided in one embodiment of the present application. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0018] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0019] In an exemplary embodiment, Figure 1 As shown, a biohazard big data analysis and monitoring early warning system is provided, including:
[0020] The graphical interface module includes the main interface, parameter configuration interface, result display interface, and status monitoring interface. The main interface displays the analysis process framework diagram, the parameter configuration interface is used to configure parameters for the analysis steps, the result display interface is used to display visual charts and log information, and the status monitoring interface is used to dynamically display the progress bar and task status.
[0021] The analysis subsystem includes a core logic module, a unit analysis module, and a microbial community integrated analysis module; the core logic module includes a quality control unit, an error correction unit, an assembly unit, and a binning unit, and is used to process and optimize sequencing data; the sequencing data is the raw data obtained from biological samples through high-throughput sequencing technology; the unit analysis module includes a pathogen analysis unit, a resistance gene unit, a virulence assessment unit, and an evolutionary development unit, and is used to analyze the sequencing data to obtain information on pathogens, resistance genes, and virulence factors in biological samples; the microbial community integrated analysis module includes a traceability analysis unit, a mutation feature unit, a transmission evolution unit, and a transformation control unit, and is used to integrate the analysis results of each unit in the unit analysis module to analyze the structure, functional characteristics, and interactions of the microbial community;
[0022] The early warning subsystem includes a risk assessment and early warning system design unit, a dynamic research and transmission analysis unit, an assessment model construction unit, a pathogen evolution and function research unit, a toxicity and transmission risk comprehensive prediction unit, a pathogen risk monitoring network unit, and an unknown pathogen and potential risk identification unit.
[0023] In this example, before use, the user needs to select the corresponding sequencing platform and sequencing strategy to ensure that the tools and parameter settings in the subsequent analysis steps are correctly matched. The selected sequencing platform (for example, whether it is a long-read platform) will directly affect the tools used in subsequent analysis steps such as error correction, assembly, and alignment. The sequencing strategy (such as metagenomic sequencing, 16S rRNA sequencing, etc.) will determine the different analysis processes, ensuring that the appropriate analysis path is selected.
[0024] In this embodiment, the graphical user interface is developed using Swing technology, and the interactive interface is designed based on the Java Swing framework. The interface mainly includes the following parts: the main interface, which shows the framework diagram of the entire analysis process, including four steps: quality control, error correction, assembly, and binning; the parameter configuration interface, which allows the user to select the corresponding software tool and configure the parameters for each analysis step; the result display interface, which is used to present the visual chart and log information of the operation results; the status monitoring interface, which dynamically displays the operation progress bar and task status, such as running, completed, error, etc. The interface triggers the operation logic through event listeners (EventListeners). Users can drag or select analysis steps through the interface to dynamically adjust the analysis process, and can predefine common analysis process templates for one-click loading.
[0025] 1. By creating independent container images in Docker for each analysis software tool (for example, mainstream tools such as FastQC (quality control), SPAdes (assembly), and Maxbin (binning)), we ensure that the installation and operating environments of mainstream metagenomic analysis software are conflict-free and provide cross-platform operation support.
[0026] 2. Docker containerized environment module: Deploy a Spring Boot server in a Docker container as the control center for the analysis software. The main functions of the server include: receiving operation instructions from the graphical interface (such as starting analysis and setting parameters), calling the analysis software inside the container to run (such as starting quality control and assembly tasks), collecting analysis results, and sending the results to the graphical interface for display. The server deployment method is to create a RESTful API based on Spring Boot, expose an HTTP interface (such as a POST request to start an analysis task), and define a service unit to receive task requests from the graphical interface. The analysis tools in the container image include: FastQC and Trimmomatic (sequencing data quality control) used by the quality control unit; FMLRC and Pilon used by the error correction unit; SPAdes or MEGAHIT used by the assembly unit.
[0027] The system in this embodiment has a dynamic expansion mechanism (dynamic expansion module) based on the Java reflection mechanism, which can dynamically load new analysis steps or tool modules. Users can use this function to introduce new functional modules (such as gene annotation tools or metabolic pathway prediction modules) into the system. These modules are directly incorporated into the existing process through automatic integration and reflection calling, without the need for additional development or adjustment of the main program. At the same time, the system provides a parameterized calling interface, which flexibly adjusts the parameter settings of the tool through configuration files or graphical interfaces, meeting the scalability requirements of different experimental needs.
[0028] The dynamic extension modules include: a plug-in management submodule, which is used to manage and load third-party plug-ins to expand the analysis function of the system; the plug-in management submodule supports the installation, update and uninstallation of plug-ins through a graphical interface or command line; the plug-in formats include JAR, Python scripts or Docker images; the resource monitoring submodule is used to monitor the system resource usage in real time, including CPU, memory and storage space; the resource monitoring submodule displays resource usage through a graphical interface and issues an early warning when resources are insufficient, and supports automatic expansion of cloud resources or manual adjustment of resource configuration.
[0029] The graphical interface module also includes: a user management interface for managing the permissions and access control of different users; the user management interface supports multi-level permission settings, including administrators, researchers and ordinary users. Administrators can configure system parameters and tool modules, researchers can perform analysis tasks and view results, and ordinary users can view public analysis results and log information.
[0030] In an exemplary embodiment, the analysis subsystem includes a core logic module, a unit analysis module and a bacterial community integrated analysis module.
[0031] The core logic module includes quality control unit, error correction unit, assembly unit and boxing unit.
[0032] The quality control unit is responsible for performing a comprehensive quality assessment and preliminary cleaning of the input raw sequencing data, generating high-quality cleaned data, and ensuring the reliability and accuracy of the data for subsequent analysis. Specific implementation method: This module uses FastQC to perform an overall check on the data quality, generate a quality assessment report (such as sequence quality distribution, GC content, sequence length distribution, etc.), and use Trimmomatic and other tools to trim low-quality reads, remove adapter sequences, and ensure that high-quality reads are retained. In addition, Cutadapt and other tools are used to detect and remove residual adapter sequences to avoid false positive interference caused by sequencing primers or adapter sequences. Ultimately, the high-quality cleaned data output by this module provides input for the error correction and assembly modules, while minimizing noise data.
[0033] The error correction unit corrects errors in the high-quality cleaned data output by the quality control unit, further improving its accuracy and completeness and generating corrected data. Error correction of high-quality cleaned data is particularly important when long and short reads are mixed. Specific implementation: This module uses FMLRC to correct random errors in long reads using information from short reads. Furthermore, it uses Pilon or similar tools to further correct mismatches or indels after genome assembly to ensure the accuracy of the assembled data. Ultimately, the error correction unit provides highly reliable data input to the assembly module.
[0034] The assembly unit integrates high-quality cleaned data and corrected data into long sequence fragments (i.e., contigs data), laying the foundation for gene prediction and functional annotation. The specific implementation is: calling tools such as SPAdes or MEGAHIT, and splicing short reads into high-quality long fragments through the overlapping layout consistency (OLC) algorithm or the de Bruijn graph algorithm. For the mixed assembly of second-generation and third-generation sequencing data, this module can dynamically adjust the assembly parameters (such as k-mer size, coverage depth, etc.) to adapt to different data types and maximize the assembly effect. The contigs data finally generated by the assembly unit provides direct support for subsequent gene prediction, classification and functional analysis.
[0035] The binning unit further extracts high-quality genomic or gene set information by binning contigs data, generating high-quality data. Specific implementation methods include calling tools such as MetaBAT2 or CheckM to bin contigs data, classifying contigs into specific microbial genomes, extracting high-quality genomic data, and assessing genome integrity and contamination rates to ensure the scientific and reliable nature of the binning results. The binning unit provides direct support for analyses such as species classification and virulence assessment.
[0036] Among them, the unit analysis module includes: pathogen analysis unit, which uses Kraken, PhlAn and Centrifuge to classify the composition of microorganisms in environmental sample data, and predicts the gene region of contigs data through Prokka, GeneMark or Prodigal;
[0037] In the resistance gene unit, resistance genes in environmental sample data are compared with integrated databases such as BLAST, RGI, CARD, ARDB, and ResFinder, and unknown resistance genes are predicted by combining Prokka and InterProScan. In the virulence assessment unit, virulence factors in contigs data are compared with integrated databases such as VFDB, MvirDB, and PATRIC, and virulence protein structures are predicted by combining Prokka and SWISSMODEL. In the evolutionary development unit, a phylogenetic tree is constructed based on the 16SrRNA gene, and gene selection pressure is evaluated by dN / dS calculation.
[0038] Existing hazard databases include pathogen databases, resistance gene databases, virulence factor databases, and other related databases. Pathogen databases such as PATRIC, NCBI Pathoge, and RefSeq provide genomic data and information on pathogenic microorganisms; resistance gene databases such as CARD, ARDB, and ResFinder contain information on a variety of resistance genes, facilitating analysis of microbial resistance characteristics; virulence factor databases such as VFDB and MvirDB contain gene data related to microbial virulence, useful for identifying potential pathogens; and other related databases such as GenBank, KEGG, and UniProt provide extensive genomic and proteomic data. These databases provide important support for pathogen identification, resistance analysis, and virulence assessment.
[0039] Among them, Figure 3As shown in the figure, the pathogen analysis unit includes: a CNN- and Transformer-based sequence feature extraction subunit for pathogenicity prediction; an RNN- or LSTM-based time series analysis subunit for modeling microbial community dynamics; and a GNN-based protein interaction network analysis subunit for modeling virulence factor network topology. The unit analysis module includes a pathogen analysis unit, a resistance gene unit, a virulence assessment unit, and an evolutionary development unit.
[0040] Specifically, the pathogen analysis unit locates potential pathogenic microorganisms in environmental samples, uses tools such as Kraken, PhlAn and Centrifuge, and classifies and annotates known pathogens by comparing them with known genome databases. It generates species abundance tables and phylogenetic trees, quickly identifies the microbial composition and species affiliation in the samples, and uses tools such as Prokka, GeneMark or Prodigal to identify the coding gene regions in the assembled sequences, and annotate rRNA, tRNA and protein-coding genes. Relying on tools such as DeepVirFinder and PAIfinder, combined with deep learning models, it analyzes key sequence features and predicts potential virulence genes and their pathogenicity.
[0041] Among them, the specific implementation process of the pathogen analysis unit is as follows: 1. Call the quality control unit to generate high-quality cleaned data. 2. Call the correction unit to remove the host sequence background by setting the allowed number of mismatches and alignment length parameters to generate correction data. 3. Call the assembly unit to integrate the high-quality cleaned data and correction data into long sequence fragments (i.e., contigs data) 4. Perform gene prediction on the contigs data. Specifically, software such as Prokka / GeneMark / Prodigal is used to predict potential genes (including protein-coding genes, rRNA, tRNA, etc.). Gene regions are identified by analyzing the sequence features of contigs, such as start and stop codons, inconsistent GC content within genes, etc. 5. Perform species classification and identify the microbial composition in environmental samples. Specifically, tools such as Kraken / PhlAn / Centrifuge are used to compare with a large number of known genome databases to identify which known species the contigs data belongs to or to what extent it is close to these species, mainly through the k-mer strategy for fast and accurate classification. 6. Call the binning unit to classify the predicted genes or clusters in order to identify potential pathogenic factors. On the one hand, there is reference classification, that is, using databases such as PATRIC, VFDB, CARD, NCBI, EBI, etc. to compare with known pathogenic bacteria information. This comparison based on the reference database can identify genes similar to known pathogens. The other type is the combination of non-reference classification and macro properties, using feature-based prediction tools such as PathogenFinder, or through collaborative analysis of environmental data (such as abundance, proportion, ecological function, etc.) to identify potential pathogenic bacteria populations.
[0042] Furthermore, the pathogen analysis unit can also be combined with artificial intelligence tools to optimize the pathogen identification process in metagenomic environmental data samples to improve accuracy and efficiency. The specific implementation method is as follows:
[0043] 1. During the sequence identification process, input data originates from raw sequence data generated by high-throughput sequencers (such as Illumina and Nanopore). After being cleaned using quality control tools such as FastQC and Trimmomatic, high-quality sequence data is obtained. GC content is calculated by directly counting the ratio of G to C bases in the sequence and is a key species-specific feature. Sequence overlap is determined by analyzing the overlapping regions (k-mer overlap) of reads using assembly tools such as SPAdes, which can be used to identify potential gene regions. Functional domain location information is annotated using tools such as HMMER to scan sequences for conserved protein function (such as Pfam and COG). These multi-layered features are extracted using a CNN to extract local features (such as pathogenic mutations), while a Transformer captures global dependencies. DeepVirFinder automatically predicts the pathogenic potential of viral sequences, while PAIfinder predicts pathogenic pathways based on protein functional domains. The final output includes high-quality sequence data, GC content distribution, functional domain annotation, and pathogenicity analysis results, providing support for subsequent analysis.
[0044] 2. Functional prediction focuses on the potential biological functions of genes and their protein-protein interaction networks (PPIs). High-quality sequence data that has undergone quality control is input, and functional annotation tools such as Prokka and eggNOG are used in combination with external databases (such as KEGG and STRING) for functional domain alignment and metabolic pathway annotation. Protein interaction information is provided by the STRING database, integrating experimental verification and prediction data, and combining deep neural network (DNN-PPI) models to predict protein functional roles and interaction patterns. The PPI network is modeled through graph neural networks (GNNs), and the network topology characteristics (such as modularity and centrality) are analyzed to reveal microbial functional modules and their mechanisms of action. The final output includes gene function annotation reports, metabolic pathway information, the structure and characteristics of the protein interaction network, and the inference of the mechanism of action of microbial functional modules.
[0045] In evolutionary analysis, input data includes quality-controlled high-quality genomic sequences and 16S rRNA sequences. Phylogenetic trees are constructed using the PhlAn tool, and the ANI (genomic similarity index) is used to quantify interspecies similarity for classification and inferring evolutionary distances. Genomic similarity is calculated using sequence alignment (e.g., MUMmer or BLAST), while the density and distribution of key nodes in the phylogenetic tree are analyzed using PhlAn to identify the complexity of evolutionary branches. Unclassified sequences are then used to predict their taxonomic status and potential evolutionary direction using machine learning models (e.g., SVM and ANI-ML). Output includes phylogenetic trees, classification predictions for unclassified sequences, evolutionary patterns, and key node analysis.
[0046] 4. To fully reveal the dynamic changes in microbial communities, metagenomic data and environmental variable data (such as temperature, pH, and pollutant concentrations) are input, and combined with sampling time series information to analyze the dynamic changes in microbial community structure over time. Metagenomic abundance data are generated using the MetaPhlAn tool, time series trends are captured by RNN or LSTM models, and random forest models are used to analyze the correlation between environmental factors and community structure. The microbial interaction network is analyzed using the GNN model to associate changes in community structure with host health status. The output includes dynamic predictions of microbial abundance, results of environmental factor association analysis, and models of association between community changes and host health, revealing the ecological role of microorganisms in time and environmental changes.
[0047] Among them, Figure 3 As shown, the resistance gene unit includes: a resistance gene subunit carried by plasmids identified by PlasmidFinder, a resistance gene key feature subunit screened by RandomForest or SVM model, and a multidrug resistance analysis subunit integrating CARD, ResFinder and AMRFinder databases.
[0048] Specifically, the resistance gene unit analyzes microbial genome data in environmental samples to identify possible antibiotic resistance genes, thereby providing strong support for the prevention and control of environmental hazards. The specific methods are as follows:
[0049] 1. Establish a common process. Before identifying antibiotic resistance genes, environmental sample data will be cleaned and filtered to remove low-quality or contaminated sequences. The cleaned, high-quality data ensures the accuracy of subsequent analysis and reduces interference from host background and environmental noise. Output includes cleaned, high-quality sequence data and a sample quality assessment report.
[0050] 2. Comparison and identification of known resistance genes. The cleaned data input is compared with resistance gene databases (CARD, ResFinder, AMRFinder) through tools such as BLAST and RGI (Resistance Gene Identifier). The CARD database provides comprehensive information on known resistance gene types and resistance mechanisms, while the RGI tool combines CARD's rules for precise annotation. ResFinder and AMRFinder focus on the comparison and identification of resistance genes for specific antibiotic categories. With these tools, known resistance genes in samples can be quickly located, and their resistance characteristics can be initially understood. The output results include a list of resistance genes identified in the sample, classification information, and an annotation report on the resistance mechanism.
[0051] 3. Prediction and functional annotation of unknown resistance genes. For potential resistance genes not recorded in the database, a combination of gene prediction, functional annotation and homology alignment is used for identification. First, Prokka is used to perform gene prediction on the assembled contigs to identify the coding sequence (CDS). Secondly, InterProScan is used to annotate gene functional domains and identify conserved functional domains related to resistance mechanisms (such as β-lactamase and transporter). Finally, through homology alignment (BLASTp or DIAMOND), the similarity between the sample sequence and the existing resistance gene database is analyzed to screen candidate sequences that may belong to new resistance genes. The output results include a list of newly discovered potential resistance genes, annotated functional domains, and correlation analysis with resistance mechanisms.
[0052] 4. Clinical relevance analysis. The identified resistance genes are further evaluated to confirm their potential threat to public health. By querying literature and databases (such as NCBIPubMed or ECDC resistance reports), it is confirmed whether the resistance genes have appeared in clinical pathogens and whether they are associated with multidrug resistance (MDR). Multidrug resistance markers are compared to identify key genes that affect the efficacy of multiple antibiotics. The output results include a clinical importance assessment report for high-risk resistance genes and an analysis of markers related to multidrug resistance.
[0053] 5. Identification and prediction of resistance gene markers. Detect characteristic sequences associated with resistance genes. Simultaneously, based on sequence data and functional annotation results, machine learning algorithms (such as random forests or support vector machines) perform pattern recognition to predict new resistance markers. This process identifies characteristic sequences and potential markers associated with resistance genes, supporting early detection and environmental monitoring. Output includes a resistance marker prediction model and key sequence features associated with resistance genes.
[0054] 6. Identification of multidrug resistance signatures. This step primarily identifies resistance genes that may exhibit multidrug resistance (MDR), meaning resistance to multiple antibiotics or antimicrobial drugs. Multidrug resistance analysis is performed on the identified resistance genes using databases such as CARD, ResFinder, and AMRFinder. By comparing resistance marker databases, it is determined whether multiple resistance signatures coexist. This information is crucial for understanding the resistance spectrum of pathogens and developing preventive strategies. Output includes a list of multidrug resistance genes and a resistance spectrum analysis report.
[0055] 7. Interpretation of resistance mechanisms. Further analyze the resistance mechanisms of identified resistance genes. Interpretation relies on functional domain and three-dimensional structure analysis. Through homology comparison, determine the relationship between resistance genes and known genes, and infer the source of their resistance mechanisms. Use InterProScan to annotate functional domains (such as β-lactamase active domains or membrane transporter domains), and use SWISS-MODEL to predict the three-dimensional structure of proteins. Combined with molecular dynamics simulation, study the interaction between resistance proteins and antibiotic molecules. Output results include inferences on resistance mechanisms, three-dimensional structural models of resistance proteins, and molecular interaction analysis.
[0056] 8. Abundance Quantification and Epidemiological Assessment. Quantify the relative abundance of resistance genes in samples and assess their spread in the environment. Calculate the relative abundance of resistance genes using MetaPhlAn or Kraken tools and analyze their distribution characteristics in different environmental samples. Combined with epidemiological information (such as sampling location and time), assess the scale of resistance gene spread and public health risks. Outputs include resistance gene abundance distribution maps, epidemiological assessment reports, and hotspot identification.
[0057] 9. Analyze the transmission potential of resistance genes by detecting mobile elements in the genome. First, insert sequence (IS) detection is performed using the ISFinder tool to assess horizontal gene transfer capacity. PlasmidFinder is used to identify plasmids carrying resistance genes. Island analysis is performed in conjunction with the IslandViewer tool to identify genomic islands coexisting with resistance genes and further infer their transmission potential. Output results include an insert sequence analysis report, the distribution of plasmid-carried resistance genes, and an analysis of the transmission potential of genomic islands.
[0058] 10. Analyze the clinical and environmental significance of resistance genes in samples based on the WHO's resistance gene classification criteria. Combining abundance data and transmission characteristics, classify resistance genes into high, medium, and low risk categories, providing a scientific basis for public health prevention and control measures. Outputs include a resistance gene classification report based on international standards and risk assessment results.
[0059] Among them, Figure 3 As shown in the figure, the virulence assessment unit includes: a sequence feature screening subunit based on kmer frequency and Shannon entropy, a virulence factor local feature extraction subunit based on Onehot encoding and CNN, and a dynamic analysis subunit that constructs a virulence factor interaction network through GNN and identifies key nodes.
[0060] Specifically, the virulence assessment unit identifies virulence factors from metagenomic environmental data samples and conducts biohazard-related research. The specific methods are as follows:
[0061] 1. Quality control and assembly of sequencing data. The data are first evaluated for quality using the FastQC tool, including base quality distribution, adapter contamination, and GC content distribution. Trimmomatic is then used to clean low-quality reads and adapter sequences to ensure that the Q value of the data meets standards such as Q30, and to remove sequence fragments that are too short. The high-quality reads after quality control are input into assembly tools such as SPAdes or MEGAHIT for assembly to generate contig sequences to represent longer genomic fragments. After assembly, QUAST is used to evaluate the assembly quality, including N50 value, coverage, and redundant sequences, to ensure the accuracy and completeness of the data. The output results include cleaned high-quality reads, assembled contig files, and assembly quality assessment reports, laying the foundation for the identification of virulence factors.
[0062] 2. Comparison and classification of virulence factors. Comparison and classification of virulence factors is a key step in the rapid identification and annotation of known virulence genes. The input contig sequence is compared with virulence factor databases (such as VFDB and PATRIC) using BLAST or DIAMOND tools. These databases contain classification information of virulence factors such as secretion system genes, toxin genes, and escape genes. The comparison results are classified and annotated according to functional categories, such as secretion, invasion, and adhesion mechanisms, to clarify the specific role of virulence factors in the virulence mechanism. In addition, the MEGARes database or a self-built virulence factor database is used to analyze the combined effects of virulence genes and resistance genes, and to evaluate the effects of the synergistic effects of drug-resistant virulence factors on the virulence level of pathogens. The output includes virulence gene classification reports, functional annotations, and virulence-resistance synergistic analysis.
[0063] 3. Functional prediction is used to identify the functional mechanisms of virulence genes and the modes of action of their encoded proteins. Input virulence gene sequences are annotated with the functions of the encoded proteins using Prokka or GeneMark. Combined with the VFDB and VICTORS databases, the mechanisms of action of virulence factors are further elucidated. For example, how secreted proteins act on host cells via the type III secretion system, or how transporters help pathogens evade host immune responses, are explored. Outputs from this stage include a functional annotation report for the virulence genes, a detailed classification of protein functions, and a description of the mechanisms of action of virulence factors in the pathogenic process.
[0064] 4. Protein structure and functional domain analysis, predict the functional domains of virulence proteins through databases such as Pfam and InterPro. The input data includes the amino acid sequence of the virulence protein, combined with the functional domain analysis tool to identify characteristic sites such as signal peptides, active centers and catalytic regions. For example, the catalytic active center of the toxin or the N-terminal signal peptide of the secretory protein. SWISS-MODEL or AlphaFold is then used to predict the three-dimensional structure of the virulence protein and analyze its possible interaction sites with host proteins. Combined with KEGG and GO functional annotations, the role of the virulence protein in the metabolic network and its function in the secretion system are clarified. The output results include the domain annotation of the virulence protein, the three-dimensional structure prediction model and the protein interaction site analysis.
[0065] 5. Evolutionary analysis: Use multiple sequence alignment tools (such as MUSCLE or MAFFT) to analyze the conservation of virulence factors across strains or species and assess their evolutionary importance. Use PhyML or RAxML to construct phylogenetic trees to infer the evolutionary origins of virulence genes, and use HGTector to predict whether they have been transmitted through horizontal gene transfer. Combined with the Codeml tool, analyze selective pressures, assess the impact of non-synonymous mutations on key functional sites, and whether these genes are driven by positive selection. Output results include a phylogenetic tree, a virulence factor conservation analysis report, and a horizontal gene transfer assessment.
[0066] 6. Network structure and dynamic topology analysis of protein-protein interaction networks. Use tools such as Cytoscape or STRING to construct protein-protein interaction networks related to virulence factors, and determine key virulence proteins and their importance in the network by analyzing network topology features (such as node centrality and modularity). Input data include functional annotations and interaction information of virulence proteins. Analyze network topology features such as node centrality and modularity to determine key virulence proteins and their importance in the network. Combined with KEGG pathway analysis, identify signal or metabolic pathways involved in virulence factors. Dynamic network tools such as DyNet can be used to explore the expression changes of virulence factors under different environmental conditions and their impact on network structure. Outputs include virulence factor interaction network models, signal pathway analysis, and dynamic expression change reports.
[0067] Therefore, the system provided in this example can comprehensively mine virulence factor information from metagenomic samples, providing solid basic data and theoretical basis for biohazard research. These analyses not only help understand the pathogenic mechanisms of pathogens but also provide support for the development of public health and biosafety strategies.
[0068] Furthermore, the virulence assessment unit, combined with artificial intelligence methods, can improve the efficiency and accuracy of identifying virulence factors from sequences and their functional and network associations. The specific implementation process is as follows:
[0069] 1. This application uses AI algorithms to process the raw data of the metagenomics, and by mining sequence feature information, preliminarily screens out contigs that may be related to virulence. The input data comes from the raw sequencing data generated by a high-throughput sequencer (such as Illumina or Nanopore). After quality control and assembly processing, contigs are generated as analysis objects. The core of feature extraction includes GC content calculation, which evaluates species or functional specificity by counting the ratio of G and C bases in each contig; k-mer frequency analysis (such as 3-mer, 5-mer) characterizes sequence features by counting the frequency of occurrence of short sequence fragments; Shannon entropy is used to quantify the complexity of the sequence, indicating the level of information content of the gene fragment. In addition, the coding length calculation identifies the length characteristics unique to virulence genes by counting the length distribution of gene fragments. Functional domain information is annotated by Pfam or InterPro, and functional domain categories and structural information are extracted from the open reading frames (ORFs) of the contigs for further analysis. Machine learning models such as support vector machines (SVM), random forests, and CGBoost are used to analyze these features. SVM captures nonlinear relationships and is used for virulence gene classification in small-scale data. Random forests assess feature importance through multiple decision trees, identifying key virulence features such as fragments of specific lengths or functional domains. CGBoost optimizes complex feature combinations to improve the accuracy and model performance of virulence factor prediction. Output includes a list of virulence-related contigs and a ranking of key features.
[0070] 2. High-dimensional feature extraction: Deep learning methods are used to extract complex, deep features from contigs, further improving the accuracy of virulence factor analysis. Convolutional neural networks (CNNs) are used to extract local feature patterns from contigs, such as conserved motifs or virulence-associated regions. Input data is converted into a numerical matrix using one-hot encoding or embedding. Sequence fragments are fed into a multi-layer convolutional network using a sliding window technique to extract features such as specific k-mer patterns. Dimensionality is then reduced through a pooling layer, preserving important information and improving computational efficiency. Recurrent neural networks (RNNs) and long short-term memory (LSTM) networks are used to process sequential dependencies in DNA sequences and capture long-range dependencies. For example, within long contigs, RNNs analyze the association between promoters and virulence genes, while bidirectional LSTMs (BiLSTMs) extract cross-regional features to predict virulence correlations. Transformer models capture global features through an attention mechanism, making them suitable for analyzing complex patterns, such as motifs within regulatory sequences or protein domains. Output includes classification results and a deep feature map of virulence factors.
[0071] 3. Regulation and interaction analysis between contigs. AI is not only used for feature extraction of a single contig, but can also simulate the regulation and interaction between contigs. Through the combined model of CNN and RNN, the feature data of multiple contigs are input to predict whether they belong to the same regulatory network. For example, the recursive neural network combined with the convolutional layer extracts the regulatory pattern between sequences and captures the coordinated expression or functional association of contigs. Transformer-based interaction analysis constructs a "contig-sequence interaction matrix" and uses a multi-head attention mechanism to explore the regulatory relationship between contigs. The dynamic analysis module uses RNN to model the impact of changes in environmental conditions (such as temperature and pH) on the expression of virulence genes. By inputting virulence gene expression data at different time points, regulatory changes are predicted and the laws of dynamic changes in virulence gene expression are revealed. The output includes a contigs regulatory network diagram, a dynamic expression model, and an explanation of environmental adaptability.
[0072] 4. Network structure analysis and dynamic topology. This application uses AI technology to conduct an in-depth analysis of the relationships and dynamic characteristics of the virulence factor network. Through graph neural networks (GNN), contigs and their interactions are represented as graph structure data to construct a virulence factor relationship network. The input includes contigs functional annotations and interaction information, and the output is a topological characteristic analysis of the protein-protein interaction (PPI) network, such as identifying key nodes and modules. Dynamic GNNs (such as Dynamic Graph Convolutional Networks, DGCN) further analyze the dynamic changes of the network under different environments, such as the impact of antibiotic selection pressure on virulence factor expression and network connectivity. The random walk algorithm is used to explore the local characteristics of the network and analyze the enrichment of virulence factors in the neighboring network of a contig. At the same time, the machine learning model can predict potential connections in the network, supplement data incompleteness, and optimize the network structure. The final output includes the virulence factor network model, key node analysis results and dynamic change report, providing data support and strategic recommendations for virulence mechanism research and pathogen control.
[0073] Among them, Figure 3 As shown, the evolutionary development unit starts with the raw data of metagenomic environmental samples and uses bioinformatics tools and methods to gradually identify gene evolutionary relationships, functional gene characteristics, and their characteristics related to environmental factors and biological hazards. The specific process is as follows:
[0074] 1. Assembly of raw sequences into contigs. Metagenomic sequences extracted from environmental samples are usually highly complex short-read high-throughput sequencing data (such as reads generated by Illumina or Nanopore sequencing). In order to study genes and evolutionary relationships, these short sequences need to be spliced into longer contigs. Tools such as SPAdes or MEGAHIT are used to construct contigs based on sequence overlap, and the assembly quality, including length, N50 value, is evaluated using the QUAST tool. Redundant sequences need to be deduplicated to ensure the independence and accuracy of the contigs. The output includes high-quality contig sequences and their assembly quality assessment reports, providing high signal-to-noise ratio data input for subsequent evolutionary analysis and functional annotation.
[0075] 2. Based on the evolutionary relationship analysis of marker genes, highly conserved marker genes (such as 16S / 18S rRNA genes) are selected as the basis for analysis. The input contig sequences are used to identify marker gene regions using gene prediction tools (such as Prokka or GeneMark), and these conserved regions are aligned and low-complexity regions are trimmed using multiple sequence alignment tools (such as MUSCLE or MAFFT). Subsequently, a phylogenetic tree is constructed using the neighbor-joining method or maximum likelihood method to observe the branching relationship and evolutionary origin of virulence genes and resistance genes among different species. Combined with homology analysis, the amplification or deletion phenomenon of genes and the laws of their functional evolution are inferred. The output includes the evolutionary tree structure, the results of virulence gene homology analysis, and the prediction of evolutionary trends.
[0076] 3. Functional gene annotation and evolutionary pressure analysis: Gene prediction is performed on contig data. The input data is the assembled contig sequence. It is compared using functional databases (such as VFDB, CARD, or eggNOG) to annotate the functional categories of resistance genes and virulence genes. In addition, by calculating the ratio of non-synonymous mutations to synonymous mutations (dN / dS), it is assessed whether the gene is under positive selection pressure. Combined with environmental data (such as antibiotic concentration or host immune characteristics), the source of selection pressure is analyzed, the mutation characteristics of functional genes are located, and their association with pathogenicity and resistance is studied. The output results include gene function annotation reports, selection pressure analysis results, and functional annotations of mutation hotspot locations.
[0077] 4. Combine the selection pressure of environmental factors with changes in gene frequency to analyze the relationship between environmental factors and gene function, and study the impact of environmental factors (such as antibiotic concentration or heavy metal pollution) on gene function and frequency changes. By inputting environmental data (such as antibiotic concentration, heavy metal pollution level) and gene abundance data, study how environmental factors affect gene function and frequency changes. For example, use the GenePop tool to analyze the distribution pattern of genes in different geographical samples, detect the recombination rate and mutation rate of genes, and infer the diffusion pattern of drug-resistant genes or metal-resistant genes driven by selection pressure. Through dynamic change data, predict the adaptive evolution path and propagation mechanism of genes in different environments. Output results include gene frequency change trend graph, selection pressure-driven analysis report and adaptive evolution prediction model.
[0078] 5. Research on adaptive evolution: assessing the selective effects of antibiotic exposure on resistance genes and investigating how antibiotic exposure promotes the selective amplification of resistance genes. Input data includes resistance gene mutation site data and gene sequences related to metabolic function. PAML or Bayes Empirical Bayes (BEB) tools are used to detect positive selection signals in gene mutations and identify genes undergoing accelerated evolution under adaptive pressure. Further analysis is conducted to identify adaptive mutation clusters in specific genomic regions and assess their function and potential in metabolic regulation and host adaptation. Output includes positive selection signal detection results, annotations of regions with adaptive mutation clusters, and analysis of gene functional adaptability.
[0079] 6. Population effect analysis, including variant gene diversity, analysis of the distribution of single nucleotide variants and small insertions / deletions, and assessment of the prevalence of variant genes in the population and their evolutionary potential. Input data includes the distribution information of single nucleotide variants (SNPs) and small insertions / deletions (InDels). Combined with tools such as STRUCTURE or ADMIXTURE, the distribution patterns of genes in different populations are analyzed to infer the evolutionary relationships of populations and the impact of gene flow on the accumulation of variants. By integrating population gene flow data with ecological and environmental information, the speed of gene transmission between populations and its impact on population adaptability are predicted. Output results include population structure analysis reports, variant gene diversity assessments, and gene flow prediction models.
[0080] 7. Detection of genetic variation under selection pressure, detection of horizontal gene transfer events (such as resistance genes acquired from other species through HGT), and inference of the time nodes of gene mutations and their evolutionary rates in different environments. The input data is the resistance gene and its adjacent genomic sequence. The HGTector tool is used to analyze the horizontal gene transfer pattern, and the molecular clock model is combined with ecological data to infer the historical trajectory of gene diffusion. The output includes analysis of horizontal gene transfer events, calculation of gene evolution rate, and inference of the historical trajectory of transmission, providing a reference for understanding the evolutionary dynamics and ecological adaptability of resistance genes. The entire process provides comprehensive support for studying the evolutionary adaptability of microorganisms and their interactions with the environment and hosts through systematic processing from raw data to complex analysis.
[0081] Metagenomic data analysis can reveal the adaptive evolutionary mechanisms of microorganisms by assembling raw sequences into contigs, constructing evolutionary relationships, and performing functional annotation and evolutionary pressure analysis, combined with environmental factors and population evolutionary effects. Furthermore, applying this information to biohazard research can predict the evolutionary direction of pathogens, identify virulence genes and resistance genes that may cause biohazards, and their evolutionary relationships. It can also investigate how gene mutations affect virulence, transmissibility, or drug resistance, predict the potential emergence of high-risk resistant strains or virulence factors, and provide data support for antibiotic development, novel therapeutic strategies, and environmental remediation.
[0082] Furthermore, the evolutionary development unit can be combined with metagenomic environmental data analysis using artificial intelligence tools to further promote the automation and accuracy of ARG identification, horizontal gene transfer prediction, evolutionary path inference, and key functional gene research, such as Figure 2 As shown, the specific details can be as follows:
[0083] 1. The identification of resistance genes depends on the extraction of potential patterns from gene sequences or protein sequences. This application integrates resistance gene databases (such as CARD and ResFinder) to construct a training data set and convert DNA or protein sequences into an input format acceptable to the model, such as one-hot encoding or embedded representation. The input data comes from high-quality reads from metagenomic sequencing, and sequence information is generated after assembly and annotation. The convolutional neural network (CNN) extracts local patterns (such as conserved regions) of resistance gene sequences through its convolutional layer, and uses the pooling layer to compress features, reducing the data dimension while retaining key information. The model outputs a classification result to determine whether the sequence belongs to a known resistance gene category. In addition, by analyzing the genomic GC content and upstream and downstream characteristics, the model can identify possible horizontal gene transfer (HGT) events, distinguish between exogenous and endogenous genes, and provide support for subsequent ecological transmission research on resistance genes.
[0084] 2. Deep learning models are used for evolutionary tree construction and path prediction. Traditional methods such as RAxML and IQ-TREE rely on multiple sequence alignments, which are computationally complex and sensitive to data quality. This application uses sequence embedding tools (such as SeqVec and ESM) to convert DNA or protein sequences into high-dimensional feature representations, and uses deep neural networks or graph neural networks to directly infer the evolutionary relationship between sequences, skipping the tedious steps of multiple sequence alignments. At the same time, combined with LSTM or RNN models, sequence change data of samples at multiple time points are analyzed to predict evolutionary paths and branch directions, providing an efficient means for understanding gene evolution and functional expansion. The output includes a high-precision prediction model of the evolutionary tree and dynamic analysis results of evolutionary branches.
[0085] 3. LSTM is used in adaptive evolution research. Long-short-term memory networks are suitable for capturing the temporal dependencies of sequence data and for studying the dynamic changes of genomes at different time points or evolutionary stages. In adaptive evolution research, the input data includes gene frequency changes under antibiotic exposure conditions and dynamic sequences of environmental samples. The LSTM model captures the temporal trends of gene frequencies and predicts the future spread of resistance genes. In addition, analysis of dynamic changes in gene expression profiles helps identify key genes associated with nutritional adaptation or host immune escape. Outputs include models predicting the spread of resistance genes and an assessment of the contribution of key genes to environmental adaptation, providing scientific support for vaccine design and environmental governance.
[0086] 4. RNN is used to analyze the potential of pathogenicity genes and can model the mutations and pathogenicity enhancement pathways that genes gradually acquire during evolution. The input data is the virulence-related gene sequences in the host genome and their associated environmental factors (such as immune pressure). By analyzing the time series of this data, RNN predicts the functional enhancement potential of virulence genes in different environments and generates a time series score to assess the risk of gradual enhancement of pathogenicity. The output includes a temporal dynamic model of virulence gene functional enhancement and a potential risk assessment.
[0087] 5. Clustering methods and population adaptability research. Unsupervised clustering algorithms (such as K-Means and DBSCAN) can partition samples into subpopulations and identify genotypic differences. Inputs include functional annotations of SNP and INDEL variant sites and sample sequence characteristics. Clustering algorithms group gene sequences by function or mutation pattern, revealing adaptive characteristics under environmental pressure. For example, the K-Means algorithm divides samples into subpopulations based on mutation patterns and analyzes how these subpopulations perform under different environmental conditions. Output includes an analysis report on the adaptive characteristics of the variant populations and a map of the mutation patterns driven by environmental pressure.
[0088] 6. Random forests are used to screen for key pathogenicity features. The random forest algorithm can effectively identify key features (such as specific mutations in virulence and resistance genes) in high-dimensional data. Input data includes sequence characteristics of virulence and resistance genes, as well as environmental variables (such as antibiotic concentration and host species). The model uses importance scores to assess the correlation between mutation sites or conserved regions and pathogenicity or drug resistance, screening for key features. The output includes a prioritized list of key genes or mutational features, providing a reference for experimental validation and drug development.
[0089] 7. RNNs are used for HGT-related analysis. RNN models can capture the patterns of HGT events along the evolutionary timeline. Input data include genomic GC content, characteristics of upstream and downstream regions, and environmental factors such as frequency of interspecies contact or duration of coexistence. Combined with time series analysis, RNN models capture the dynamic patterns of HGT events along the evolutionary timeline and generate HGT risk assessment maps. Outputs include models predicting species or environments likely to experience HGT in the future, as well as a report analyzing HGT temporal dynamics.
[0090] 8. Self-supervised learning and representation learning can build models using unlabeled data. Self-supervised learning can learn the underlying patterns of sequences through pseudo-labeling tasks. This application uses self-supervised learning to pseudo-label DNA or protein sequences, extracting functional annotations and evolutionary pattern features. Representation learning also embeds high-dimensional gene sequences into a low-dimensional space, preserving key information for population classification or propagation network analysis. The output includes low-dimensional embedding features of genomic data and predictions of key nodes in the propagation network, providing unprecedented depth and breadth of support for evolutionary analysis and ecological adaptability research.
[0091] In an exemplary embodiment, the integrated microbial community analysis module includes: a traceability analysis unit, which is used to detect SNP / Indel through GATK or bcftools, and to construct a pathogen geographical distribution map in combination with GIS tools; a mutation feature unit, which is used to detect point mutations and structural variations through GATK and Manta, and to perform functional annotation in combination with ANNOVAR; a transmission evolution unit, which is used to construct a phylogenetic tree through IQtree and to analyze the temporal dynamic transmission path in combination with BEAST; a transformation control unit, which is used to perform binning classification through MetaBAT2 and predict the origin of the plasmid using PlasFlow.
[0092] Specifically, such as Figure 3 As shown in Figure 1, the traceability analysis unit, from sequencing data to evolutionary traceability and then to biohazard research, deeply explores the data potential at each step by integrating different methods and tools, combining artificial intelligence, geographic information system (GIS), big data analysis and multidisciplinary cross-disciplinary research, such as Figure 2 As shown, the following is the specific implementation of this application:
[0093] 1. After the sample sequencing is completed, the first raw data input is usually a FASTQ format file generated by a high-throughput sequencer (such as Illumina or Nanopore). These files contain sequences and their quality scores. Use tools such as FastQC and MultiQC to evaluate data quality and generate reports on indicators such as base quality distribution, GC content, and adapter contamination ratio. Next, use Trimmomatic or Cutadapt to remove low-quality bases (below Q20 or Q30), sequencing adapters, and short sequences less than a specified threshold (such as 50bp) in length, and output them as filtered high-quality sequence data. To address the problem of host DNA contamination, align the sample data to the reference host genome (such as the human or animal genome), use Bowtie2 or BBMap to remove non-target sequences, and generate a host-removed data file. Subsequently, the high-quality non-host sequence is input into SPAdes or MEGAHIT for de novo assembly to generate Contigs or Scaffolds. Finally, gene prediction and functional annotation of these contigs are performed through Prokka or Prodigal, potential functional genes and related annotations are marked, and structured standardized data are output to provide high-quality input for tracing analysis and subsequent research.
[0094] 2. Population genetic variation analysis: Use GATK or bcftools to align to the reference genome to detect variants such as SNPs and indels, and combine with VCFtools or Annovar for annotation. First, use tools such as BWA or Bowtie2 to align the cleaned sequences to the reference genome to generate SAM / BAM files. Use GATK or bcftools to detect single nucleotide variants (SNPs) and insertions / deletions (Indels) and generate variant call files (VCF format). Subsequently, use VCFtools or Annovar to annotate the variant data, marking the location of the variant in the genome and its possible functional impact (such as non-synonymous mutations and splice site variants). After inputting these variant site data, analyze the variation frequency distribution across samples to assess population genetic diversity. Haplotype network analysis (such as PopART) is also used to visualize the evolutionary relationships of the samples. Combined with random forest or neural network models, feature selection and prediction are performed on the variant data to analyze the association between genotype and phenotype. Output includes a variation frequency distribution map, a list of key variation sites, and a haplotype network that illustrates the evolutionary relationships of the samples.
[0095] 3. Combine GIS and other geographic information to trace the source of pathogens, collect the geographical source information of samples, and the input data include the geographical source information of metagenomic samples (such as the latitude and longitude of the sampling site), environmental condition data (such as climate, land use), and pathogen genome data detected in the samples. Use GIS tools (such as ArcGIS or QGIS) to construct a geographical distribution map of pathogens to show the regional transmission pattern of pathogens. Subsequently, these geographical distribution data are superimposed with environmental factors (such as climate, population density, etc.) and correlation analysis is performed to explore the impact of environmental conditions on the distribution of pathogens. Use diffusion models (such as MaxEnt or spreadR) to simulate the spread potential of pathogens under different geographical conditions, and combine multivariate statistical models (such as GAM) to identify transmission drivers.
[0096] 4. Analysis of time series characteristics and transmission dynamics. Input data includes time series data of pathogens (such as changes in the frequency of mutation sites over time) and sample collection time. Use TimeSeries in R language or statsmodels library in Python for time series modeling to analyze the dynamic trend of mutations. Combine epidemiological models (such as SEIR model) to simulate the transmission process of pathogens, and use anomaly detection algorithms (such as LOF or DBSCAN) to identify abnormal transmission events (such as sudden epidemics). In addition, by constructing a network transmission model, study the diffusion pattern of pathogens in complex networks (such as host-host networks or regional transportation networks). Combined with Granger causal analysis, explore the correlation between environmental factors (such as temperature, antibiotic use) and transmission dynamics. Output includes time series graphs of mutation frequencies, reports of abnormal transmission events, and causal analysis results of environmental factors and transmission dynamics.
[0097] 5. Combine historical information for traceability analysis. The input data includes sample time information (such as sampling year) and variant site data, and use this information to generate a time-calibrated dataset. Use the BEAST tool to construct a time-calibrated evolutionary tree of the pathogen to estimate the pathogen's evolutionary rate, origin time, and transmission path. Combine historical outbreak records or host migration paths to verify the credibility of BEAST's inference results. Use posterior probability analysis to assess the likelihood of different transmission paths. Combine modern samples with ancient DNA samples to analyze the long-term evolutionary history of the pathogen and explore the relationship between host adaptability and pathogen evolution. Output includes a time-calibrated evolutionary tree, a credibility assessment of the transmission path, and a comprehensive report on the long-term evolutionary history of the pathogen.
[0098] 6. Rapidly identify potential pathogens in metagenomic samples using the GMI platform. The input data consists of high-quality sequence data from metagenomic samples. The GMI (Global Microbial Identifier) platform integrates core genome and pan-genome information, and the Pathogenome tool constructs pathogen gene functional networks. Deep learning techniques are used to analyze the association between genomic and epidemiological data, predicting pathogen transmission pathways and key transmission chains. Combined with gene functional annotations, the transmission mechanisms of pathogens and potential public health threats are explored. Outputs include pathogen core genome and pan-genome analysis reports, gene functional networks, and transmission chain prediction models, providing data support and decision-making for pathogen prevention and control.
[0099] Among them, the mutation feature unit conducts a comprehensive study on the mutation information in metagenomic data, covering the entire process from raw sequence processing to functional annotation, pathogenicity research and spatial distribution analysis. Figure 2 The specific method is as follows:
[0100] 1. Metagenomic mutation studies first require high-quality data input. Input data comes from raw sequence data generated by high-throughput sequencers (such as Illumina or Nanopore) and is typically stored in FASTQ format. These data are first quality-checked using FastQC or MultiQC to assess base quality distribution, GC content, adapter contamination, and reproducibility. The results are used to determine whether the data require further processing. Trimmomatic or Cutadapt are then used to remove low-quality reads (usually using a Q20 or Q30 threshold) and sequencing adapters, while also removing short sequences of insufficient length (e.g., less than 50 bp). Next, CD-HIT is used to de-redundant the sequences, removing repetitive regions to reduce analytical redundancy and computational complexity. High-quality, de-redundant sequence data are input into SPAdes or MEGAHIT for assembly, generating contigs or scaffolds. These assembly results are then quality-assessed using tools such as QUAST to ensure the data are suitable for subsequent mutation analysis. Output includes quality-controlled, high-quality contig files and an assembly quality assessment report.
[0101] 2. For different types of mutations, this application uses multiple tools for collaborative analysis. Input assembled Contig or Scaffold data for in-depth analysis. For point mutations (such as SNPs) and small fragment insertions / deletions (Indels), the GATK tool is used for variation detection, providing a standardized processing flow from preliminary mutation calls to genotype calibration; FreeBayes uses a Bayesian model to process variation data of complex samples, which is particularly suitable for mixed samples; and DeepVariant uses deep learning methods to detect mutations with high precision from low-quality sequences. For structural variations (such as large fragment insertions, deletions, and duplications), Manta or Delly is used for analysis, and CNVnator or Control-FREEC is combined to detect copy number variations in the genome. Through collaborative analysis of these tools, point mutations and structural variations can be fully covered. The output includes mutation data such as SNPs, Indels, and CNVs, usually stored in VCF format for subsequent functional annotation and ecological analysis.
[0102] 3. Functional annotation and pathogenicity study. Functional annotation of detected mutations is a key step in understanding their biological significance. Input the detected mutation site data (VCF format) into the ANNOVAR and SnpEff tools to annotate their gene location, coding region variation effects (such as non-synonymous mutations) and non-coding region function predictions. Combined with databases such as EggNOG, KEGG and Pfam, the mutation annotation results are mapped to functional pathways to analyze the impact of mutations on biological processes. This application further integrates multiple mutation features (such as conservation, frequency, functional impact) through PVSI and PMI tools to calculate a comprehensive pathogenicity score. Combined with the CADD score, the pathogenicity prediction results of the mutation are generated; PolyPhen-2 and PROVEAN are used to evaluate the potential impact of non-synonymous mutations on protein function. Compare the mutation data of environmental samples and clinical samples to assess the public health risk of the mutation, and combine with drug resistance gene studies (such as CARD or ResFinder databases) to analyze the role of mutations in the spread of drug resistance. The output includes an annotated mutation function table, pathogenicity score and drug resistance mutation transmission assessment.
[0103] 4. Distribution characteristics and variation pattern analysis, the frequency distribution of mutations and their patterns in samples can reveal the ecological significance of mutations. Input mutation data and environmental information of samples (such as water bodies, soil, and atmosphere) to analyze the differences in the environmental distribution of mutations and their selection pressures. Visualize mutation distribution data through dimensionality reduction techniques such as PCA to explore clustering relationships between samples. If mutation data containing long-term sampling time points is input, the temporal dynamic characteristics of mutation patterns can be analyzed to identify trends in evolution or adaptation. Combined with GIS tools, mutation distribution data can be superimposed with geographic information (such as the latitude and longitude of the sampling site or the regional pollution level) to study the driving effect of environmental factors (such as pollutant concentrations) on mutation distribution. Outputs include mutation frequency distribution maps, temporal dynamic analysis results, and geographic distribution characteristic models.
[0104] 5. Use IGV to view the location and functional impact of mutations in the genome, and use Circos diagrams to display mutation distribution patterns. Input mutation data and its genomic location information, and display and analyze them through visualization tools. Use IGV (Integrative Genomics Viewer) to view the distribution and functional impact of mutations in the genome; use Circos diagrams to display the genome-wide distribution pattern of mutations. Use visualization libraries in R or Python (such as ggplot2, matplotlib) to draw mutation frequency distribution maps and dynamic charts of their changes over time and space. Combine Shiny or Dash frameworks to develop interactive tools to provide researchers with a dynamic perspective on mutation analysis. In addition, 3D visualization technology is used to display the correlation between mutations and geographic space to enhance the intuitiveness of mutation ecology research. Outputs include mutation distribution maps, dynamic interactive visualization tools, and genome function annotation charts.
[0105] 6. Combined with the functional annotation results, this application combines the functional annotation results and focuses on analyzing the impact of mutations on protein structure and function. The input data includes the protein sequence and annotation information of the gene where the mutation is located. Combined with pathogenicity scores (such as CADD scores), mutation frequencies, and conservation analysis, the potential role of mutations in protein active sites or functional domains is evaluated. By training the Random Forest or XGBoost model, the mutation characteristics and annotation data are used as input to predict the pathogenicity of the mutation. Combined with clinical case data, the actual hazards of mutations are verified, and the relationship between environmental mutations and characteristics such as drug resistance and virulence factors is explored to provide data support for assessing public health risks. The output includes protein structure modeling results, mutation pathogenicity prediction models and risk assessment reports.
[0106] 7. Compare the mutations in environmental samples with drug-resistant gene libraries (such as CARD and ResFinder) to study the transmission pathways and hotspots of drug-resistant genes. Input data include mutation sites, metagenomic annotation results, and geographic sample information of environmental samples. Combine the high-frequency mutations in the samples with the homology of related genes in the human microbiome to study the potential risk of gene transmission. Assess the mutation selection pressure of environmental pollution (such as antibiotic residues or heavy metal pollution) on the microbial population through mutation pattern analysis, and predict the risk of biological hazards in the environment. Combine the mutation characteristics of different geographical regions to study their relationship with local pathogen outbreaks, and use macroecological methods to analyze the role and driving factors of mutations in the transmission dynamics of microbial populations. Outputs include the transmission pathways of drug-resistant genes, identification of hotspots, and analysis of the ecological driving forces of mutations.
[0107] Through the above process, this application can comprehensively mine mutation characteristics in metagenomic data and analyze their potential impact on ecosystems and public health. Functional annotation, pathogenicity analysis, distribution pattern research, and spatial association analysis are used to reveal the biological significance of mutations.
[0108] Among them, Figure 3 As shown in the figure, the transmission evolution unit integrates multi-dimensional research methods from raw sequence processing to transmission chain analysis, host adaptation research and ecological niche prediction, covering all aspects of basic to applied research. The specific steps are as follows:
[0109] 1. The basis of metagenomic transmission evolution analysis is high-quality data. The input data is usually a FASTQ file generated by high-throughput sequencing, and FastQC is used to evaluate the base quality distribution, GC content, and adapter contamination ratio of the data. Trimmomatic is then used to remove low-quality reads, adapter sequences, and sequences that are too short to ensure that the data quality meets the analysis standards. SPAdes or MEGAHIT is used for assembly to generate Contig or Scaffold sequences, and Prokka or InterProScan is used to functionally annotate the genome to identify functional regions and open reading frames (ORFs). Next, pan-genome analysis is performed using Roary or PanX to divide the genome into core genomes and variable genomes, and metabolic, virulence, and drug resistance gene modules are annotated using databases such as KEGG and EggNOG to provide comprehensive functional information and gene distribution maps for transmission evolution analysis. The output results include annotated genome files, core-variable gene distribution reports, and functional classification results.
[0110] 2. Phylogeny and transmission chain analysis: the input data is the core gene sequence or the whole genome alignment result. The phylogenetic tree is constructed by IQtree or RAxML, and the phylogenetic relationship is calculated based on model selection (such as maximum likelihood or Bayesian inference). MrBayes is used for Bayesian statistical inference to generate an evolutionary tree with confidence intervals. At the same time, SNP analysis and inter-genomic distances are combined to infer the transmission chain and evolutionary relationship between samples. The divergence time and transmission dynamics of pathogens are analyzed by combining spatiotemporal transmission models (such as EpiPhylo) with time correction tools (such as TimeTree). Integrate geographic data to generate pathogen transmission maps, reveal the spatial and temporal laws of transmission chains, and analyze cross-species transmission events in combination with host information. The output includes phylogenetic trees, transmission chain models, and time-corrected transmission paths.
[0111] 3. Genetic diversity and variation analysis, the input data includes the genome sequence of the sample and the reference genome sequence. Use Snippy or GATK to detect SNPs and Indels and generate a variation call file (VCF format). By analyzing the variation frequency distribution between samples, the variation patterns and hotspot areas are revealed. Combined with tools such as HGTector to detect horizontal gene transfer (HGT) events, and ICEborg to analyze the role of integrons and mobile elements in genome evolution. Further combined with protein structure prediction tools (such as AlphaFold), the functional consequences of mutations are verified and the significance of specific mutations in adaptive evolution is evaluated. PhyloSNP is used to analyze the variation spectrum between strains to infer the direction of gene flow and its impact on fitness. Outputs include variation detection reports, HGT event analysis, and mutation function annotation results.
[0112] 4. Transmission dynamics and spatiotemporal dynamics analysis: Input data includes time information (such as sampling date), genome sequence, and alignment data. BEAST is used to construct a time scale to infer the pathogen's evolutionary rate and temporal dynamics. Nextstrain is used to dynamically visualize the pathogen's transmission paths and evolutionary patterns, generating a dynamic map of its geographic distribution. The impact of environmental factors (such as climate change and antibiotic use) on transmission patterns is studied, revealing the correlation between pathogen transmission patterns and environmental changes. Output includes a time-scale evolutionary tree, a dynamic map of geographic transmission, and analysis of the impact of environmental factors.
[0113] 5. Host adaptability and ecological research: Input data includes pathogen genome sequences and host sample information. Gene-host relationship networks are constructed using Cytoscape to analyze the impact of genetic variation on host adaptability. Ecological niche modeling tools are used to predict the distribution trends of pathogens in different environments, and the Shannon index or Alpha / Beta diversity indicators are used to assess the ecological diversity of host-pathogen systems. Macroecological data (such as precipitation and land use) are combined to analyze the impact of environmental factors on transmission and niche differentiation. Output includes gene-host network models, niche prediction maps, and diversity analysis reports.
[0114] 6. Transmission pathway analysis and gene flow research: Input data includes genomic variation information of pathogen populations, geographic data, and human activity data. Use fastSimCoal2 or TreeMix to infer the direction of gene flow between populations, construct a transmission network, and refine the transmission pathways based on environmental factors. Identify high-risk transmission nodes by integrating human activity networks. Combine multi-omics data (such as transcriptomes and metabolomes) to analyze the dynamic relationship between gene expression and transmission, revealing key gene flow pathways and their driving effects on transmission. Output includes a transmission network model, gene flow direction, and risk node analysis.
[0115] 7. Pathogenicity and evolution analysis: Input data includes mutation information of virulence genes and drug-resistance genes, as well as host genome samples. Positive selection analysis is performed using tools such as PAML to study the evolutionary patterns of virulence factors or drug-resistance genes and assess their contributions to host adaptability and pathogenicity. The role of virulence genes in evolution is analyzed in conjunction with HGT events, and the impact of cross-host transmission on pathogenic evolution is studied. Co-evolution analysis methods are used to study the mutual adaptation mechanisms between host and pathogen genomes, revealing the long-term impact of pathogen evolution on host ecosystems. Outputs include analysis of virulence gene evolutionary patterns, reports of mutations driven by positive selection, and co-evolution results.
[0116] 8. Ecological diversity and prediction models, combined with environmental metagenomic data, study niche differentiation and co-evolution among species. Input data include environmental metagenomic sequences, mutation data and environmental factors (such as climate and pollutant concentrations). Use CANOCO to analyze the principal components of environmental factors and genetic diversity, and explore the ecological significance of mutation hotspots. Use prediction models (such as Bayesian Skyline or Eco-evolutionary Dynamics Models) to simulate the future spread range of pathogens and their evolutionary trends in combination with climate change scenarios. Evaluate the potential impact of selection pressure on transmission and evolution, and reveal the role of ecological factors in driving pathogen evolution. Outputs include principal component analysis of ecological diversity, predictions of future transmission and evolutionary trends, and selection pressure assessment reports.
[0117] Among them, Figure 3 As shown in the figure, the transformation control unit starts with the processing of environmental samples, and then goes through binning analysis, gene function screening, transformation feature identification, plasmid transfer assessment, phylogenetic tree construction, functional domain and tolerance exploration, etc., to comprehensively analyze the metagenomic data in environmental samples and identify potential biosafety risks. The specific process is as follows:
[0118] 1. After obtaining the contig sequences of environmental samples, the input data is the contig sequences of environmental samples, which are usually derived from high-throughput sequencing assembly results (such as contig files generated by SPAdes or MEGAHIT). First, binning analysis is performed on the contigs using tools such as MetaBAT2, MaxBin, or CONCOCT. The mixed contig data are clustered into different bins based on features such as k-mer frequency, GC content, and read coverage, preliminarily separating the genomic data at the species level. Subsequently, the integrity and contamination level of each bin are evaluated using CheckM or BUSCO to generate an integrity score and contamination assessment report, and the binning results are adjusted based on these indicators. The results of the DAS Tool are further integrated to optimize the classification by combining the advantages of multiple binning algorithms. In the case of multiple sampling, binning analysis is performed on samples from different time points. By comparing the distribution and variation of the same species or genome at different time points, the dynamic changes in the microbial community are analyzed. The output includes optimized binning results, an assessment report, and a microbial community dynamic change map.
[0119] 2. Combine CARD and ResFinder to screen resistance genes, use IslandViewer to locate toxicity islands, and the input data is the binned genome sequence. Combine CARD and ResFinder databases, use alignment tools (such as DIAMOND or BLAST) to screen resistance genes, and annotate their functional categories and resistance mechanisms. Use IslandViewer to locate toxicity islands and analyze their distribution patterns and functional types of carried genes. Combine PlasmidFinder to annotate plasmid functional genes and identify transposon nested regions, and use the ConjScan tool to analyze whether toxic gene islands are transmitted through plasmids or transposons. Use evolutionary analysis tools (such as RAxML or IQ-TREE) to compare the conservation of toxic gene islands in different strains, infer their horizontal gene transfer history and diffusion paths, and provide transmission paths and risk reports for biosafety assessments. Outputs include toxicity island distribution maps, transmission path analysis, and horizontal transfer history assessments.
[0120] 3. In the identification of transformation features, focus on analyzing exogenous insertion features such as CRISPR-Cas, artificial promoters, reporter genes, and codon optimization. The input data is the target genome sequence. Use CRISPRCasFinder or CRISPRDetect to detect the CRISPR cluster and identify whether it contains spacer sequences derived from exogenous DNA. By comparing the spacer sequence with the reference database, it is inferred whether its source is related to the engineered DNA. At the same time, common artificially designed markers are identified, including specific promoter (such as T7, lac) sequences, reporter genes (such as gfp), and resistance marker genes (such as lacZ, hygR). By analyzing the codon usage frequency of the target sequence, regions that are significantly different from the host preference are detected, and potential traces of artificial transformation are inferred. Combined with a deep learning model, a data set based on artificially designed sequence features is trained to enhance recognition accuracy. The output includes transformation feature annotation, codon usage analysis results, and artificially designed region predictions.
[0121] 4. Plasmid identification and transmission risk assessment are important links in transformation management and control. The input data is short read length or contig sequence. Potential plasmid fragments are assembled by PlasmidSPAdes, and the deep learning model is used in combination with PlasFlow to predict the plasmid or chromosome origin of each contig. ConjScan is used to analyze whether the plasmid carries conjugated transfer elements (such as the Tra gene cluster) and evaluate the transfer ability of the plasmid. The plasmid data from multiple environmental samples are integrated to construct a plasmid transfer network model, and the resistance genes, toxic genes carried by the plasmid and their transfer ability are analyzed. Combining these data, a quantitative assessment model for transmission risk is constructed to provide a scientific basis for the risk of plasmid transmission. The output includes a plasmid transfer network, a risk assessment model, and an identification of high transmission risk areas.
[0122] 5. Construction of a phylogenetic tree is a key step in confirming the relationship between the target strain and known pathogens. The input data is the 16S rRNA or core genome sequence in the target genome. Sequence alignment is performed using MAFFT or Clustal Omega, and a phylogenetic tree is constructed using IQ-TREE or RAxML to analyze the relationship between the target strain and known pathogens. Combined with plasmid analysis, gene transfer events between the target strain and the pathogen are inferred, specific toxicity gene clusters are located, and their genomic similarities are analyzed. The laboratory validation step combines model host experiments to evaluate the infectivity of the target strain and its toxic effects on different hosts. Outputs include a high-resolution phylogenetic tree, gene transfer event analysis, and a biosafety risk assessment report.
[0123] 6. Use InterProScan and Pfam to annotate toxicity-related functional domains (such as T3SS secretion system, RTX toxin). The input data is the genome or transcriptome sequence of the target strain. Use InterProScan and Pfam to annotate toxicity-related functional domains (such as T3SS secretion system, RTX toxin), combine STRING to construct a toxic protein interaction network, and analyze the toxicity mechanism and the functional association between key proteins. Combine qPCR and transcriptome data to analyze the gene expression change pattern of the target strain under antibiotic, heavy metal or high salinity pressure. Use KEGG or MetaCyc to annotate metabolic pathways and explore the environmental adaptability mechanism of the target strain. Outputs include toxicity mechanism annotation report, gene expression change map and adaptive metabolic pathway model.
[0124] Among them, in some embodiments, the early warning subsystem includes a risk assessment and early warning system design unit, a dynamic research and transmission analysis unit, an assessment model construction unit, a pathogen evolution and function research unit, a toxicity and transmission risk comprehensive prediction unit, a pathogen risk monitoring network unit, and an unknown pathogen and potential risk identification unit.
[0125] like Figure 2 As shown, the risk assessment and early warning system design unit integrates a variety of information output from previous analyses, including annotation results of resistance genes and virulence factors (from CARD and VFDB database comparisons), the distribution of toxicity islands and gene interaction network analysis results, as well as gene abundance data and temporal dynamic change information. Gene function annotations (such as toxin genes and resistance genes) provide basic assessment information for potential gene risks. Combined with the transmission path analysis of HGT elements and the prediction of plasmid transfer ability, the potential threat of gene spread is further quantified. The system integrates this information through a unified interface, uses deep learning models (such as Transformer) to analyze the evolutionary relationships and emerging pathogenicity possibilities of pathogens in environmental samples, and generates quantitative risk level reports and visual outputs to support the core algorithms of the early warning system.
[0126] The dynamic research and transmission analysis unit collects samples at multiple time points to study the dynamic changes of microbial communities and functional genes, analyzes the relationship between environmental parameters (such as temperature, humidity, and chemical pollutant concentrations) and the abundance of pathogens, and explores the impact of seasonal or sudden events on the spread of pathogens. Combined with traceability analysis technology, the transmission path and diffusion pattern of pathogens are simulated through genome mutation patterns and diversity analysis. Further combined with socioeconomic data (such as population density and land use) to assess the transmission risk, and predict the diffusion trend of pathogens in different environments, providing dynamic risk assessment and real-time early warning. Inputs include the community structure dynamics of environmental samples at multiple time points (outputs from binning analysis and time series models), the diffusion pattern of toxic gene islands and resistance genes, and socioeconomic data (such as population density and land use). Based on the phylogenetic tree and transmission path model (results generated by IQ-TREE and the transmission path prediction module), the transmission trajectory of the pathogen is simulated, and its correlation with the abundance of pathogens is analyzed in combination with environmental parameters (such as temperature and humidity). Further combining gene abundance dynamics (from dynamic sampling analysis) and gene diffusion networks (such as plasmid network models) can predict the impact of environmental emergencies on microbial transmission and provide dynamically updated transmission trend analysis for real-time early warning.
[0127] The assessment model construction unit integrates key indicators such as microbial species, virulence factors, and resistance gene abundance to construct a multidimensional risk assessment model. Combined with the expert scoring system, the risk factors are quantified into a matrix (pathogen species × resistance level × environmental concentration × time). The model construction integrates the key data previously output, including the abundance, distribution and transmission path of resistance genes and toxic genes (such as the gene transfer ability assessment and gene abundance data generated in the risk assessment module), host adaptability analysis results (gene interaction network constructed by GNN), and gene expression potential prediction (through DeepGOPlus annotation functional domain). Based on these data, machine learning algorithms (such as random forests and support vector machines) are used to train the risk prediction model. Combined with multi-time point sampling data (such as dynamic changes in microbial community structure and environmental parameters), the model's ability to identify high-risk areas and key pathogens is optimized to generate a multidimensional risk assessment matrix. By building a real-time online system, the analysis results are dynamically updated and a visual risk assessment report is provided to users.
[0128] The pathogen evolution and function research unit combines molecular evolutionary trees and phylogenetic trees to study the evolutionary history of emerging pathogens, identify mutation hotspots (such as virulence islands and resistance islands), and analyze the potential impact of mutations on function. The input data comes from the distribution analysis of toxicity islands and resistance islands, the evolutionary trajectory of HGT elements, and the host interaction pattern (from phylogenetic trees and gene transfer pathway analysis). The molecular evolutionary tree (generated by RAxML or BEAST) and the results of mutation hotspot identification are used to study the evolutionary history of emerging pathogens. Combined with host adaptability analysis (through gene co-expression network modeling) and gene mutation function prediction (such as protein structure generated by AlphaFold), the potential impact of mutations on pathogenicity and transmission ability is evaluated. Integrate exogenous data in the geographic model (such as climate change and population mobility) with transmission risk prediction to infer the future spread trend of pathogens and provide intervention recommendations. Outputs include mutation function evaluation, evolutionary pattern analysis, and risk prediction reports.
[0129] The comprehensive prediction unit for toxicity and transmission risk combines epidemiological transmission models with deep learning technology (such as DeepSIR) to predict the diffusion trajectory of toxic genes or resistance genes. Input data include the diffusion trajectory of toxic genes and resistance genes (from the DeepSIR model), environmental dynamic variables (such as pollutant concentrations, climatic conditions), toxic island transmission risk analysis, and modeling results of the correlation between gene diffusion rate and environmental factors. The risk level of toxic genes to the environment and hosts is comprehensively quantified through decision trees and random forest models, and a comprehensive risk assessment framework is constructed in combination with dynamic time series data (such as abundance dynamics analysis). Model outputs include the transmission pattern of toxic genes, future diffusion predictions, and hazard range analysis.
[0130] The pathogen risk monitoring network unit integrates all the above output data by integrating AI technology with metagenomic data, including gene annotation results (CARD, VFDB), functional domain prediction (InterProScan, DeepGOPlus), transmission risk assessment (toxic gene transmission model, phylogenetic tree analysis), artificially designed feature recognition (DeepSynBio analysis), dynamic transmission paths (generated by LSTM and GRU) and environmental variable influences (generated by GIS and distribution modeling). Using Transformer and multimodal deep learning frameworks, gene sequences, time dynamic data, geographic spatial information and functional annotations are unified into a model to construct a pathogen risk monitoring system. The system supports real-time monitoring, dynamic threshold adjustment and automatic early warning, generates full life cycle risk assessment reports, and provides timely intervention recommendations and policy support for decision makers. Outputs include real-time risk monitoring results, dynamic transmission analysis and a comprehensive risk assessment framework.
[0131] The input information of the unknown pathogen and potential risk identification unit includes the distribution and transmission network analysis results of toxic genes, resistance genes, and HGT elements (generated by GNN and toxic transmission risk modules), as well as the identification results of artificial modification traces (through DeepSynBio analysis of codon optimization and artificial design markers). Deep learning models (such as Transformer) identify unknown pathogenic microorganisms and their interaction patterns with host microbiome by learning gene sequences and functional domain features. Combining GAN to generate artificial gene sequences and compare them with natural sequences further improves the recognition accuracy of artificially designed features. Reinforcement learning models optimize the prediction of epidemic transmission paths and generate spatiotemporal dynamic infection risk maps for simulating the functional impact of gene mutations and environmental adaptability analysis.
[0132] By combining AI technology with metagenomic analysis, the early warning subsystem of this application achieves a seamless transition from basic research to practical application. The automated analysis process can rapidly identify potential pathogens and high-risk genes, and combine environmental variables and spatiotemporal dynamics to predict transmission and quantify risk. Especially when faced with complex pathogenic microbial communities and dynamic environments, the deep integration of AI greatly enhances the system's analytical efficiency and predictive capabilities, providing strong technical support for global biosafety monitoring.
[0133] The risk assessment mechanism is implemented by constructing a multidimensional model. This model is based on key indicators such as microbial species and virulence factors, and is trained and optimized using machine learning algorithms to identify high-risk areas and key pathogens. The pathogen evolution research and toxicity prediction unit further enriches the dimensions of risk assessment. The warning trigger conditions rely on the outputs of the dynamic research and transmission analysis unit and the pathogen risk monitoring network unit. These units monitor pathogen risks in real time by collecting samples and integrating AI technology, and trigger warnings when the data exceeds preset thresholds. The automated analysis process of the assessment model construction unit ensures the accuracy and timeliness of risk assessments. The pathogen risk monitoring network unit is responsible for dynamically adjusting the warning threshold to improve the sensitivity of the warning system and reduce false alarms.
[0134] In summary, the technical effects of this application are mainly reflected in the following points:
[0135] 1) We have built a fully integrated metagenomic analysis platform, enabling a complete analytical process from quality control, assembly, and functional annotation of second- and third-generation sequencing data to biohazard assessment. By integrating multiple analysis steps and tools, we eliminate the need for users to switch between different software programs, ensuring a seamless workflow.
[0136] 2) Developed hybrid processing capabilities tailored to the characteristics of second-generation and third-generation sequencing data, and innovatively developed new algorithms and optimized processes for data integration, assembly, and correction, taking into account the high accuracy of second-generation data and the long read length advantages of third-generation data.
[0137] 3) In response to the needs of biohazard monitoring and prevention and control, dedicated analysis modules have been developed, including pathogen tracking, resistance gene screening, and functional annotation of genes related to pollutant degradation.
[0138] 4) Machine learning and intelligent algorithms have been introduced, and automated parameter optimization functions and analysis recommendation systems have been developed to assist users in selecting the most appropriate analysis process based on data characteristics.
[0139] 5) It provides the integration and visualization function of data results, realizes the integration and dynamic visualization of multi-dimensional results, enables the intuitive display of analysis results (such as microbial community structure, functional gene distribution, pathogen abundance, etc.), and supports the generation of biohazard assessment reports.
[0140] 6) A modular architecture ensures a highly compatible design, facilitating flexible configuration of analysis functions by users while ensuring compatibility with existing tools (such as SPAdes, MEGAHIT, etc.) and databases (such as CARD, KEGG, etc.).
[0141] 7) It provides a user-friendly graphical user interface that supports visual analysis process configuration and result display, reducing the difficulty of use for non-professional users.
[0142] 8) Provides a risk assessment module for biological hazards and environmental pollution, generates intuitive assessment reports based on analysis results, and provides decision support for users.
[0143] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A biological hazard big data analysis and monitoring early warning system, characterized in that: include: Graphical interface module, including main interface, parameter configuration interface, result display interface and status monitoring interface; The main interface displays the analysis process framework diagram, the parameter configuration interface is used to configure parameters for the analysis steps, the result display interface is used to display visual charts and log information, and the status monitoring interface is used to dynamically display the progress bar and task status; The analysis subsystem includes a core logic module, a unit analysis module, and a microbial community integrated analysis module; the core logic module includes a quality control unit, an error correction unit, an assembly unit, and a binning unit, and is used to process and optimize sequencing data; the sequencing data is the raw data obtained from biological samples through high-throughput sequencing technology; the unit analysis module includes a pathogen analysis unit, a resistance gene unit, a virulence assessment unit, and an evolutionary development unit, and is used to analyze the sequencing data to obtain information on pathogens, resistance genes, and virulence factors in biological samples; the microbial community integrated analysis module includes a traceability analysis unit, a mutation feature unit, a transmission evolution unit, and a transformation control unit, and is used to integrate the analysis results of each unit in the unit analysis module to analyze the structure, functional characteristics, and interactions of the microbial community; The early warning subsystem includes a risk assessment and early warning system design unit, a dynamic research and transmission analysis unit, an assessment model construction unit, a pathogen evolution and function research unit, a toxicity and transmission risk comprehensive prediction unit, a pathogen risk monitoring network unit, and an unknown pathogen and potential risk identification unit.
2. A biological hazard big data analysis and monitoring early warning system according to claim 1, characterized in that: Also includes: The Docker containerized environment module is used to build independent container images for each analysis unit in the unit analysis module, and deploy a server as the control center in each independent container image; The server receives instructions from the graphical interface through the application program interface and calls the analysis tool in the container image; The analysis tools in the container image include: FastQC, Trimmomatic and nanopore used in the quality control unit, FMLRC and Pilon used in the error correction unit, SPAdes and MEGAHIT used in the assembly unit, and MetaBAT2 and CheckM used in the binning unit.
3. A biological hazard big data analysis and monitoring early warning system according to claim 2, characterized in that: The core logic module includes: A quality control unit is used to generate a quality assessment report using FastQC and to denoise the quality assessment report using Trimmomatic and Cutadapt to obtain denoised data; the quality assessment report includes sequence quality distribution, GC content, and sequence length distribution; The error correction unit is used to correct long read errors in the denoised data through FMLRC, and to correct mismatches or insertions after genome assembly in combination with Pilon to obtain corrected data; The assembly unit is used to integrate the corrected data and the denoised data based on the OLC algorithm or the deBruijn graph algorithm to generate contigs data; The binning unit is used to call MetaBAT2 or CheckM tools to perform binning operations on contigs data, classify contigs data into specific microbial genomes, and evaluate the integrity and contamination rate of the genome in the contigs data.
4. A biological hazard big data analysis and monitoring early warning system according to claim 3, characterized in that: The unit analysis module includes: Pathogen analysis unit, which classifies the composition of microorganisms in environmental sample data into species using Kraken, PhlAn, and Centrifuge, and predicts gene regions of contigs data using Prokka, GeneMark, or Prodigal; Resistance gene unit, resistance genes in environmental sample data were compared by BLAST, RGI and CARD databases, and unknown resistance genes were predicted in combination with Prokka and InterProScan; Virulence assessment unit, which compares virulence factors in contigs data with the VFDB database and predicts virulence protein structures using Prokka and SWISSMODEL; Evolutionary developmental units, phylogenetic trees were constructed based on the 16S rRNA gene, and gene selection pressure was evaluated by dN / dS calculation.
5. A biological hazard big data analysis and monitoring early warning system according to claim 4, characterized in that: The pathogen analysis unit comprises: A sequence feature extraction subunit based on CNN and Transformer; the sequence feature extraction subunit is used for pathogenicity prediction; A time series analysis subunit based on RNN or LSTM, which is used for modeling microbial community dynamics; A GNN-based protein interaction network analysis subunit, wherein the protein interaction network analysis subunit is used for virulence factor network topology modeling.
6. A biological hazard big data analysis and monitoring early warning system according to claim 4, characterized in that: The resistance gene unit comprises: Plasmid-borne resistance gene subunits were identified through PlasmidFinder, key feature subunits of resistance genes were screened through RandomForest or SVM models, and multidrug resistance analysis subunits were integrated with CARD, ResFinder and AMRFinder databases.
7. The biological hazard big data analysis and monitoring early warning system according to claim 4, characterized in that: The toxicity assessment unit comprises: The subunits include a sequence feature screening subunit based on kmer frequency and Shannon entropy, a virulence factor local feature extraction subunit based on Onehot encoding and CNN, and a dynamic analysis subunit that constructs a virulence factor interaction network through GNN and identifies key nodes.
8. The biological hazard big data analysis and monitoring early warning system according to claim 1, characterized in that: The microbial community integrated analysis module includes: The traceability analysis unit is used to detect SNPs / Indels using GATK or bcftools and construct pathogen geographic distribution maps in combination with GIS tools; Mutation signature unit, used to detect point mutations and structural variations using GATK and Manta, and to perform functional annotation in conjunction with ANNOVAR; The propagation evolution unit is used to construct phylogenetic trees using IQtree and analyze temporal dynamic propagation paths in combination with BEAST; The control unit was modified to perform binning classification using MetaBAT2 and predict plasmid origin using PlasFlow.
9. The biological hazard big data analysis and monitoring early warning system according to claim 1, characterized in that: The graphical interface module also includes: The user management interface is used to manage the permissions and access control of different users; the user management interface supports multi-level permission settings, including administrators, researchers and ordinary users. Administrators can configure system parameters and tool modules, researchers can perform analysis tasks and view results, and ordinary users can view public analysis results and log information.
10. The biological hazard big data analysis and monitoring early warning system according to claim 1, characterized in that: Also includes dynamic expansion modules; The dynamic expansion module includes: The plugin management submodule is used to manage and load third-party plugins to expand the system's analysis capabilities. The plugin management submodule supports installing, updating, and uninstalling plugins through a graphical interface or command line. Plugins are available in JAR, Python script, or Docker image formats. The resource monitoring submodule is used to monitor the usage of system resources in real time, including CPU, memory and storage space. The resource monitoring submodule displays resource usage through a graphical interface and issues warnings when resources are insufficient. It supports automatic expansion of cloud resources or manual adjustment of resource configuration.
Citation Information
Cited By
Single molecule signal decoding method based on artificial intelligence algorithm
CN121148509A
Detection method of coronavirus sample
CN121171339A
Quantitative assessment method and system for pathogen transmission risk
CN121171645A
A method and system for quantitatively assessing the risk of pathogen transmission.
CN121171645B
Data management and quality control system for virus discovery and pollution traceability
CN121215039A