Antibiotic resistance gene end-to-end detection system and method suitable for multiple sequencing platforms
By combining one-stop analysis of metagenomic data from multiple platforms with deep learning models, the accuracy and standardization issues of antibiotic resistance gene detection have been resolved, enabling efficient and accurate identification and mechanism analysis of resistance genes, which is suitable for rapid diagnosis and monitoring in clinical and public health fields.
Patent Information
- Application Number
- CN202511508795.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-02-24
AI Technical Summary
Existing antibiotic resistance gene detection methods suffer from insufficient ability to identify novel resistance genes and a lack of standardization and uniformity in the analysis process, resulting in insufficient accuracy and reliability, making it difficult to meet the needs of rapid clinical diagnosis and precise treatment.
We employ a one-stop analysis process for metagenomic data across multiple platforms and a deep learning-based method for detecting antibiotic resistance genes, including adaptive processing workflows, containerized deployment, and quality monitoring. By combining a multi-scale residual neural network module and a bilinear classification head, we can achieve simultaneous prediction of antibiotic resistance categories and mechanisms.
It significantly improves the accuracy and efficiency of drug resistance gene detection, can identify drug resistance in novel and unknown pathogens, provides standardized analysis procedures and highly sensitive detection results, is suitable for resource-constrained environments, and supports personalized clinical treatment and public health surveillance.
Smart Images

Figure CN121565252A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of deep learning and bioinformatics, specifically to the field of antibiotic resistance gene detection, and particularly to a one-stop method for detecting antibiotic resistance genes based on a deep learning framework for metagenomic data. Background Technology
[0002] Antibiotic resistance has become a major challenge in the treatment of infectious diseases worldwide. It not only increases the difficulty of treating infectious diseases, prolongs the course of illness, and raises medical costs, but can also lead to the serious consequence of having no effective treatment options, posing a grave threat to human health. Therefore, accurate, rapid, and low-risk detection of antibiotic resistance genes (ARGs) is crucial for understanding rational drug use in clinical practice, controlling the spread of drug-resistant bacteria, and developing effective infection control strategies.
[0003] In recent years, significant progress has been made in next-generation sequencing (NGS) technologies, represented by the Illumina platform, and third-generation sequencing (TGS) technologies, represented by the Nanopore and PacBio platforms. NGS technologies, with their high precision and high throughput, have played a crucial role in pathogen identification and drug resistance gene detection; while third-generation sequencing technologies, with their ultra-long read lengths, offer the possibility of solving complex genome structure problems. The continuous development of these technologies has significantly enhanced the potential of metagenomics in pathogen identification and drug resistance gene detection.
[0004] Metagenomic sequencing, as an emerging detection method, has brought new hope to the detection of antibiotic resistance genes. This technology directly performs high-throughput sequencing on all genetic material in clinical samples (such as blood, body fluids, and tissues), without relying on traditional microbial isolation and culture processes. It can comprehensively analyze the composition of pathogenic microorganisms and the characteristics of their antibiotic resistance genes in the samples. This method provides important technical support for the rapid diagnosis of clinical infectious diseases, pathogen tracing, and drug resistance monitoring.
[0005] Metagenomic sequencing technology, through high-throughput sequencing of all genetic material in clinical samples (such as blood, body fluids, and tissues), can comprehensively analyze the composition of pathogenic microorganisms and the characteristics of their antibiotic resistance genes (ARGs) without relying on traditional microbial isolation and culture processes. This method provides important technical support for the rapid diagnosis of clinical infectious diseases, pathogen tracing, and drug resistance monitoring. In recent years, the development of next-generation sequencing (NGS) technology, represented by the Illumina platform, and third-generation sequencing (TGS) technology, represented by the Nanopore and PacBio platforms, has significantly enhanced the potential of metagenomics in pathogen identification and drug resistance gene detection due to its continuously improving sequencing throughput, read length (especially for TGS), and cost-effectiveness.
[0006] However, existing metagenomic sequencing technologies still face a series of significant challenges when applied to antibiotic resistance gene detection in clinical settings. One of these challenges is the integration and standardization of data from different sequencing platforms. Different high-throughput sequencing platforms (such as Illumina's short-read, high-precision data and Nanopore's long-read, but relatively high-error-rate data) exhibit significant differences in sequencing read length, error profiles, sequencing depth, and coverage uniformity. This makes it difficult to effectively integrate and compare data from different platforms or batches of experiments, limiting the universality of analytical results and standardization in clinical applications. Furthermore, current antibiotic resistance gene detection primarily relies on sequence alignment (e.g., using tools like BLAST) between the sequenced reads or assembled gene sequences and known antibiotic resistance gene databases (such as CARD and ResFinder). Clearly, such methods can only detect ARGs that are highly similar to known ARG sequences, while many functionally similar ARGs that do not exhibit significant sequence similarity remain unidentified. Furthermore, while drug resistance gene databases collect a large number of known drug resistance gene sequences, there are numerous types of drug resistance genes in nature, and new drug resistance genes are constantly emerging. The sequences currently in the databases represent only a small fraction and cannot cover all potentially existing drug resistance genes. These factors lead to a high false-negative rate in drug resistance gene testing, posing numerous risks to clinical diagnosis and treatment.
[0007] To overcome the bottlenecks of traditional analytical methods, deep learning (DL) technology offers new possibilities for extracting useful information from complex genomic data. Deep learning models, especially deep neural networks (DNNs), have the ability to automatically learn and extract abstract features from raw sequence data, theoretically reducing complete reliance on prior knowledge (such as known databases of drug resistance genes). Although deep learning has shown great potential in bioinformatics, its application in metagenomic drug resistance gene detection is still in the exploratory stage and faces some inherent problems, such as: the scarcity of high-quality, large-scale labeled datasets, which directly affects the training effect and generalization ability of models; deep learning models are often regarded as "black boxes," with their internal decision-making processes lacking sufficient biological interpretability, making it difficult for clinicians to fully trust them; and training and running complex deep learning models usually requires a lot of computational resources and time, which may limit their application in resource-constrained environments.
[0008] Therefore, there is an urgent need to develop a novel end-to-end antibiotic resistance gene detection technology that can be applied to sequencing data from different sequencing platforms, effectively overcome the limitations of traditional alignment methods, and utilize advanced algorithms (such as optimized and interpretable deep learning models) to construct an efficient, accurate, and standardized antibiotic resistance gene detection process, thereby better meeting the needs of rapid clinical diagnosis and precision treatment. Summary of the Invention: This invention aims to address the following problems existing in current methods for detecting antibiotic resistance genes: 1) Insufficient ability to identify novel drug resistance genes: Currently, mainstream drug resistance gene detection methods mainly rely on sequence alignment with known drug resistance gene databases. However, this method has significant limitations when detecting novel drug resistance genes or distantly homologous genes with low similarity to known gene sequences. Specifically, it results in a low detection rate, a high false negative rate, and an inability to accurately identify all potential drug resistance genes, thus affecting the accuracy and reliability of drug resistance detection and the formulation and implementation of drug resistance control strategies.
[0009] 2) Lack of Standardization and Uniformity in Analytical Processes: Currently, there is no standardized, end-to-end bioinformatics analysis workflow for data generated by different sequencing platforms (e.g., second-generation sequencing, third-generation sequencing) or for different analytical objectives (e.g., identification of known drug resistance genes, prediction of drug resistance categories). In practice, users often need to combine multiple bioinformatics tools, which not only makes the process complex and cumbersome but also results in poor comparability of the final analysis results due to the lack of compatibility and standardization between different tools, making it difficult to meet the needs of large-scale, efficient, and accurate drug resistance analysis.
[0010] To address the shortcomings of existing technologies, such as insufficient identification capabilities for novel antibiotic resistance genes and a lack of standardization and uniformity in analytical procedures, this invention provides a one-stop method for detecting antibiotic resistance genes. This method includes a novel antibiotic resistance gene prediction algorithm based on deep learning, and a standardized antibiotic resistance gene analysis process based on metagenomic data.
[0011] The specific technical solution is as follows: An end-to-end detection method for antibiotic resistance genes applicable to multiple sequencing platforms includes a one-stop analysis step for multi-platform metagenomic data and a deep learning-based antibiotic resistance gene detection step, wherein... The one-stop analysis steps for multi-platform metagenomic data include: adaptively determining the sequencing platform based on the input sequencing data format and selecting the appropriate processing flow: for second-generation sequencing data, the sequence is preprocessed, the preprocessed sequence is aligned to the host reference genome to remove host contamination, and initial contigs are generated using a short read assembly tool; for third-generation sequencing data, after sequence preprocessing, the preprocessed sequence is self-aligned to generate an overlap correction file, then an overlap cluster backbone is constructed, and finally three iterations of correction are performed to generate initial contigs. After obtaining the initial contigs, species-level classification and annotation were performed using the Kraken2 tool to provide species background information for antibiotic resistance gene detection. Then, the Prodigal tool was used to identify prokaryotic ORFs. After obtaining the corresponding ORFs, the translated protein sequences were used in the antibiotic resistance gene detection step. The deep learning-based antibiotic resistance gene detection step includes: taking the protein sequence as input, fine-tuning a pre-trained protein language model, which is combined with a multi-scale residual neural network module to improve the model's ability to extract complex global and local high-dimensional features related to drug resistance; and using a bilinear classification head to achieve simultaneous prediction of antibiotic resistance categories and antibiotic resistance mechanisms.
[0012] Preferably, the one-stop analysis step of multi-platform metagenomic data is implemented using the following method: A workflow engine is built, using Snakemake to construct a task dependency graph, and dynamically scheduling sequence preprocessing, sequence assembly, and ORF prediction processes based on the input data format. Containerized deployment, with each module packaged into a Docker image, supports cross-platform, dependency-free operation; Quality monitoring automatically generates quality reports at key steps.
[0013] Preferably, the fine-tuned pre-trained protein language model includes, The input layer takes one-hot encoded protein sequences as input. The feature extraction layer inserts LoRA modules into each Transformer layer of the ESM2 model. By adding low-rank update paths to query, key, and value projections, parameter fine-tuning is achieved, resulting in a fine-tuned ESM2 model. The embedded representation of the fine-tuned ESM2 model is then input into the multi-scale residual neural network module. The global average pooling and normalization layer performs global average pooling on the sequence output by the feature extraction layer along the sequence dimension, and then performs normalization through LayerNorm to output dimensionality-reduced sequence features. The classification layer uses a bilinear classification head to process the dimensionality-reduced features in parallel, and outputs the drug resistance category probability and the drug resistance mechanism probability, respectively.
[0014] Preferably, training the fine-tuned pre-trained protein language model includes, A dataset was constructed by integrating and manually correcting five databases, including CARD, HMD-ARG, PLM-ARGDB, NCBI Bacterial Antimicrobial Resistance Reference Gene Database, and NCBI β-lactam Family Reference Gene Catalog, to form a model training set containing known antibiotic resistance gene sequences. A loss function is constructed. During the model training process, a Focal Loss function based on class weights and prediction confidence is used to optimize the task loss for predicting antibiotic resistance categories and resistance mechanisms, respectively, so as to improve the prediction accuracy of low-frequency categories. The formula for calculating Focal Loss is as follows:
[0015] Where is the positive sample weight parameter for drug resistance category or resistance mechanism, and γ is the focusing parameter. The total loss is calculated as follows:
[0016] During model training, the two tasks are jointly optimized through backpropagation to update the model parameters.
[0017] Preferably, after obtaining the contigs sequence, the Kraken2 tool is used for species-level classification annotation to provide species background information for antibiotic resistance gene detection; then, the Prodigal tool is used to identify prokaryotic ORFs; after obtaining the corresponding ORFs, the translated protein sequence is used for antibiotic resistance gene detection.
[0018] Preferably, the sequence preprocessing includes removing low-quality sequences using a dynamic parameter filtering tool. For second-generation sequencing data, the dynamic parameter filtering tool is Trimmomatic and / or FastP; for third-generation sequencing data, the filtering tool is a cascaded filtering tool, using PoreChop to remove adapters and NanoFilt for quality filtering, respectively.
[0019] This invention also provides a multi-platform end-to-end detection system for antibiotic resistance genes, which uses the method described above to simultaneously predict antibiotic resistance categories and mechanisms. It includes a one-stop analysis module for multi-platform metagenomic data and a deep learning-based antibiotic resistance gene detection module. The multi-platform metagenomic data one-stop analysis module adaptively determines the sequencing platform and selects the appropriate processing flow based on the input sequencing data format: for second-generation sequencing data, the sequence is preprocessed, the preprocessed sequence is aligned to the host reference genome to remove host contamination, and initial contigs are generated using a short read assembly tool; for third-generation sequencing data, after sequence preprocessing, the preprocessed sequence is self-aligned to generate an overlap correction file, then an overlap cluster backbone is constructed, and finally three iterations of correction are performed to generate initial contigs. After obtaining the initial contigs, species-level classification annotation was performed using the Kraken2 tool to provide species background information for antibiotic resistance gene detection. Then, the Prodigal tool was used to identify prokaryotic ORFs. After obtaining the corresponding ORFs, the translated protein sequences were input into the antibiotic resistance gene detection module. The deep learning-based antibiotic resistance gene detection module takes the protein sequence as input and fine-tunes a pre-trained protein language model. This model, combined with a multi-scale residual neural network module, improves the model's ability to extract complex global and local high-dimensional features related to resistance. Furthermore, a bilinear classification head is used to simultaneously predict antibiotic resistance categories and mechanisms.
[0020] Preferably, the present invention also provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multi-platform antibiotic resistance gene end-to-end detection method as described above.
[0021] Preferably, the present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the multi-platform antibiotic resistance gene end-to-end detection method as described above.
[0022] Preferably, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the multi-platform antibiotic resistance gene end-to-end detection method as described above.
[0023] The one-stop antibiotic resistance gene detection method provided by this invention significantly improves the accuracy, efficiency, and comprehensiveness of resistance gene detection by integrating multi-platform metagenomic sequencing data processing with deep learning-based resistance prediction and mechanism identification technologies. It is particularly suitable for detecting resistance to emerging and unknown pathogens and has the following beneficial effects: 1) Standardized and Automated Analysis Workflow: This invention constructs a standardized, end-to-end analysis workflow for second- and third-generation sequencing data, integrating steps such as data quality control, host removal, sequence assembly, gene annotation, and species classification through the Snakemake automated toolchain. This workflow eliminates the complexity of manually combining multiple tools required in existing technologies, lowers the operational threshold, and significantly improves the comparability of results between different laboratories, providing reliable technical support for large-scale drug resistance monitoring.
[0024] 2) Enhancing the Detection Capability of Novel Drug Resistance Genes: Traditional alignment methods have limitations in identifying drug resistance genes because they heavily rely on known databases, making it difficult to effectively identify novel drug resistance genes or those with low similarity to known genes. This invention employs a deep learning model to directly extract features from protein sequences, eliminating the dependence on sequence alignment and enabling efficient detection of novel and distantly related drug resistance genes. Specifically, it fine-tunes the pre-trained ESM2 protein language model, combines it with the DRSN module to extract drug resistance-related local and global features, and then uses a bilinear classification head to infer the drug resistance category and mechanism in parallel. Simultaneously, a Focal Loss loss function is introduced for optimization to address the problem of imbalanced sample distribution and maintain high model sensitivity. This invention achieves accurate identification of non-classical drug resistance mechanisms, reduces the false negative rate, and provides more accurate detection evidence for drug resistance prevention and control.
[0025] 3) Multi-task learning and efficient parameter optimization: This invention employs a bilinear classification head to process drug resistance category and mechanism prediction tasks in parallel, achieving multi-task learning through a shared underlying feature extraction network, thus avoiding the overfitting risk of single-task models. Simultaneously, a LoRA low-rank adaptation module is introduced into the ESM2 model, enabling efficient fine-tuning with a minimal number of trainable parameters (rank r=4). This significantly reduces computational resource consumption while maintaining the generalization ability of the pre-trained model, making it suitable for resource-constrained clinical or field testing scenarios.
[0026] 4) Significant Clinical Application Value: By improving the sensitivity of drug resistance gene detection and the comprehensiveness of mechanism analysis, the method of this invention provides more precise medication guidance for clinical anti-infective treatment. For example, in rapidly identifying multidrug-resistant strains or novel drug resistance mechanisms, this invention can assist doctors in developing personalized treatment plans and reducing the risk of antibiotic overuse. In the field of public health, its standardized procedures and high-throughput analysis capabilities help build regional drug resistance monitoring networks, providing data support for the prevention and control of drug-resistant bacteria transmission.
[0027] In summary, this invention has solved the key bottlenecks in the existing technology through technological innovation, and has made breakthroughs in the accuracy, efficiency and depth of mechanism analysis of drug resistance gene detection, which has significant clinical application value and social benefits. Attached Figure Description
[0028] Figure 1 Table of main open-source analysis tools used in this invention, their version numbers, and functional classifications. Figure 2 Flowchart of the technical method of this invention Figure 3 : Schematic diagram of the structure of the deep learning model for drug resistance prediction used in this invention Detailed Implementation
[0029] This application contains software versions (such as...) Figure 1 (As shown), different versions may result in different data processing results; please download and install the provided version whenever possible. The following will describe the implementation scheme of this application in detail with reference to examples; however, those skilled in the art will understand that the examples below are for illustrative purposes only and should not be considered as limiting the scope of this application. Unless otherwise defined below, all technical and scientific terms used in the specific embodiments of this application are intended to have the same meaning as commonly understood by those skilled in the art.
[0030] Figure 2 This is a system flowchart illustrating the sub-processes of preprocessing, assembly, annotation, and prediction of second-generation (Illumina) and third-generation (Nanopore) sequencing data, as well as Snakemake as the scheduling engine for overall process management. The present invention is implemented using the following architecture.
[0031] 1. End-to-end metagenomic drug resistance monitoring system architecture This system provides a one-stop processing workflow from raw data to gene annotation for metagenomic sequencing data generated by second-generation (represented by the Illumina sequencing platform) and third-generation (represented by Nanopore sequencing technology) sequencing platforms. The workflow is built using Snakemake and automatically selects the appropriate analysis workflow based on the sequencing data format input by the user. Specifically, it consists of the following modules: Data input module: Receives raw sequencing data (FastQ format) generated by second-generation (such as Illumina) or third-generation (such as Nanopore) sequencing platforms. Multi-platform sequencing data processing module: performs quality control, host genome removal, metagenomic assembly, and ORF prediction on sequencing data to generate candidate gene sequences; The deep learning-based drug resistance prediction and mechanism identification module: Based on a deep learning multi-task model, it classifies candidate genes into drug resistance categories and infers drug resistance mechanisms.
[0032] 2. Metagenomic Data Processing Workflow Processing workflow for second-generation metagenomic sequencing data: Data quality control: Fastp tool is used for connector identification and low-complexity sequence filtering, and Trimmomatic is used to remove low-quality sequences (bases with a Phred quality value <20) to ensure data quality.
[0033] Host removal: Bowtie2 was used to align the quality-controlled sequences to the host reference genome (such as the human reference genome GRCh38, T2T-CHM13), and the aligned sequences were removed to reduce host sequence interference.
[0034] Sequence assembly: Run the MEGAHIT tool to assemble metagenomic data based on the iterative k-mer De Bruijn Graph algorithm to generate contig sequences.
[0035] Open reading frame (ORF) prediction: Kraken2 tool was used for species-level classification annotation to provide species background information for drug resistance gene detection. Prodigal was used to predict the ORF of the assembled contigs, filtering out ORFs with a length <100aa to reduce noise data and generate standardized protein sequences.
[0036] Processing workflow for third-generation metagenomic sequencing data: Data preprocessing: PoreChop was used to remove connector sequences, and NanoFilt was used to filter short reads (<1000bp) and low-quality sequences (average quality value <10) to improve data accuracy.
[0037] Sequence contig construction: Minimap2 is used for self-alignment to generate contig correction files, and Miniasm is used to construct the sequence contig skeleton.
[0038] Sequence correction: Iterative correction using Racon is applied to reduce the base error rate and improve sequence quality. The final contig sequence is then generated.
[0039] Open reading frame (ORF) prediction: Prodigal is used to predict the ORF of the assembled contigs, filtering out ORFs with a length <100aa to reduce noisy data and generate standardized protein sequences.
[0040] Standardized protein sequences are input into the antibiotic resistance gene detection module to infer the type of resistance gene and the resistance mechanism.
[0041] 3. Design of Deep Learning Model for Drug Resistance Prediction Step 3.1: Design the multi-task model architecture The model employs a dual-task network structure, such as Figure 3 As shown, the process structure is illustrated from protein sequence input, ESM2 encoding, LoRA fine-tuning, DRSN feature extraction, to parallel dual classifier output of antibiotic category and resistance mechanism, including resistance category classification (such as β-lactamases, macrolides, etc.) and resistance mechanism inference (such as antibiotic efflux and inactivation, etc.).
[0042] Specifically, the model includes the following structure: Input layer: protein sequence; Feature Extraction Layer: The pre-trained protein language model fine-tuned in this invention is based on the pre-trained protein language model (ESM2_t30_150M_UR50D), with a LoRA module inserted into each Transformer layer of the ESM2 model. The LoRA module achieves efficient parameter fine-tuning by adding a low-rank update path (rank r=4, scaling factor α=16) to the query, key, and value projections, avoiding the overfitting problem caused by overall fine-tuning.
[0043] The fine-tuned ESM2 model is embedded into a multi-scale residual neural network module (DRSN). This module contains three cascaded residual blocks, each employing a 1D convolution (kernel size = 3), and retaining semantic information from previous layers through residual connections. By utilizing the Trasformer layer in ESM2 and the convolutional layers in DRSN, the model can effectively extract local conserved regions and long-range dependencies, improving its ability to express complex resistant structural features.
[0044] Global average pooling and normalization: After the feature representation is processed by the DRSN module, global average pooling is performed along the sequence dimension, followed by LayerNorm normalization to improve the robustness of downstream classification tasks.
[0045] Classification layer: The reduced-dimensional features are processed in parallel using a bilinear classification head. The resistance category classifier outputs a 32-dimensional vector representing the probability of resistance to 32 antibiotics (e.g., β-lactams 0.02, macrolides 0.97). Simultaneously, the resistance mechanism classifier outputs an 8-dimensional vector representing the probability of 8 resistance mechanisms (e.g., efflux pump mechanism 0.01, antibiotic inactivation mechanism 0.96). The two prediction results are output in parallel.
[0046] Step 3.2: Model Training and Optimization Dataset: Resistance gene information from multiple public datasets was integrated and manually corrected, including CARD, HMD-ARG, PLM-ARGDB, the NCBI Bacterial Antimicrobial Resistance Reference Gene Database (BioProject ID: 313047), and the NCBI β-lactam family reference gene catalog (2024-07-22.1). Categories containing fewer than 10 training sequences were excluded. The final dataset contains 30,338 resistance gene sequences and tags, covering 32 resistance categories and 7 resistance mechanisms. This constitutes the model training set containing 4,293 known antibiotic resistance gene sequences.
[0047] Loss function: This includes a weighted sum of prediction losses for different drug resistance gene categories and drug resistance mechanisms. Both losses use the Focal Loss function, and the different drug resistance categories and mechanisms are empirically defined.
[0048] Focal Loss was defined for both the drug resistance category classification and drug resistance mechanism classification tasks, with a positive sample weight of 0.75. The total loss is the sum of the Focal Losses of the two classes, used to optimize the parameters of the bi-class classifier. The formula for calculating Focal Loss is as follows:
[0049] in, γ represents the positive sample weight parameter for drug resistance category or mechanism, and γ is the focusing parameter. The total loss is calculated using the following formula:
[0050] During model training, backpropagation is used to jointly optimize the two tasks and update the model parameters. The trained model will then be applied to the final antibiotic resistance gene detection task.
[0051] Fine-tuning strategy: Use the LoRA adapter to fine-tune the protein language model (ESM-2) to reduce the number of parameters (rank r=4, scaling factor α=16).
[0052] Step 3.3: Model Prediction Input the protein sequence into the trained drug resistance prediction model, and output a file containing the drug resistance category and the probability of the drug resistance mechanism. Categories or mechanisms with a probability greater than 0.5 are the prediction results.
[0053] 4. The one-stop process of this invention achieves full automation through the following methods: 1) Workflow Engine: Snakemake is used to build a task dependency graph and dynamically schedule sequence preprocessing, sequence assembly and ORF prediction (second generation sequencing data processing workflow or third generation sequencing data processing workflow) according to the input data format.
[0054] 2) Containerized deployment: Each module is packaged into a Docker image, supporting cross-platform (Linux / macOS) operation without dependencies.
[0055] 3) Quality monitoring: Automatically generate quality reports (such as sequencing data quality, assembly quality indicators, and model confidence distribution) after key steps (such as assembly and prediction). 5. The present invention will be further described below with reference to specific examples. Example 1: Taking Pasteurella multocida (SRR18605042) isolated from cattle as an example, this invention demonstrates the antibiotic resistance detection steps for Nanopore sequencing data: 1) Raw data quality assessment (step 2.1): Use the FastQC tool (adapted to Nanopore mode) to assess the quality of the raw sequencing files and generate a report including read length distribution, average quality value, and base error rate. The input file is sample_raw.fastq, and the output quality assessment report is sample_qc_report.html.
[0056] 2) Adapter removal and sequence filtering (step 2.1): Sequencing adapters are identified and removed using PoreChop. Then, NanoFilt is used to filter low-quality reads, removing sequences with a Q value <20 or a length <500bp.
[0057] 3) Sequence Assembly and Correction (Step 2.3): First, the filtered reads are self-aligned using minimap2 (parameter -x ava-ont) to generate an overlap relationship file for subsequent assembly. Then, initial contigs are constructed based on the overlap relationships using miniasm. Finally, Racon is used for three iterative corrections, with the error rate gradually decreasing after each correction, ultimately outputting a high-precision contigs file.
[0058] 4) Open reading frame identification (step 2.3): Use the Prodigal tool (parameter -p meta) to predict the prokaryotic ORF of the corrected contigs and generate protein sequence files.
[0059] 5) Antimicrobial resistance category and mechanism prediction (step 3.3): The predicted protein sequences were input into a deep learning-based antimicrobial resistance prediction model. After prediction, a total of 4659 ORFs were predicted for this bacterium, of which only one was predicted as an ARG (antimicrobial resistance category probability > 0.5, antimicrobial resistance mechanism probability > 0.5). The annotated antimicrobial resistance category was cephalosporin, and the predicted antimicrobial resistance mechanism was antibiotic inactivation. This result is consistent with data provided in related articles, indicating the reliability of the procedure.
[0060] Example 2: Using a carbapenem-resistant Gram-negative clinical isolate from South Korea (WGS number: SRR8289540) as the research subject, this example demonstrates a streamlined analysis process for antibiotic resistance detection based on its single-end next-generation sequencing data. 1) Raw Data Quality Assessment (Step 2.1): The raw sequencing data SRR8289540.fq.gz was quality controlled using the FASTP tool. FASTP automatically detected and removed adapter sequences and filtered out low-quality regions. The generated data file after quality control was SRR8289540_clean.fq.gz, and a quality assessment report, SRR8289540_qc.html, was also output. The report results showed that the effective data ratio was approximately 99%, indicating good sequencing quality.
[0061] 2) Host sequence removal (step 2.2): The quality-controlled sequences were aligned to the human genome (GRCh38) using the bowtie2 tool to remove host-derived sequences. The resulting non-host data was stored in the file SRR8289540_nonhost.fq. The host sequence removal rate was approximately 95%, significantly improving the specificity of subsequent analyses.
[0062] 3) Sequence Assembly (Step 2.3): Metagenomic assembly of non-host sequences was performed using MEGAHIT. The generated assembly result file is SRR8289540_contigs.fasta 4) Open Reading Frame Identification (Step 2.4): The Prodigal tool was used to predict open reading frames (ORFs) in the contigs file, outputting the protein sequence file SRR8289540_orfs.faa. Approximately 6119 complete gene ORFs with start and stop codons were identified.
[0063] 5) Drug Resistance Category and Mechanism Prediction (Step 3.3): The predicted protein sequences were input into a deep learning-based drug resistance prediction model. A total of 215 drug resistance genes were predicted, with the main drug resistance categories being aminoglycoside, tetracycline, fluoroquinolone, and carbapenem, and the main resistance mechanisms being antibiotic efflux and antibiotic inactivation. The predicted results are consistent with wet experimental results and provide more comprehensive drug resistance information.
[0064] This invention can be applied to the following scenarios, including but not limited to: 1) Rapid drug resistance testing for pathogens causing infectious diseases; 2) Diagnosis of bloodstream infection pathogens in patients in the intensive care unit (ICU); 3) Monitoring of pathogen resistance during public health emergencies; 4) Monitoring of environmental microbial resistance.
[0065] Furthermore, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described multi-platform antibiotic resistance gene end-to-end detection method. A computer-readable storage medium storing a computer program thereon is characterized in that, when executed by a processor, the computer program implements the above-described multi-platform antibiotic resistance gene end-to-end detection method. A computer program product includes a computer program, characterized in that, when executed by a processor, the computer program implements the above-described multi-platform antibiotic resistance gene end-to-end detection method.
[0066] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for end-to-end detection of antibiotic resistance genes applicable to multiple sequencing platforms, characterized in that, This includes a one-stop analysis process for multi-platform metagenomic data and a deep learning-based antibiotic resistance gene detection process. The one-stop analysis steps for multi-platform metagenomic data include: adaptively determining the sequencing platform based on the input sequencing data format and selecting the appropriate processing flow: for second-generation sequencing data, the sequence is preprocessed, the preprocessed sequence is aligned to the host reference genome to remove host contamination, and initial contigs are generated using a short read assembly tool; for third-generation sequencing data, after sequence preprocessing, the preprocessed sequence is self-aligned to generate an overlap correction file, then an overlap cluster backbone is constructed, and finally three iterations of correction are performed to generate initial contigs. After obtaining the initial contigs, species-level classification annotation was performed using the Kraken2 tool to provide species background information for antibiotic resistance gene detection. Then, the Prodigal tool was used to identify prokaryotic ORFs and filter out ORFs with a length <100aa. The translated protein sequences were then used in the antibiotic resistance gene detection step. The deep learning-based antibiotic resistance gene detection step includes: taking the protein sequence as input, fine-tuning a pre-trained protein language model, which is combined with a multi-scale residual neural network module to improve the model's ability to extract complex global and local high-dimensional features related to drug resistance; and using a bilinear classification head to achieve simultaneous prediction of antibiotic resistance categories and antibiotic resistance mechanisms.
2. The method according to claim 1, characterized in that, The one-stop analysis steps for multi-platform metagenomic data are implemented using the following method. A workflow engine is built, using Snakemake to construct a task dependency graph, and dynamically scheduling sequence preprocessing, sequence assembly, and ORF prediction processes based on the input data format. Containerized deployment, with each module packaged into a Docker image, supports cross-platform, dependency-free operation; Quality monitoring automatically generates quality reports at key steps.
3. The method according to claim 1, characterized in that, The fine-tuned pre-trained protein language model includes... The input layer takes one-hot encoded protein sequences as input. The feature extraction layer inserts LoRA modules into each Transformer layer of the ESM2 model. By adding low-rank update paths to query, key, and value projections, parameter fine-tuning is achieved, resulting in a fine-tuned ESM2 model. The embedded representation of the fine-tuned ESM2 model is then input into the multi-scale residual neural network module. The global average pooling and normalization layer performs global average pooling on the sequence output by the feature extraction layer along the sequence dimension, and then performs normalization through LayerNorm to output dimensionality-reduced sequence features. The classification layer uses a bilinear classification head to process the dimensionality-reduced features in parallel, and outputs the drug resistance category probability and the drug resistance mechanism probability, respectively.
4. The method according to claim 3, characterized in that, Training the fine-tuned pre-trained protein language model includes, A dataset was constructed by integrating and manually correcting five databases, including CARD, HMD-ARG, PLM-ARGDB, NCBI Bacterial Antimicrobial Resistance Reference Gene Database, and NCBI β-lactam Family Reference Gene Catalog, to form a model training set containing known antibiotic resistance gene sequences. A loss function is constructed. During the model training process, a Focal Loss function based on class weights and prediction confidence is used to optimize the task loss for predicting antibiotic resistance categories and resistance mechanisms, respectively, so as to improve the prediction accuracy of low-frequency categories. The formula for calculating Focal Loss is as follows:
5. Among them, γ represents the positive sample weight parameter for drug resistance category or mechanism, and γ is the focusing parameter. The total loss is calculated using the following formula:
6. During model training, the two tasks are jointly optimized through backpropagation to update the model parameters.
7. The method according to claim 1, characterized in that, After obtaining the contig sequences, the Kraken2 tool was used for species-level classification and annotation to provide species background information for antibiotic resistance gene detection. Then, the Prodigal tool was used to identify prokaryotic ORFs. After obtaining the corresponding ORFs, the translated protein sequences were used for antibiotic resistance gene detection.
8. The method according to claim 1, characterized in that, Sequence preprocessing includes removing low-quality sequences using dynamic parameter filtering tools. For second-generation sequencing data, the dynamic parameter filtering tools are Trimmomatic and / or FastP; for third-generation sequencing data, the filtering tools are cascaded filtering tools, using PoreChop to remove adapters and NanoFilt for quality filtering.
9. A multi-platform end-to-end detection system for antibiotic resistance genes, employing the method described in any one of claims 1-6 to simultaneously predict antibiotic resistance categories and mechanisms, characterized in that... It includes a one-stop analysis module for multi-platform metagenomic data and a deep learning-based antibiotic resistance gene detection module, among which... The multi-platform metagenomic data one-stop analysis module adaptively determines the sequencing platform and selects the appropriate processing flow based on the input sequencing data format: for second-generation sequencing data, the sequence is preprocessed, the preprocessed sequence is aligned to the host reference genome to remove host contamination, and initial contigs are generated using a short read assembly tool; for third-generation sequencing data, after sequence preprocessing, the preprocessed sequence is self-aligned to generate an overlap correction file, then an overlap cluster backbone is constructed, and finally three iterations of correction are performed to generate initial contigs. After obtaining the initial contigs, species-level classification and annotation were performed using the Kraken2 tool to provide species background information for antibiotic resistance gene detection. Then, the Prodigal tool was used to identify prokaryotic ORFs and filter out ORFs with a length <100 aa. The translated protein sequences were then input into the antibiotic resistance gene detection module. The deep learning-based antibiotic resistance gene detection module takes the protein sequence as input and fine-tunes a pre-trained protein language model. This model, combined with a multi-scale residual neural network module, improves the model's ability to extract complex global and local high-dimensional features related to resistance. Furthermore, a bilinear classification head is used to simultaneously predict antibiotic resistance categories and mechanisms.
10. A computer device, comprising: The memory and processor contain a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the multi-platform antibiotic resistance gene end-to-end detection method according to any one of claims 1-6.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the multi-platform antibiotic resistance gene end-to-end detection method according to any one of claims 1-6.
12. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the multi-platform antibiotic resistance gene end-to-end detection method according to any one of claims 1-6.