Genome analysis method for mining potential natural strain repository of drug-resistant gene tetX

By constructing a tetX multigenotyping database of drug resistance genes and comparing nucleic acid data, combined with analysis at the genomic sequence and sequencing read levels, the problems of incomplete drug resistance gene discovery and limited host association in existing technologies have been solved, enabling efficient mining of potential natural strain repositories.

CN121459944APending Publication Date: 2026-02-03CHINA PHARM UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511531125.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively discover uncharacterized drug resistance genes tetX and their hosts in the environment, and existing processes are complex and inefficient, failing to fully explore potential drug resistance gene repositories.

Method used

We constructed a comprehensive database of tetX drug resistance gene multigenotypes, combined with multiple public databases and literature, and screened and verified potential tetX drug resistance gene sequences through nucleic acid alignment and phylogenetic analysis. We also identified host information and gene locations by combining genomic sequence and sequencing read level analysis.

Benefits of technology

It expands the types and number of drug resistance gene databases, provides an integrated analysis strategy, and improves the efficiency and accuracy of mining potential natural strain repositories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459944A_ABST
    Figure CN121459944A_ABST
Patent Text Reader

Abstract

The invention discloses a genome analysis method for mining a potential natural strain repository of a drug-resistant gene tetX. According to the method, aiming at a database of a drug-resistant gene detection core, a plurality of representative databases (such as CARD, Resfinder and SARG) and drug-resistant gene sequences annotated in an NCBI public database are combined on the construction composition of the database; various types which are not included in but reported in the database at present are considered, so that the types and the quantity of corresponding drug-resistant gene databases (such as tetX genes) are greatly expanded; an optimized and integrated analysis strategy and a verification strategy are provided for mining and analyzing a potential natural strain repository of tetX drug-resistant genes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to bioinformatics analysis techniques for bacterial genomes, specifically to a genome analysis method for mining a potential natural strain reservoir of the drug resistance gene tetX. Background Technology

[0002] Antibiotic resistance (AMR) has become a major threat to global public health. The spread of resistance genes (ARGs) in pathogens makes many infections difficult or even untreatable. Understanding the origin, evolution, and transmission mechanisms of ARGs is crucial for developing new countermeasures. Extensive evidence suggests that microbial communities in natural environments (such as soil, water bodies, animal guts, and plant surfaces) are a vast and diverse reservoir of ARGs. Bacteria in these environments (including non-pathogenic bacteria) may carry “primitive” or “silent” ARGs not yet found in clinical pathogens, constituting a potential source for ARG evolution and transmission to pathogens. Tetracycline antibiotics are key drugs for treating Gram-negative bacterial infections, and the TetX resistance gene, through enzymatic modification that inactivates tetracycline, has become one of the main mechanisms of clinical resistance. Currently, TetX exhibits high genetic diversity, meaning that genotyping is associated with resistance, with different variants showing significant differences in resistance intensity to tetracyclines. For example, the resistance of tetX3 and tetX4 to tigecycline (MIC ≥ 16 μg / mL) is 4 times that of tetX2 (MIC ~ 4 μg / mL). Therefore, systematic identification and discovery of different tetX variants and potential host sources are of great scientific significance and application value.

[0003] Environmental resistance genome (ARG) research based on metagenomic sequencing and bioinformatics analysis is currently the closest existing technology. It mainly involves collecting samples from specific environments (such as soil, sewage, and farms); extracting total environmental DNA; performing metagenomic shotgun sequencing; identifying and annotating ARGs using bioinformatics workflows (such as read-based or contig / MAG-based methods) (mainly relying on comparisons with existing databases); and analyzing the abundance, diversity, distribution patterns of ARGs and their relationship with environmental factors.

[0004] Preliminary taxonomic annotation of microorganisms carrying ARGs (e.g., through 16S rRNA genes or genome-wide marker genes).

[0005] Limited Discovery of Drug Resistance Genes: Current technologies rely on comparing against known ARG databases. However, they cannot effectively discover novel / uncharacterized ARGs, and many drug resistance genes not included in databases, with low sequence similarity, or possessing novel structures are missed. The bacterial microbiome is a huge potential source of novel ARGs, but current methods have limited ability to discover them.

[0006] Limitations of Microbial Host Classification: Annotation methods for detecting microbial classification are based on assembling into contigs. However, a single contig may contain gene fragments from different microorganisms (chimeric assembly), or there may be taxonomic annotation errors, leading to incorrect ARG host associations. Constructing high-quality, complete MAGs is very difficult in complex environmental samples, especially for low-abundance or genomically complex microorganisms. Many microorganisms carrying ARGs may not be successfully binned or have high-quality MAGs obtained, resulting in missing host information.

[0007] Ignoring potential drug resistance genes: Existing processes tend to identify known, high-abundance ARGs and their hosts, while under-exploring the large number of "silent" ARGs in the environment that have not yet been activated or expressed (i.e., the real "potential" reservoir).

[0008] Complex and inefficient processes: Existing processes typically involve the interconnection of multiple independent software tools, resulting in complex processes, high computational resource consumption, and insufficient and inefficient information integration between different steps. There is a lack of a unified analysis strategy specifically optimized for the goal of "mining potential natural strain reservoirs". Summary of the Invention

[0009] Purpose of the invention: The purpose of this invention is to provide a genomic bioinformatics analysis method for mining potential natural strains containing the drug resistance gene tetX, in order to solve the problems of incomplete drug resistance gene mining and limited association with host microorganisms in existing technologies.

[0010] Technical solution: The method for establishing a potential natural strain bank for mining the drug resistance gene tetX, as described in this invention, includes the following steps: (1) Construction of a complete database of tetX drug resistance gene polytyping: Search and / or collect tetX gene sequences from existing public drug resistance databases, novel tetX variant gene sequences from literature, and tetX coding sequences annotated by NCBI; The genotyping information of the tetX coding sequence annotated by NCBI was identified and verified, and the identified and verified tetX gene sequence was obtained. We collected tetX gene sequences from existing public drug resistance databases, novel tetX variant gene sequences from literature, and identified and verified tetX gene sequences, integrated and deduplicated them, and constructed a complete database of tetX multi-type drug resistance genes. (2) Screening of potential strains with tetX resistance gene: Search for potential tetX resistance gene sequences: Using the complete database of tetX multi-type resistance genes constructed in step (1) as the index database and the NCBI NT database as the query database, perform nucleic acid alignment; further screening is then performed, with the following screening conditions: alignment length ≥ minimum sequence length * 80%; nucleic acid identity ≥ minimum sequence consistency; coverage ≥ 90%; Extract gene sequences from the alignment region and identify host information, location information, and gene type of the gene sequences; Statistical calculation of key indicators: The chromosomal localization rate of bacterial resistance genes and the proportion of chromosomal localized resistance gene variants in bacteria are statistically analyzed according to different genera.

[0011] The method for establishing a potential natural strain bank for mining the drug resistance gene tetX includes the following steps in step (1): (1.1) Collect gene sequences from existing public drug resistance databases; the existing public drug resistance databases include CARD, Resifinder, and SARG; (1.2) Search the literature to obtain novel variant gene sequences; (1.3) Collect tetX encoded sequences annotated by NCBI; (1.4) Identify and verify the typing information of the tetX encoded sequences annotated by NCBI; (1.5) Construct a complete database of tetX gene polymorphisms: Collect the sequences obtained in steps (1.1), (1.2), and (1.4) and remove sequences with the same Accession ID, Start site, and End site information; (1.6) Perform multiple sequence alignment of the nucleic acid sequences of the gene set using MUSCLE; generate a phylogenetic likelihood tree using IQ-TREE.

[0012] The method for establishing a potential natural strain bank for mining the drug-resistant gene tetX, in step (2), the host information is the bacterial classification information corresponding to the bacterial species corresponding to the gene sequence; the location information is divided into chromosome and plasmid; gene type: gene sequences with >90% consistency with the gene set sequence are identified as gene types that are the same as the corresponding alignment sequence; while gene sequences with <90% consistency with the gene set sequence are identified as potential gene types that are the same as the corresponding alignment sequence.

[0013] The method for establishing a potential natural strain bank for mining the drug resistance gene tetX, wherein the minimum sequence length in step (2) is 1060~1080bp and the minimum sequence consistency is 80%~90%. Preferably, the minimum sequence length is 1077bp and the minimum sequence consistency is 80.9%.

[0014] The method for establishing a potential natural strain repository for the drug resistance gene tetX, wherein step (2) involves nucleic acid alignment using the NCBI BLASTn tool.

[0015] The method for establishing a potential natural strain repository for the drug resistance gene tetX, wherein the alignment parameter for nucleic acid alignment in step (2) is set to -outfmt 6 output format.

[0016] The method for establishing a potential natural strain bank for mining the drug-resistant gene tetX, step (2) of which involves extracting the gene sequence of the comparison region, specifically involves downloading the FASTA and summary files of these sequences in batches using the BatchEntrez tool on the NCBI Nucleotide page based on the sequence ID of the screening results. Then, based on the sequence ID, alignment start point, and alignment end point information in the alignment result file, Bedtools is used to extract the sequence of the alignment region; MAFFT is used to adjust the orientation uniformly.

[0017] The potential natural strain reservoir for the drug resistance gene tetX was obtained using the method described above.

[0018] For the verification method of potential natural strains with the drug resistance gene tetX, choose one or both of (a) and (b): (a) BLAST protocol based on the level of assembled contigs of genome sequence Step 1: Download the target strain genome: Download all assembled contig-level genome sequence data of the target strain in batches; Step 2: Identify drug resistance genes in the genomic sequence: Using the complete database of the drug resistance gene tetX multi-genotyping constructed in step (1) as the index database, and the genomic sequence of the target strain as the query database, perform nucleic acid alignment; Screening conditions: Alignment length ≥ minimum sequence length; Nucleic acid identity ≥ 90%; Coverage ≥ 90%; (b) Mapping scheme based on WGS sequencing reads Step 1: Download target strain sequencing data: Download all sequencing library sequence data of the target strain in batches; Step 2: Construction of the index database: Using bowtie2, construct the index database for subsequent bowtie2 alignments based on the complete database of tetX multi-genotyping of the drug resistance gene constructed in step (1); Step 3: Alignment of raw sequencing data; Step 4: Alignment file conversion and processing: Use samtools to convert the alignment result SAM file to BAM format and sort it by genomic coordinates. Then use samtools to convert the sorted BAM file into a readable text format. Finally, use bedtools to calculate the sequencing coverage depth of each genomic location. Step 5: Calculate mapping parameters Based on the sequencing coverage depth result file obtained in step 4, calculate the following parameters: a. Average depth = Sum of depths at all locations of the gene / Length of the gene (number of locations) b. Coverage ratio = Number of locations with depth greater than 0 / Total number of locations for this gene c. Maximum coverage gap: Calculate the maximum length of a continuous gap at a depth of 0 on this gene. d. Continuous Coverage: Calculate the maximum length of consecutive non-zero values ​​on the gene. e. Continuous coverage: Calculated as the maximum length of consecutive non-zero values ​​on the gene divided by the length of the gene (number of positions). Step 6: Filter mapping results Based on the calculation parameters in step 5, set the following thresholds: average depth threshold: >= 5x; coverage ratio threshold: >= 80%; maximum gap threshold: <= 100bp. Sequencing data that meets the above thresholds contain the target drug resistance gene tetX.

[0019] The method for verifying potential natural strains of the drug resistance gene tetX, wherein the sequence length is at least 1060~1080 bp, preferably 1077 bp.

[0020] Preferably, the following method is used:

[0021] Option 1: Construction of a complete database of tetX resistance gene polygenotyping Step 1: Review and collect gene sequences from existing public drug resistance databases, including: CARD database (Comprehensive Antibiotic Resistance Database, https: / / card.mcmaster.ca); The Resifinder database (https: / / cge.cbs.dtu.dk / services / ResFinder); SARG database (Structured Antibiotic Resistance Genes, https: / / smile.hku.hk / SARGs); Step 2: Obtain novel variant gene sequences that are missing from existing public drug resistance databases through recent literature searches. Step 3: Collect tetX encoded sequences annotated by NCBI Search for relevant sequences in the NCBI (National Center for Biotechnology Information, https: / / www.ncbi.nlm.nih.gov) Nucleotide database using "selected drug resistance gene" as the search term, download all Coding Sequences, and collect sequences whose names contain both [gene=drug resistance gene name] and [protein=drug resistance gene protein name].

[0022] Step 4: Identify and verify the genotyping information of the tetZ coding sequence annotated by NCBI. To filter out incorrectly annotated sequences and avoid false positives, the following steps will be used for identification and verification: First, the gene sequences obtained in steps 1 and 2 are integrated into a tetX resistance gene multigenotype set. The sequences obtained from NCBI were compared and identified using BLASTn with the aforementioned multi-type gene set of drug resistance genes. The aligned sequences were then adjusted using the –adjustdirection parameter of MAFFT to uniformly adjust the orientation of the sequences obtained from NCBI, and then verified by multiple sequence alignment with the above tetX resistance gene multi-type gene set. Step 5: Construct a complete database of tetX resistance gene polygenotyping. The sequences obtained in steps 1, 2, and 4 are collected, and sequences with the same Accession ID, Start site, and End site information are removed to obtain a complete database of tetX resistance gene polytyping.

[0023] Option 2: Screening of a potential strain bank containing tetX resistance genes Step 1: Confirming the filtering criteria Gene similarity calculation: Using the NCBI BLASTn tool, with the gene set itself as the query and alignment database (indexed using makeblastdb), nucleic acid (blastn) alignment is performed. The alignment parameters are set to -outfmt 6 output format (including fields such as qseqid / qlen / pident / length / evalue), and -num_alignments 188 to ensure full coverage. Gene length calculation: Use TBtools-II to calculate sequence information, including minimum length / maximum length / average length / median length; The screening criteria were determined based on the minimum gene similarity and minimum gene length. The subsequent BLAST screening criteria were: alignment length ≥ 80% of the minimum sequence length; nucleic acid identity ≥ the minimum sequence consistency; and coverage ≥ 90%.

[0024] Step 2: Search for potential tetX resistance gene sequences Using the NCBI BLASTn tool, the complete tetX resistance gene multigenotyping database constructed in Scheme 1 was used as the index database (makeblastdb was used to create the index), and the NCBI NT database (Nucleotide Sequence Database) was used as the query database to perform nucleic acid (blastn) alignment. The alignment parameters were set to -outfmt 6 output format (including qseqid / qlen / qstart / qend / sseqid / stitle / pident / length / evalue / staxid fields), and -num_alignments 5, which means that the maximum number of sequences that are successfully aligned with the complete resistance gene multigenotyping database will be output. Then, the result files are filtered again based on the filtering criteria confirmed in step 1.

[0025] Step 3: Extract gene sequences from the alignment region Based on the sequence IDs of the filtered results, i.e. the ACCESSION IDs of the sequences, use the Batch Entrez tool on the NCBI Nucleotide page to download the FASTA and summary files of these sequences in batches. Then, based on the first, third, and fourth columns of the alignment result file, namely the sequence ID, alignment start point, and alignment end point information, Bedtools is used to extract the sequence of the alignment region; the MAFFT's –adjustdirection parameter is used to uniformly adjust the direction of the sequence.

[0026] Step 4: Identify host information from gene sequences Obtain the bacterial species-level host information corresponding to each gene sequence from the summary file downloaded in step 3; Use the TaxonKit tool to identify bacterial species and their corresponding bacterial classification information, such as phylum, class, order, family, and genus.

[0027] Step 5: Identify the location information of the gene sequence First, based on the summary file downloaded in step 3, we can obtain preliminary information about this sequence, including: Chromosome, complete genome / plasmid, complete sequence / complete cds / genomic sequence; For the above sequence information, this step identifies the gene location of sequences with "Chromosome, complete genome" information as chromosomes; and the gene location of sequences with "plasmid, complete sequence" information as plasmids. For sequences with "genomic sequence" information, this step uses PlasFlow to identify the presence of chromosomes / plasmids in these sequences.

[0028] Step 6: Identification of the genotype of potential tetX resistance genes Based on the alignment result file obtained in step 3, the gene type of the gene sequence obtained in step 3 is identified according to the fifth and seventh columns of the alignment result file, namely the gene variant sequence name of the gene set and the alignment consistency. The specific identification conditions are as follows: gene sequences with >90% consistency with the gene set sequence are identified as the same gene type as the corresponding aligned sequence (such as the tetX3 gene); while gene sequences with <90% consistency with the gene set sequence are identified as potential gene types of the same gene type as the corresponding aligned sequence (such as the potential tetX3 gene).

[0029] Step 7: Statistical Calculation of Key Indicators In our analytical workflow, two key indicators are used to screen for the existence of potential natural hosts: the chromosomal localization rate of bacterial resistance genes and the proportion of chromosomally located resistance gene variants. Therefore, this step requires combining the bacterial genera and chromosomal localization sequence information from steps 3 and 4 to statistically analyze these two indicators for different bacterial genera.

[0030] Option 3: Validation of a potential strain bank for tetX resistance genes This scheme uses two schemes based on the difference in the assembly level of the sequences: a BLAST scheme based on the level of assembled contigs of the genome sequence; and a mapping scheme based on the level of WGS sequencing reads.

[0031] 1. BLAST scheme based on pre-assembled contig / scaffold genome sequences Step 1: Download the genome of the target strain Use datasets to batch download all assembled contig-level genome sequence data of the target strain.

[0032] Step 2: Identify drug resistance genes in the genomic sequence Using the NCBI BLASTn tool, the complete tetX multigenotyping database constructed in Scheme 1 was used as the index database (indexed using makeblastdb), and the genomic sequence of the target strain was used as the query database for nucleic acid (blastn) alignment. The alignment parameters were set to -outfmt 6 output format (including qseqid / qlen / qstart / qend / sseqid / stitle / pident / length / evalue / staxid fields), and -num_alignments 5, meaning a maximum of 5 sequences successfully aligned with the complete tetX multigenotyping database were output. The selection criteria were: alignment length ≥ minimum sequence length; nucleic acid identity ≥ 90%; coverage ≥ 99%.

[0033] 2. Mapping scheme based on WGS sequencing reads Step 1: Download the sequencing data of the target strain Use sra-tools to download all sequence data of the target strain in batches.

[0034] Step 2: Building the Index Database The bowtie2-build command of bowtie2 is used to build an index database for subsequent bowtie2 alignments from a complete database of drug resistance gene multigenotypes.

[0035] Step 3: Alignment of raw sequencing data Depending on the sequencing library and sequencing method, single-end sequencing uses the "-U" input file parameter command in bowtie2 v2.5.4; paired-end sequencing uses the "-1" and "-2" input file parameter commands in bowtie2 v2.5.4.

[0036] Step 4: Compare and convert the files Use samtools to convert the alignment results SAM file to BAM format and sort it by genomic coordinates. Then use samtools to convert the sorted BAM file into a readable text format. Finally, use bedtools to calculate the sequencing coverage depth (i.e., how many sequencing fragments cover each base) for each genomic location.

[0037] Step 5: Calculate mapping parameters Based on the final result file obtained in step 4, which is the sequencing coverage depth result file for each genomic location, the following parameters are calculated: a. Average depth = Sum of depths at all locations of the gene / Length of the gene (number of locations) b. Coverage ratio = Number of locations with depth greater than 0 / Total number of locations for this gene c. Maximum coverage gap: Calculate the maximum length of a continuous gap at a depth of 0 on this gene. d. Continuous Coverage: Calculate the maximum length of consecutive non-zero values ​​on the gene. e. Continuous coverage: Calculate the maximum length of consecutive non-zero values ​​on the gene / the length of the gene (number of positions).

[0038] Step 6: Filter mapping results Based on the calculation parameters in step 5, set the following thresholds: average depth threshold: >= 5x; coverage ratio threshold: >= 80%; maximum gap threshold: <= 100bp. Sequencing data that meets these thresholds contains the target drug resistance gene.

[0039] Beneficial effects: Compared with the prior art, the present invention has the following advantages: For the core database of drug resistance gene detection, the present invention combines multiple representative databases (such as CARD, Resfinder, SARG) and annotated drug resistance gene sequences from the NCBI public database in its construction composition, and takes into account various subtypes that are not included in the current database but have been reported, which greatly expands the types and number of corresponding drug resistance gene databases (such as tetX gene); it provides an optimized and integrated analysis strategy and validation strategy for mining and analyzing the "potential natural strain reservoir" of tetX drug resistance gene. Attached Figure Description

[0040] Figure 1 This is a technical roadmap of the analysis process of this invention; Figure 2 Phylogenetic maximum likelihood tree for the tetX gene multigenotyping database; Figure 3 Nucleic acid and amino acid similarity matrix among different tetX subtypes in the screening of potential strains for the tetX gene family; Figure 4 The gene set sequence length range for screening potential strains of the tetX gene family; Figure 5 The species and genera of different bacterial families in the screening of potential strains for the tetX gene family; Figure 6 A diagram showing the distribution of gene variants in the screening of potential strains for the tetX gene family; Figure 7 Chromosomal localization rate of the tetX gene in different fungal families during the screening of potential strains for the tetX gene family reservoir; Figure 8Chromosomal mapping of the top 8 families with the most tetX gene counts in the screening of potential strains for the tetX gene family; Figure 9 The percentage of chromosome / plasmid-localized genes in the tetX gene identified at the assembly level of Weeksellaceae genome contigs; Figure 10 The number of genome sequences of different genera in the family Weeksellaceae and the percentage of tetX-positive bacteria were identified at the assembly level of Weeksellaceae genome contigs. Figure 11 The types and proportions of Weeksellaceae gene variants in the tetX gene identified at the assembly level of Weeksellaceae genome contigs; Figure 12 The percentage of multi-copy genomes in the tetX gene of the Weeksellaceae family identified at the assembly level of Weeksellaceae genome contigs; Figure 13 Co-location map of Weeksellaceae family multicopy genome gene variants in the tetX gene identified at the assembly level of Weeksellaceae family genome contigs; Figure 14 Coverage depth of the tetX gene in the library identified at the Weeksellaceae genomic read level; Figure 15 The percentage of continuous coverage of the tetX gene in the library identified at the genomic read level for the Weeksellaceae family; Figure 16 The number of organisms containing tetX gene libraries identified at the genomic read level in the Weeksellaceae family; Figure 17 The percentage of tetX genes identified at the genomic read level in different library classifications of the Weeksellaceae family. Detailed Implementation

[0041] The technical solution of the present invention will be further described below with reference to the accompanying drawings. A roadmap of the present invention is shown below. Figure 1 .

[0042] Example 1 Construction of a complete database of tetX gene polymorphisms Step 1: Review and collect gene sequences from existing public drug resistance databases, including: The CARD (Comprehensive Antibiotic Resistance Database, https: / / card.mcmaster.ca) database contains the sequences of six genes: tetX / tetX1 / tetX3 / tetX4 / tetX5 / tetX6. The Resifinder database (https: / / cge.cbs.dtu.dk / services / ResFinder) contains 7 sequences for tetX / tetX3 / tetX4 / tetX5 / tetX6; The SARG (Structured Antibiotic Resistance Genes, https: / / smile.hku.hk / SARGs) database includes 46 gene sequences of tetX / tetX1 / tetX2 / tetX3 / tetX4 / tetX5 / tetX6 gene variants.

[0043] Step 2: Obtain novel variant gene sequences through recent literature searches (a detailed list of literature is attached at the end of the instructions). In 2020, Gasparrini et al. [1] identified and detected new variants of tetX7 ~ tetX13. In 2020, Cheng et al. [2] discovered a new variant of the tetX family, tetX14. In 2020, Liu et al. [3] detected four subtypes of tet (X6), and in 2021, Xu et al. [4] renamed subtypes II, III, and IV as tet (X15) ~ tet (X17). In 2021, Umar et al. [5] identified 26 new variants tet(X18) ~tet(X44) in the duck plague bacterium. In 2021, Zhang et al. [6] identified three new tet(X) genes: tet(X45), tet(X46), and tet(X47). In 2025, Yao et al. [7] identified the tetX48-tetX57 gene variant sequence. Step 3: Collect tetX encoded sequences annotated by NCBI Using "tetX" as the search term, a search was conducted in the NCBI (National Center for Biotechnology Information, https: / / www.ncbi.nlm.nih.gov) Nucleotide database, yielding 771 entries. All Coding Sequences were downloaded, and sequences containing "[gene=tetX]" and "[protein=tetracycline resistance protein / Flavin-dependentmonooxygenase]" were collected, resulting in 267 sequences.

[0044] Step 4: Identify and verify the genotyping information of the tetX encoded sequences annotated by NCBI. To filter out incorrectly annotated sequences and avoid false positives, the following steps will be used for identification and verification: First, the gene sequences obtained in steps 1 and 2 are integrated into a tetX multigenotype gene set; The 267 sequences obtained from NCBI were compared and identified with the above-mentioned tetX multi-type gene set using BLASTn v2.15.0; The aligned sequences were then aligned with the 267 sequences obtained from NCBI using the –adjustdirection parameter of MAFFT v7.526, and then verified by multiple sequence alignment with the above tetX multi-type gene set. Ultimately, after identifying and verifying the gene sequences, 96 tetX gene sequences were obtained, and the corresponding genotype names of the tetX genotypes with the highest consistency values ​​compared with BLASTn were matched.

[0045] Step 5: Construct a complete database of tetX genotyping The sequences obtained in steps 1, 2, and 4 were collected, and sequences with the same Accession ID / Start site / End site information were removed. The final tetX gene database includes 56 gene variants and 188 gene sequences.

[0046] Step 6: Perform multiple sequence alignment on the nucleic acid sequences of the gene set using MUSCLE v5.1.linux64; generate a phylogenetic maximum likelihood tree using IQ-TREE 2.3.6 (parameters: -m MFP -T 20 -B 1000).

[0047] Phylogenetic maximum likelihood tree of the tetX genotyping database, such as Figure 2 As shown.

[0048] Example 2 Screening of potential strains from the tetX gene family 1. Preliminary screening of potential strain banks Step 1: Confirming the filtering criteria Gene similarity calculation: Using the NCBI BLASTn v2.15.0 tool, with the gene set itself as the query and alignment database (indexed using makeblastdb), nucleic acid (blastn) alignment was performed. The alignment parameters were set to -outfmt 6 output format (including fields such as qseqid / qlen / pident / length / evalue), and -num_alignments 188 to ensure full coverage. Gene length calculation: Use TBtools-Ⅱ v2.310 to calculate sequence information, including minimum length / maximum length / average length / median length.

[0049] Using the above methods, the similarity among the tetX gene polymorphism gene sets can be calculated to be greater than 80%, with a minimum of 80.9% and a maximum of 100%. The similarity matrix among the tetX gene polymorphism gene sets is as follows: Figure 3 As shown. The gene set has a minimum sequence length of 1077 bp, a maximum of 1167 bp, an average of 1158 bp, and a median of 1167 bp. The gene set sequence length range is as follows. Figure 4 As shown. Therefore, based on the minimum gene similarity and minimum gene length, the screening criteria were confirmed, namely, the subsequent BLAST screening criteria are: alignment length ≥ 861 (minimum sequence length * 80%); nucleic acid identity ≥ 80.9% (minimum sequence consistency); coverage ≥ 90%.

[0050] Steps 2-3: Search for and obtain potential drug resistance gene sequences Using the NCBI BLASTn v2.15.0 tool, with the complete tetX multigenotyping database constructed in Scheme 1 as the index database (indexed using makeblastdb), and the NCBI NT database (Nucleotide SequenceDatabase) as the query database, nucleic acid (blastn) alignment was performed. The alignment parameters were set to -outfmt 6 output format (including qseqid / qlen / qstart / qend / sseqid / stitle / pident / length / evalue / staxid fields), and -num_alignments 5, meaning a maximum of 5 sequences successfully aligned with the complete tetX multigenotyping database were output. The result files were then further filtered according to the screening criteria confirmed in Step 1. Based on the sequence IDs (ACCESSION IDs) of the filtered results, the FASTA and summary files of these sequences were downloaded in batches using the Batch Entrez tool on the NCBI Nucleotide page. Then, based on the first, third, and fourth columns of the alignment results file—namely, sequence ID, alignment start point, and alignment end point—the alignment region sequences were extracted using Bedtools v2.31.1; the `--adjustdirection` parameter of MAFFT v7.526 was used to uniformly adjust the sequence orientation. Using the above search method, 692 tetX-like genes, distributed across 574 bacterial strains, were retrieved from the Nt database.

[0051] Steps 4-5: Identify the location of the gene sequence and host information First, based on the summary file downloaded in step 3, we can obtain preliminary information about this sequence, including: Chromosome, complete genome / plasmid, complete sequence / complete cds / genomic sequence; For the above sequence information, this step identifies the gene location of sequences with "Chromosome, complete genome" information as chromosomes; sequences with "plasmid, complete sequence" information are identified as plasmids; and for sequences with "genomic sequence" information, PlasFlow v.1.1 is used to determine the presence of chromosomes / plasmids. Similarly, host information at the bacterial species level is obtained from the summary file for each gene sequence; TaxonKit v0.17.0 is used to identify the bacterial classification information, such as phylum, class, order, family, and genus, corresponding to the bacterial species.

[0052] Strains carrying the gene sequence are distributed across 33 genera and 16 families. Strains carrying the tetX-like gene are also distributed across 33 genera and 16 families. Enterobacteriaceae is primarily spread via plasmids: 31.5% of tetX genes are found in this family, mainly in the genera *Escherichia* (187 strains) and *Klebsiella* (27 strains). Among these, tetX4 accounts for 97.4% of the genes in this family, and 95.7% are located on plasmids. The family Weeksellaceae is predominantly chromosomally transmitted, with 19.5% of the tetX gene residing within this family. The genus *Riemerella* (71.9%) is the primary vector, carrying various variants including tetX10 (30 strains), tetX18 (13 strains), tetX19 (13 strains), tetX20 (11 strains), and tetX45.3 (26 strains). Furthermore, 89.6% of these gene variants are located on chromosomes. The genera and gene locations within different bacterial families are shown below. Figure 5 As shown.

[0053] Step 6: Identification of the genotype of the potential drug resistance gene tetX Based on the alignment results file obtained in step 3, the gene types of the gene sequences obtained in step 3 were identified according to the fifth and seventh columns of the alignment results file, namely the gene variant sequence names of the gene set and the alignment consistency. Specific identification criteria are as follows: gene sequences with >90% consistency with the gene set sequence were identified as gene types identical to their corresponding alignment sequences; while gene sequences with <90% consistency with the gene set sequence were identified as potential gene types identical to their corresponding alignment sequences. Ultimately, approximately 97.1% (670 / 690) of these genes were identified as having greater than 90% consistency with genes in the tetX gene database, and only 2.90% (20 / 692) of tetX-like genes had 80-90% consistency with the database. Of the 670 tetX genes, 33 variants were identified, with tetX4 (34.8%), tetX10 (23.4%), and tetX3 (16.0%) being the dominant variants. The vast majority of variants, namely tetX3, tetX4, and tetX5 (13 variants in total, 85.5%), exhibited chromosome-plasmid coexistence. However, among the three dominant variants, tetX4 accounted for 94.6% of the plasmid, demonstrating significant "plasmid adaptability," while tetX10 accounted for 75.9% of the chromosome but only 7.1% of the plasmid, showing "chromosome bias." The distribution of gene variants is shown in the figure below. Figure 6 As shown.

[0054] Step 7: Statistical Calculation of Key Indicators In our analytical workflow, two key indicators were used to screen for the existence of potential natural hosts: the chromosomal localization rate of the bacterial tetX gene and the proportion of chromosomal localized drug-resistant gene variants and their numbers. Based on these two key indicators, the Weeksellaceae family was identified as exhibiting the highest number of gene chromosomal localizations (135), a high gene chromosomal localization rate (89.6%), and the most chromosomal localized gene variant types (14) and numbers (121), indicating that this method can screen for potential natural hosts of drug-resistant genes (such as tetX). The chromosomal localization rate of the tetX gene in different bacterial families is shown below. Figure 7 As shown; the gene types of the top 8 fungal families with the most tetX genes located on chromosomes are as follows: Figure 8 As shown.

[0055] Example 3 Validation of a potential strain bank for drug resistance genes This approach uses two schemes based on the different levels of sequence assembly: a BLAST scheme based on the assembled contig level of the genome sequence; and a mapping scheme based on the WGS sequencing read level. 1. BLAST scheme based on pre-assembled contig / scaffold genome sequences Using the datasets tool v16.35.0, all assembled contig-level genome sequences of the target strain were downloaded in batches, totaling 2793 genomes. Using the NCBI BLASTn v2.15.0 tool, with the complete tetX multi-genotyping database constructed in Example 1 as the index database (indexed using makeblastdb), and the genome sequences of the target strain as the query database, nucleic acid (BLASTn) alignment was performed. The alignment parameters were set to -outfmt 6 output format (including qseqid / qlen / qstart / qend / sseqid / stitle / pident / length / evalue / staxid fields), and -num_alignments 5, meaning a maximum of 5 sequences successfully aligned with the complete tetX multi-genotyping database were output. The selection criteria were: alignment length ≥ 1077; nucleic acid identity ≥ 90%; coverage ≥ 90%.

[0056] Based on the above genome mining scheme, 719 gene sequences were mined from 2793 strain genomes, distributed across 544 genomes. Plasflow was used to distinguish between plasmid and chromosomal sequences, revealing that chromosomal genes accounted for 80.3%, while plasmids accounted for only 1.3%. This data also corroborates the reliability of the screening analysis process in Part II. The proportion of chromosomal / plasmid-localized genes is as follows: Figure 9 As shown. Of the 24 genera in this family, 12 genera (50%) tested positive for tetX, including: *Riemerella* (positive rate 92%, 362 / 392); *Weeksella* (positive rate 50%, 4 / 8), etc. The number of genome sequences from different genera within the *Weeksellaceae* family and the percentage of tetX-positive bacteria are shown in the figure. Figure 10 As shown. Among the gene variant types, the major gene variant types present in the genomes of Weeksellaceae bacteria are tetX10 (192, 26.7%), tetX45.3 (125, 17.4%), and tetX19 (98, 13.6%). (See table for Weeksellaceae gene variant types and their percentages.) Figure 11As shown. Furthermore, this invention discovered multiple copies of the tetX gene (i.e., multiple copies of the tetX gene detected in the genome of the same strain) in several genera within the Weeksellaceae family, mainly concentrated in the genera: Bergeyella, Riemerella, Empedobacter, Chryseobacterium, and Elizabethkingia. Statistical analysis of specific genotypes of these multiple-copy genomes revealed frequent co-localization between tetX19 and tetX45.3 (20%, 28 / 140). The percentage of multiple-copy genomes in the Weeksellaceae family is as follows: Figure 12 As shown. The co-location map of multicopy genome variants in the Weeksellaceae family is shown below. Figure 13 As shown.

[0057] 2. Mapping scheme based on WGS sequencing reads Step 1: Download the sequencing data of the target strain Use sra-tools v3.0.0 to batch download all Illumina sequencing library sequence data of the target strain, totaling 1644 SRA libraries.

[0058] Step 2: Building the Index Database The bowtie2-build command of bowtie2 v2.5.4 was used to construct an index database for subsequent bowtie2 alignments from the complete database of tetX multigenotyping of the drug resistance gene obtained in Example 1. Step 3: Alignment of raw sequencing data Depending on the sequencing library and sequencing method, single-end sequencing uses the "-U" input file parameter command in bowtie2 v2.5.4; paired-end sequencing uses the "-1" and "-2" input file parameter commands in bowtie2 v2.5.4.

[0059] Step 4: Compare and convert the files The alignment results SAM files were converted to BAM format using samtools v1.20 and sorted by genomic coordinates. Then, samtools was used to convert the sorted BAM files into a readable text format. Finally, bedtools v2.31.1 was used to calculate the sequencing coverage depth (i.e., how many sequencing fragments cover each base) for each genomic location.

[0060] Step 5: Calculate mapping parameters Based on the final result file obtained in step 4, which is the sequencing coverage depth result file for each genomic location, the following parameters are calculated: a. Average depth = Sum of depths at all locations of the gene / Length of the gene (number of locations) b. Coverage ratio = Number of locations with depth greater than 0 / Total number of locations for this gene c. Maximum coverage gap: Calculate the maximum length of a continuous gap at a depth of 0 on this gene. d. Continuous Coverage: Calculate the maximum length of consecutive non-zero values ​​on the gene. e. Continuous coverage: Calculated as the maximum length of consecutive non-zero values ​​on the gene divided by the length of the gene (number of positions). Step 6: Filter mapping results Based on the calculation parameters in step 5, set the following thresholds: average depth threshold: >= 5x; coverage ratio threshold: >= 80%; maximum gap threshold: <= 100bp. Sequencing data that meets these thresholds contains the target drug resistance gene.

[0061] Following the steps above, the tetX gene was detected in 52 out of 1644 SRA libraries. The coverage depth of the tetX gene in the libraries is as follows: Figure 14 As shown; the continuous coverage ratio of the tetX gene in the library is, for example... Figure 15 As shown in the figure. Of these, 20 libraries originated from the human gut, 16 from *Elizabethkingia* bacteria, 8 from *Riemerella* bacteria, 3 libraries each from *Chryseobacterium* and *Weeksella*, and only one library each from *Ornithobacterium* and *Cloacibacterium*. The number of biological taxa containing the tetX gene library is shown in the figure. Figure 16 As shown. The percentage of the tetX gene in different library classifications of the Weeksellaceae family is as follows: Figure 17 As shown.

[0062] This application retrieved a list of literature containing novel variant gene sequences: 1. Gasparrini, AJ, et al., Tetracycline-inactivating enzymes from environmental, human commensal, and pathogenic bacteria cause broad-spectrum tetracycline resistance. Communications Biology, 2020.3(1): p. 241. 2. Cheng, Y., et al., Identification of novel tetracycline resistance gene tet(X14) and its co-occurrence with tet(X2) in a tigecycline-resistant and colistin-resistant Empedobacter stercoris. Emerg Microbes Infect, 2020.9(1): p. 1843-1852. 3. Liu, D., et al., Identification of the novel tigecycline resistance gene tet(X6) and its variants in Myroides, Acinetobacter and Proteus of food Animal origin. J Antimicrob Chemother, 2020.75(6): p. 1428-1431. 4. Xu, Y., et al., Co-production of Tet(X) and MCR-1, two resistance Enzymes are produced by a single plasmid. Environ Microbiol, 2021.23(12): p. 7445-7464. 5. Umar, Z., et al., The poultry pathogen Riemerella anatipestifer appears as a reservoir for Tet(X) tigecycline resistance. Environ Microbiol,2021.23(12): p. 7465-7482. 6. Zhang, R.M., et al., Source Tracking and Global Distribution of the Tigecycline Non-Susceptible tet(X). Microbiol Spectr, 2021.9(3): p. e0116421. 7.Yao, C., et al., Unraveling the evolution and global transmission of high level tigecycline resistance gene tet(X). Environment International,2025.199: p. 109499。

Claims

1. A method for establishing a potential natural strain bank for mining the drug resistance gene tetX, characterized in that, Includes the following steps: (1) Construction of a complete database of tetX drug resistance gene polytyping: Search and / or collect tetX gene sequences from existing public drug resistance databases, novel tetX variant gene sequences from literature, and tetX coding sequences annotated by NCBI; The genotyping information of the tetX coding sequence annotated by NCBI was identified and verified, and the identified and verified tetX gene sequence was obtained. We collected tetX gene sequences from existing public drug resistance databases, novel tetX variant gene sequences from literature, and identified and verified tetX gene sequences, integrated and deduplicated them, and constructed a complete database of tetX multi-type drug resistance genes. (2) Screening of potential strains with tetX resistance gene: Search for potential tetX resistance gene sequences: Using the complete database of tetX multi-type resistance genes constructed in step (1) as the index database and the NCBI NT database as the query database, perform nucleic acid alignment; further screening is then performed, with the following screening conditions: alignment length ≥ minimum sequence length * 80%; nucleic acid identity ≥ minimum sequence consistency; coverage ≥ 90%; Extract gene sequences from the alignment region and identify host information, location information, and gene type of the gene sequences; Statistical calculation of key indicators: The chromosomal localization rate of bacterial resistance genes and the proportion of chromosomal localized resistance gene variants in bacteria are statistically analyzed according to different genera.

2. The method for establishing a potential natural strain bank for mining the drug resistance gene tetX according to claim 1, characterized in that, Step (1) specifically includes the following steps: (1.1) Collect gene sequences from existing public drug resistance databases; the existing public drug resistance databases include CARD, Resifinder, and SARG; (1.2) Search the literature to obtain novel variant gene sequences; (1.3) Collect tetX encoded sequences annotated by NCBI; (1.4) Identify and verify the typing information of the tetX encoded sequences annotated by NCBI; (1.5) Construct a complete database of tetX gene polymorphisms: Collect the sequences obtained in steps (1.1), (1.2), and (1.4) and remove sequences with the same Accession ID, Start site, and End site information; (1.6) Perform multiple sequence alignment of the nucleic acid sequences of the gene set using MUSCLE; generate a phylogenetic likelihood tree using IQ-TREE.

3. The method for establishing a potential natural strain bank for mining the drug resistance gene tetX according to claim 1, characterized in that, The host information mentioned in step (2) is the bacterial classification information corresponding to the bacterial species corresponding to the gene sequence; the location information is divided into chromosome and plasmid; gene type: gene sequences with >90% consistency with the gene set sequence are identified as gene types that are the same as the corresponding alignment sequence; while gene sequences with <90% consistency with the gene set sequence are identified as potential gene types that are the same as the corresponding alignment sequence.

4. The method for establishing a potential natural strain bank for mining the drug resistance gene tetX according to claim 1, characterized in that, The minimum sequence length in step (2) is 1060~1080bp, and the minimum sequence consistency is 80%~90%.

5. The method for establishing a potential natural strain bank for mining the drug resistance gene tetX according to claim 1, characterized in that, The nucleic acid alignment in step (2) uses the NCBI BLASTn tool.

6. The method for establishing a potential natural strain bank for mining the drug resistance gene tetX according to claim 1, characterized in that, The alignment parameters for nucleic acid alignment in step (2) are set to -outfmt 6 output format.

7. The method for establishing a potential natural strain bank for mining the drug resistance gene tetX according to claim 1, characterized in that, Step (2) involves extracting gene sequences from the alignment region by using the Batch Entrez tool on the NCBINucleotide page to download the FASTA and summary files of these sequences in batches according to the sequence IDs of the screening results. Then, based on the sequence ID, alignment start point, and alignment end point information in the alignment result file, Bedtools is used to extract the sequence of the alignment region; MAFFT is used to adjust the orientation uniformly.

8. A repository of potential natural bacterial strains for mining the drug resistance gene tetX, characterized in that, Obtained by any one of the methods described in claims 1-7.

9. A method for verifying potential natural strains of the drug resistance gene tetX, characterized in that, Choose one or both of (a) and (b): (a) BLAST protocol based on the level of assembled contigs of genome sequence Step 1: Download the target strain genome: Download all assembled contig-level genome sequence data of the target strain in batches; Step 2: Identify drug resistance genes in the genomic sequence: Using the complete database of the drug resistance gene tetX multi-genotyping constructed in step (1) as the index database, and the genomic sequence of the target strain as the query database, perform nucleic acid alignment; Screening conditions: Alignment length ≥ minimum sequence length; Nucleic acid identity ≥ 90%; Coverage ≥ 90%; (b) Mapping scheme based on WGS sequencing reads Step 1: Download target strain sequencing data: Download all sequencing library sequence data of the target strain in batches; Step 2: Construction of the index database: Using bowtie2, construct the index database for subsequent bowtie2 alignments based on the complete database of tetX multi-genotyping of the drug resistance gene constructed in step (1); Step 3: Alignment of raw sequencing data; Step 4: Alignment file conversion and processing: Use samtools to convert the alignment result SAM file to BAM format and sort it by genomic coordinates. Then use samtools to convert the sorted BAM file into a readable text format. Finally, use bedtools to calculate the sequencing coverage depth of each genomic location. Step 5: Calculate mapping parameters Based on the sequencing coverage depth result file obtained in step 4, calculate the following parameters: a. Average depth = Sum of depths at all locations of the gene / Length of the gene (number of locations) b. Coverage ratio = Number of locations with depth greater than 0 / Total number of locations for this gene c. Maximum coverage gap: Calculate the maximum length of a continuous gap at a depth of 0 on this gene. d. Continuous Coverage: Calculate the maximum length of consecutive non-zero values ​​on the gene. e. Continuous coverage: Calculated as the maximum length of consecutive non-zero values ​​on the gene divided by the length of the gene (number of positions). Step 6: Filter mapping results Based on the calculation parameters in step 5, set the following thresholds: average depth threshold: >= 5x; coverage ratio threshold: >= 80%; maximum gap threshold: <= 100bp. Sequencing data that meets the above thresholds contain the target drug resistance gene tetX.

10. The method for verifying potential natural strains of the drug resistance gene tetX according to claim 9, characterized in that, The minimum sequence length is 1060~1080bp.