Prediction method and system for antibiotic resistance gene based on multi-channel Transform

By combining protein sequence and structural information with a multi-channel Transformer architecture, the problem of low accuracy in existing antibiotic resistance gene prediction methods is solved, and high-precision antibiotic resistance gene identification and classification is achieved.

CN120954513APending Publication Date: 2025-11-14UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511074606.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing methods for predicting antibiotic resistance genes rely on amino acid or DNA sequence information, failing to fully utilize protein structural information and solvent accessibility. This results in low accuracy of models in identifying antibiotic resistance genes with low sequence homology, and traditional methods are time-consuming or complex to operate.

Method used

Employing a multi-channel Transformer architecture, combining protein sequence, secondary structure, and residue surface accessibility information, this study extracts multimodal features and optimizes attention weight distribution through a multi-head self-attention mechanism and a dual-constraint regularization strategy to capture key functional sites.

Benefits of technology

It significantly improved the identification accuracy and functional site resolution of antibiotic resistance genes, achieving an AUC-ROC of 99.23% for binary classification tasks, increasing the Matthews correlation coefficient to 92.74%, and achieving a prediction accuracy of 92.42% for multi-class classification tasks, while reducing the false negative rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954513A_ABST
    Figure CN120954513A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for predicting an antibiotic resistance gene based on a multi-channel Transform, and belongs to the field of bioinformatics. According to the method, a secondary structure sequence of a protein sequence and residue surface accessibility and relative accessibility are firstly obtained, after standardization processing, a multi-channel Transform architecture is used for extracting modal features, features are fused through average pooling and a full connection layer, then double-constraint regularization strategy optimization is introduced, and finally performance is verified and evaluated through an independent test set. According to the method, protein sequences, secondary structures and solvent accessibility information are integrated, a multi-head self-attention mechanism and double-constraint regularization are utilized, capture of key functional sites is enhanced, multi-modal feature collaboration is achieved, and a new scheme is provided for drug-resistant gene recognition and classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics, and in particular to a method and system for predicting antibiotic resistance genes based on a multi-channel Transformer. Background Technology

[0002] With the widespread use of antibiotics in hospitals, communities, and the environment, the global spread of antibiotic resistance (AR) has become a major public health challenge in the 21st century. According to the World Health Organization's (WHO) 2022 Global Antimicrobial Resistance (AMR) Surveillance Report, drug-resistant infections cause more than 1.27 million deaths annually, a number projected to climb to 10 million by 2050, with related healthcare costs increasing to $100 trillion. Antibiotic resistance genes (ARGs), as the core carriers of resistance transmission, spread among pathogens, environmental microorganisms, and even human commensal bacteria through horizontal gene transfer mechanisms, forming complex "resistome" networks. This process not only accelerates the emergence of multidrug-resistant (MDR) and extensively drug-resistant (XDR) strains but also seriously threatens the therapeutic efficacy of first-line antibiotics in clinical practice. Therefore, developing high-precision ARG identification tools is of urgent significance for resistance monitoring, novel antibiotic development, and nosocomial infection control.

[0003] Traditional antibiotic susceptibility testing (AST) directly reflects the phenotypic resistance characteristics of strains by measuring the growth inhibition of microorganisms at specific antibiotic concentrations in vitro. However, AST is not only time-consuming but also cannot identify recessive resistance genes that have not yet expressed a resistance phenotype, and it cannot be applied to the analysis of unculturable microbial communities. To overcome the limitations of AST, functional metagenomics, through steps such as constructing metagenomic DNA libraries, recombinant expression, and screening, directly captures functional ARGs from environmental or clinical samples. However, this method is complex and requires a high level of experimental skill.

[0004] Sequence similarity-based alignment tools (such as BLAST, Bowtie, and DIAMOND) align target sequences and predict possible ARG categories by setting sequence similarity thresholds and alignment length requirements. For example, the CARD database relies on manually compiled antibiotic resistance ontology (ARO) and BLAST alignments to annotate ARGs. However, because these methods heavily depend on databases of known ARGs, they struggle to effectively identify novel ARGs with significant sequence variations. Furthermore, a fixed global sequence homology threshold may lead to some true ARGs being misclassified as non-ARGs, increasing the false negative rate.

[0005] In recent years, deep learning-based classification models have demonstrated high accuracy in ARG prediction by automatically extracting sequence features through neural networks. For example, DeepARG uses convolutional neural networks (CNNs) for classification and annotation of metagenomic data; HMD-ARG employs a hierarchical multi-task deep learning framework, combining information on antibiotic resistance categories, resistance mechanisms, and gene mobility for ARG annotation; ARGNet combines unsupervised learning with an autoencoder and multi-classification CNNs for ARG prediction; and ARG-SHINE integrates information such as sequence similarity, protein domains, protein families, and sequence motifs to improve the ability to identify ARGs. Although deep learning methods have performed well in ARG prediction tasks, most current models still rely primarily on amino acid or DNA sequence information, failing to fully utilize protein structural information and its relative solvent accessibility (RSA) information. This may limit the model's generalization ability, especially when identifying ARGs with low sequence homology, making them susceptible to sequence variations and thus reducing prediction accuracy. Summary of the Invention

[0006] The purpose of this invention is to overcome one or more shortcomings of the prior art and provide a method and system for predicting antibiotic resistance genes based on multi-channel Transformer.

[0007] The objective of this invention is achieved through the following technical solution:

[0008] A method for predicting antibiotic resistance genes based on a multi-channel Transformer, the method comprising:

[0009] S1. Cluster the protein sequences of antibiotic resistance genes, and divide them into training, validation, and test sets based on the clusters; use the protein sequences to obtain the relative accessibility of their secondary structure sequences and residue surface accessibility continuous values;

[0010] S2. The raw data is standardized, and the amino acid sequence and secondary structure are encoded into index vectors respectively. The RSA value is normalized to retain its surface exposure characteristics, and the sequence length is unified by zero padding.

[0011] S3. A multi-channel Transformer architecture is used to extract features of each modality. Sequence and structural information are transformed by embedding layers and positional encoding. A multi-head self-attention mechanism is combined to capture global dependencies, and a feedforward network is used to further mine nonlinear features.

[0012] S4. The feature fusion stage integrates the three-channel semantic information through average pooling and a fully connected layer to generate a comprehensive representation;

[0013] S5. A dual-constraint regularization strategy is introduced. On the one hand, the attention is forced to focus on key residues by minimizing entropy, thus suppressing noise interference. On the other hand, the attention weights are clustered in local continuous regions by constraining the Gaussian kernel, which strengthens the modeling of the spatial regularity of α-helices and β-sheets.

[0014] S6. The method performance was evaluated using cross-validation on an independent test set, and the impact of different loss weight combinations on the method performance was assessed through ablation experiments.

[0015] Further, step S1 includes setting a sequence identity threshold of 90% for unique antibiotic resistance genes and performing clustering. The data is then strictly partitioned according to the clusters: the clusters are randomly assigned to the training set, validation set, and test set in an 8:1:1 ratio. The original protein sequence is used to obtain its secondary structure information using SCRATCH-1D: H=α-helix, G=310-helix, I=π-helix, E=β-strand, B=β-bridge, T=turn, S=bend, C=coil, and the relative accessibility sequence of continuous values ​​for residue surface accessibility. This information is used for encoding in step S2.

[0016] Furthermore, in step S2, the protein sequence and secondary structure sequence are digitized using an amino acid dictionary and a secondary structure dictionary, and relative accessibility is normalized. The amino acid dictionary contains 20 standard amino acids and a terminator. The secondary structure dictionary contains 8 DSSP secondary structure types. .

[0017] Furthermore, in step S2, the RSA sequence is normalized by mini-max scaling to the [0,1] interval, and the amino acid sequence, secondary structure sequence and RSA sequence are uniformly lengthened using a zero-filling strategy.

[0018] Furthermore, in step S3, feature extraction is performed using the multi-head attention mechanism and feedforward neural network in the Transformer. The multi-head attention mechanism is used to calculate the relationship between each position in the normalized protein sequence and secondary structure sequence and other position representations, generating new feature representations that incorporate contextual information, thereby capturing the dependencies between different positions in the input sequence. Then, the nonlinear transformation of the feedforward neural network is used to extract more complex features of the sequence, and the output of the self-attention mechanism is processed. Through the stacking of these layers, deep features of the input protein sequence, secondary structure, and RSA information are gradually extracted.

[0019] Furthermore, in step S4, each channel is independently subjected to feature extraction and average pooling, and finally the features of the three channels are spliced ​​and fused.

[0020] Furthermore, in step S5, the proposed regularization method enhances the model's ability to capture key features through a dual constraint mechanism. First, entropy minimization is used, which will cause the attention distribution to exhibit a peak feature. According to information theory, entropy minimization corresponds to the maximum amount of information. When attention is concentrated in a few positions, the model can more clearly capture the interaction between key residues. Then, local continuity constraint is used, which is achieved by maximizing the similarity between the attention weights and the Gaussian kernel. The secondary structure of proteins is usually formed by continuous residues, which forces the attention weights to concentrate within a local window centered on the current position.

[0021] Furthermore, in step S5, a balance is achieved between attention focus and local smoothness by adjusting the weights of the entropy minimization loss function and the local continuity loss function. Entropy minimization forces the model to focus on a few key positions, while entropy regularization reduces attention to irrelevant positions and reduces noise interference. The α-helix structure depends on the synergistic effect of continuous residues, and continuity regularization makes it easier for the model to capture such patterns. Amino acid sequences, secondary structures, and RSA have different characteristics, and independent regularization allows each channel to retain its own characteristics, ultimately integrating information through the fusion layer.

[0022] Furthermore, in step S6, the method performance is evaluated using an independent test set, and an early stopping mechanism is set during training to prevent overfitting. The impact of different loss weight combinations on the method performance is evaluated through an ablation experiment system, and the cross-entropy loss weight is finally determined. The evaluation system covers six indicators: precision, recall, area under the AUC-ROC curve, accuracy, Matthews correlation coefficient, and F1 score. The evaluation is carried out in the binary and multi-classification tasks of antibiotic resistance. At the same time, in order to verify the effectiveness of multimodal feature fusion, four sets of control experiments are designed: single-channel, dual-channel, complete three-channel model, and three-channel + dual-constraint regularization model, which confirms the key role of multimodal information synergy in drug resistance prediction.

[0023] A prediction system for antibiotic resistance genes based on a multichannel Transformer, which is used to perform a prediction method for antibiotic resistance genes based on a multichannel Transformer.

[0024] The beneficial effects of this invention are:

[0025] (1) By integrating protein sequence, secondary structure and solvent accessibility (RSA) information, the traditional model’s dependence on single sequence features is broken, and the ARGs recognition accuracy and functional site resolution capability are significantly improved.

[0026] (2) A multi-head self-attention mechanism is used to mine the global dependencies of multimodal features, and a dual-constraint regularization strategy (entropy minimization and local continuity constraint) is introduced to optimize the attention weight distribution and enhance the ability to capture key functional sites.

[0027] (3) In the binary classification task, the AUC-ROC reached 99.23% and the Matthews correlation coefficient (MCC) increased to 92.74%. In the multi-class classification task, the prediction accuracy for 15 classes of antibiotics reached 92.42% and the AUC-PR value was 99.65%, which also verified the effectiveness of multimodal feature fusion. It has stronger predictive ability than existing antibiotic resistance genes and provides a new solution for the identification and classification of drug resistance genes. Attached Figure Description

[0028] Figure 1 This is a model architecture diagram of the antibiotic resistance gene prediction method using the multi-channel Transformer of the present invention.

[0029] Figure 2 This is a flowchart of the method for predicting antibiotic resistance genes using the multi-channel Transformer of the present invention.

[0030] Figure 3 The performance of the multi-channel Transformer method for predicting antibiotic resistance genes in a binary classification task under different models is shown in the figure.

[0031] Figure 4 The performance of the multi-classification task prediction method for antibiotic resistance genes using the multi-channel Transformer under different models is shown in the figure.

[0032] Figure 5 Performance graphs of a multi-channel Transformer method for predicting antibiotic resistance genes in a binary classification task under different optimization strategies.

[0033] Figure 6 The graph shows the multi-classification prediction performance of the multi-channel Transformer method for predicting antibiotic resistance genes under different optimization strategies. Detailed Implementation

[0034] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] A method for predicting antibiotic resistance genes based on a multi-channel Transformer is provided, the method comprising:

[0036] S1. Cluster the protein sequences of antibiotic resistance genes, and divide them into training, validation, and test sets based on the clusters; use the protein sequences to obtain the relative accessibility of their secondary structure sequences and residue surface accessibility continuous values;

[0037] S2. The raw data is standardized, and the amino acid sequence and secondary structure are encoded into index vectors respectively. The RSA value is normalized to retain its surface exposure characteristics, and the sequence length is unified by zero padding.

[0038] S3. A multi-channel Transformer architecture is used to extract features of each modality. Sequence and structural information are transformed by embedding layers and positional encoding. A multi-head self-attention mechanism is combined to capture global dependencies, and a feedforward network is used to further mine nonlinear features.

[0039] S4. The feature fusion stage integrates the three-channel semantic information through average pooling and a fully connected layer to generate a comprehensive representation;

[0040] S5. A dual-constraint regularization strategy is introduced. On the one hand, the attention is forced to focus on key residues by minimizing entropy, thus suppressing noise interference. On the other hand, the attention weights are clustered in local continuous regions by constraining the Gaussian kernel, which strengthens the modeling of the spatial regularity of α-helices and β-sheets.

[0041] S6. The method performance was evaluated using cross-validation on an independent test set, and the impact of different loss weight combinations on the method performance was assessed through ablation experiments.

[0042] Further, step S1 includes setting a sequence identity threshold of 90% for unique antibiotic resistance genes and performing clustering. The clusters are then strictly partitioned: randomly assigned to the training, validation, and test sets in an 8:1:1 ratio. The original protein sequence is used to obtain its secondary structure information using SCRATCH-1D: H=α-helix, G=310-helix, I=π-helix, E=β-strand, B=β-bridge, T=turn, S=bend, C=coil, and the relative accessibility sequence of continuous values ​​for residue surface accessibility. This information is used for encoding in step S2.

[0043] Furthermore, in step S2, the protein sequence and secondary structure sequence are digitized using an amino acid dictionary and a secondary structure dictionary, and relative accessibility is normalized. The amino acid dictionary contains 20 standard amino acids and a terminator. The secondary structure dictionary contains 8 DSSP secondary structure types. .

[0044] Furthermore, in step S2, the RSA sequence is normalized by mini-max scaling to the [0,1] interval, and the amino acid sequence, secondary structure sequence and RSA sequence are uniformly lengthened using a zero-filling strategy.

[0045] Furthermore, in step S3, feature extraction is performed using the multi-head attention mechanism and feedforward neural network in the Transformer. The multi-head attention mechanism is used to calculate the relationship between each position in the normalized protein sequence and secondary structure sequence and other position representations, generating new feature representations that incorporate contextual information, thereby capturing the dependencies between different positions in the input sequence. Then, the nonlinear transformation of the feedforward neural network is used to extract more complex features of the sequence, and the output of the self-attention mechanism is processed. Through the stacking of these layers, deep features of the input protein sequence, secondary structure, and RSA information are gradually extracted.

[0046] Furthermore, in step S4, each channel is independently subjected to feature extraction and average pooling, and finally the features of the three channels are spliced ​​and fused.

[0047] Furthermore, in step S5, the proposed regularization method enhances the model's ability to capture key features through a dual constraint mechanism. First, entropy minimization is used, which will cause the attention distribution to exhibit a peak feature. According to information theory, entropy minimization corresponds to the maximum amount of information. When attention is concentrated in a few positions, the model can more clearly capture the interaction between key residues. Then, local continuity constraint is used, which is achieved by maximizing the similarity between the attention weights and the Gaussian kernel. The secondary structure of proteins is usually formed by continuous residues, which forces the attention weights to concentrate within a local window centered on the current position.

[0048] Furthermore, in step S5, a balance is achieved between attention focus and local smoothness by adjusting the weights of the entropy minimization loss function and the local continuity loss function. Entropy minimization forces the model to focus on a few key positions, while entropy regularization reduces attention to irrelevant positions and reduces noise interference. The α-helix structure depends on the synergistic effect of continuous residues, and continuity regularization makes it easier for the model to capture such patterns. Amino acid sequences, secondary structures, and RSA have different characteristics, and independent regularization allows each channel to retain its own characteristics, ultimately integrating information through the fusion layer.

[0049] Furthermore, in step S6, the method performance is evaluated using an independent test set, and an early stopping mechanism is set during training to prevent overfitting. The impact of different loss weight combinations on the method performance is evaluated through an ablation experiment system, and the cross-entropy loss weight is finally determined. The evaluation system covers six indicators: precision, recall, area under the AUC-ROC curve, accuracy, Matthews correlation coefficient, and F1 score. The evaluation is carried out in the binary and multi-classification tasks of antibiotic resistance. At the same time, in order to verify the effectiveness of multimodal feature fusion, four sets of control experiments are designed: single-channel, dual-channel, complete three-channel model, and three-channel + dual-constraint regularization model, which confirms the key role of multimodal information synergy in drug resistance prediction.

[0050] A prediction system for antibiotic resistance genes based on multichannel Transformer is provided, which is used to perform a prediction method for antibiotic resistance genes based on multichannel Transformer.

[0051] See Figure 1 This embodiment proposes a multi-channel Transformer framework that combines protein sequence, secondary structure, and RSA information to improve the accuracy of ARG classification. This method comprehensively mines the sequence-structure-function relationships of ARGs through multi-channel feature extraction, self-attention mechanism modeling, and regularization strategy optimization, providing a new solution for the identification and classification of drug resistance genes. Specific steps are as follows: Figure 2 As shown:

[0052] S1. Data Acquisition:

[0053] Antibiotic resistance gene sequences were collected and organized from six publicly available antibiotic resistance gene (ARG) databases, including the Comprehensive Antibiotic Resistance Database (CARD), AMRFinder, ResFinder, DeepARG, MEGARes, and HMD-ARG, yielding a total of 48,615 ARG amino acid sequences. To remove redundant sequences, CD-HIT (with a sequence identity threshold of 100% and full-length alignment coverage of 100%) was used for sequence clustering, ultimately resulting in 27,022 non-redundant ARG sequences. The most prevalent resistance categories in the database and their proportions were: β-lactams (8,080 sequences, 29.9%), multidrug-resistant strains (5,148 sequences, 19.1%), and bacitracins (4,219 sequences, 15.6%). Compared to the integrated database used in this study, two larger reference databases—DeepARG-DB (DeepARG database, 12,279 sequences) and HMD-ARG-DB (HMD-ARG database, 17,282 sequences)—cover 28 and 33 ARG categories, respectively. Of the 43 known resistance categories, 26 categories are present in both this study's database, DeepARG-DB, and HMD-ARG-DB.

[0054] Given the widespread existence of highly similar sequence variants among ARGs (such as members of the β-lactamase family), to avoid biased model performance evaluation due to excessive similarity (i.e., data leakage) between the training, validation, and test sets, a secondary clustering process was performed on the aforementioned non-redundant ARG sequences (with a sequence identity threshold set to 90%). This step generated 6,200 clusters. Dataset partitioning was strictly performed on a cluster-by-cluster basis: clusters were randomly assigned to the training, validation, and test sets in an 8:1:1 ratio. This strategy ensures that all sequences belonging to the same cluster appear only in a single subset of the data, thereby minimizing the risk of data leakage.

[0055] To construct a high-quality negative sample dataset (representing non-antibiotic-resistant proteins), 70,000 protein sequences were randomly extracted from the Swiss-Prot database. First, internal redundancy was removed using CD-HIT (with a 90% sequence identity threshold). Then, strict homology filtering was performed using the DIAMOND tool (parameters: 30% sequence identity threshold, 80% query sequence coverage, E-value 1e-10) to exclude sequences significantly similar to the self-built ARG core database (27,022 sequences). Finally, the remaining sequences were scanned using the InterProScan functional annotation system, and all sequences carrying known antibiotic resistance-related functional domains (e.g., β-lactamase catalytic domain, tetracycline efflux pump domain, etc.) were manually removed. After this screening process, 55,347 rigorously selected high-confidence non-ARG protein sequences were obtained, forming a balanced negative sample control dataset for model training and evaluation.

[0056] The sequence was then standardized, and the original sequence contained a protein sequence consisting of 20 standard amino acids. A sequence of secondary structures composed of eight DSSP secondary structure types (H=α-helix, G=310-helix, I=π-helix, E=β-strand, B=β-bridge, T=turn, S=bend, C=coil). , representing the relative accessibility (RSA) sequence of continuous values ​​of residue surface accessibility.

[0057] S2, Sequence Standardization:

[0058] First, use the amino acid field. and two-level structure dictionary Sequence digitization of protein sequences and secondary structure sequences. Normalization is performed:

[0059] ;

[0060] ;

[0061] in, For sequence length,

[0062] ;

[0063] It contains 22 standard amino acids and a terminator. It includes 8 types of DSSP secondary structure.

[0064] Then, the RSA sequence is normalized using a min-max scaling normalization to the [0,1] interval:

[0065] ;

[0066] A zero-padding strategy is used to unify the sequence length. :

[0067] ;

[0068] S3, Feature Extraction:

[0069] The Transformer is a sequence processing method that utilizes an attention mechanism. It is widely used in various fields and has achieved excellent performance. This paper utilizes the multi-head attention mechanism and feedforward neural network in the Transformer for feature extraction.

[0070] First, a multi-head attention mechanism is used to compute the relationship between each position in the normalized protein sequence and secondary structure sequence and other position representations, generating new feature representations that incorporate contextual information, thereby capturing the dependencies between different positions in the input sequence:

[0071] ;

[0072] ;

[0073] ;

[0074] ;

[0075] ;

[0076] ;

[0077] ;

[0078] in, , It is the embedded dimension. , , It is the position in the sequence, and the range is... , It is the embedded dimension. Initialize to , Initialize to , Initialize to , , , They are , , Learnable weight matrices, all of which have the following shape: , , For attention head count

[0079] Then, a nonlinear transformation of the feedforward neural network is used to extract more complex features of the sequence, and the output of the self-attention mechanism is processed.

[0080] ;

[0081] ;

[0082] in, , For the number of attention heads, , , , and This is a bias term.

[0083] By stacking these layers, deep features of the input protein sequence, secondary structure, and RSA information are gradually extracted.

[0084] ;

[0085] ;

[0086] S4, Feature Fusion:

[0087] Feature extraction is performed independently for each channel:

[0088] ;

[0089] in, , This represents the average pooling operation along the sequence dimension.

[0090] S5. Regularization optimization:

[0091] Our proposed regularization method enhances the model's ability to capture key features through a dual-constraint mechanism, and its complete loss function is defined as:

[0092] ;

[0093] Attention matrix for each channel definition:

[0094] ;

[0095] in, This is a numerically stable term. The loss function exhibits sparsity-induced characteristics and gradient regulation mechanisms. When When it approaches a uniform distribution, Reaching the maximum value; minimizing the loss will cause the attention distribution to exhibit a peaked characteristic. According to information theory, minimizing entropy corresponds to maximizing information content. When attention is focused on a few locations, the model can more clearly capture the interactions between key residues. By calculating the gradient:

[0096] ;

[0097] when When approaching a uniform distribution, When the gradient is large, the absolute value of the gradient is small, allowing attention weights at important locations to be retained. When the gradient is small, the absolute value of the gradient is large, which forces a reduction in attention weights at irrelevant locations.

[0098] Define the position-related Gaussian similarity kernel matrix :

[0099] ;

[0100] Where the normalization factor make sure This avoids the sequence length affecting the loss scale. Controlling the width of the Gaussian kernel, The larger the value, the wider the allowed local range.

[0101] The local continuity loss is achieved by maximizing the similarity between the attention weights and the Gaussian kernel:

[0102] ;

[0103] The secondary structure of proteins (such as α-helices and β-sheets) is typically formed by consecutive residues, maximizing... This forces attention weights to be focused on the local window centered on the current location.

[0104] By adjusting and This achieves a balance between attentional focus and local smoothness. Entropy minimization forces the model to focus on a few key locations, through entropy regularization. Reduce focus on irrelevant positions and decrease noise interference. Structures such as α-helices rely on the synergistic effect of consecutive residues; continuity regularization makes the model more likely to capture such patterns. Amino acid sequences, secondary structures, and RSAs have different characteristics; independent regularization allows each channel to retain its own characteristics, ultimately integrating information through a fusion layer.

[0105] S5. Experimental verification:

[0106] The model's performance was evaluated on training, validation, and test sets strictly divided in an 8:1:1 ratio. An early stopping mechanism (patience value of 20) was implemented during training to prevent overfitting. Ablation experiments were conducted to systematically evaluate the impact of different loss weight combinations on model performance, ultimately determining the cross-entropy loss weights. Locating loss weights ( , The optimizer uses the Adam algorithm with an initial learning rate of 0.001, a fixed batch size of 64, and a maximum training epoch of 100. The hardware platform is configured with an NVIDIA RTX 4090 GPU and a CUDA 12.6 acceleration environment. Performance metrics include precision, recall, accuracy, and F1 score.

[0107] like Figure 3 As shown, MCT-ARG outperforms all benchmark methods across all evaluation metrics for binary classification tasks. Its F1-score reaches 0.9487, significantly outperforming DeepARG (0.7320) and RGI (0.6062), and significantly outperforming ARG-SHINE (0.4087). Although RGI achieves a precision of 1.0000, its recall is only 0.4349, reflecting a high false negative rate due to its highly conservative prediction strategy. DeepARG exhibits a similar precision-recall imbalance. ARG-SHINE fails entirely (Accuracy = 0.3099), indicating insufficient generalization ability and overfitting risk. These results highlight MCT-ARG's superior generalization ability and robustness, maintaining both high precision (0.9760) and recall (0.9230), thus accurately identifying ARGs while minimizing false negatives.

[0108] In more challenging multi-class classification tasks, MCT-ARG continues its outstanding performance. Figure 4 As shown, MCT-ARG's F1 score is 0.9240, an absolute improvement of 26.2% over the best baseline DeepARG (0.7320); this indicates that the model can distinguish subtle functional differences between resistance mechanisms. In contrast, RGI reproduces the imbalance between high precision (0.9279) and low recall (0.3915), consistent with its strict homology-based criteria and conservative detection strategy. ARG-SHINE performed the worst on this task, with all metrics stagnating at 0.3332, indicating insufficient class-level resolution.

[0109] MCT-ARG demonstrates superior performance in both binary and multi-class classification tasks, exhibiting strong robustness and domain adaptability. Unlike models that rely solely on sequence homology (e.g., RGI) or are prone to overfitting (e.g., ARG-SHINE), MCT-ARG combines deep representation learning with an efficient generalization strategy. This results in higher recall, lower false positive rates, and improved accuracy in distinguishing different resistance types for detecting novel drug resistance genes, making it a promising tool for drug resistance gene monitoring and resistance mechanism characterization.

[0110] To systematically evaluate the effectiveness of the proposed multi-feature fusion strategy and dual-constraint regularization mechanism, this study designed four sets of control experiments, covering: (1) a single-channel model (using only amino acid sequences as input features); (2) a dual-channel model (fusing amino acid sequences and protein secondary structure information); (3) a three-channel model (introducing relative solvent accessibility (RSA) information based on the dual-channel model); and (4) a three-channel + dual-constraint regularization model (integrating entropy minimization and local continuity constraints in the three-channel model). Performance metrics included precision, recall, AUC, accuracy, MCC, and F1 score.

[0111] like Figure 5 As shown, in binary classification tasks, model performance systematically improves with the increase of the number of input feature channels. Compared to the single-channel model, the three-channel model shows significant improvements in key metrics such as AUC (0.9932), F1 score (0.9439), and MCC (0.9196), confirming that fusing secondary structure and structural information such as RSA significantly enhances the model's ability to identify adversarial key sites. Crucially, after introducing dual-constraint regularization, the model achieves optimal performance in comprehensive metrics such as Precision (0.9760), Accuracy (0.9691), and MCC (0.9274), strongly validating that this structure-guided regularization strategy can effectively improve the model's attention to adversarial regions and its robustness in discrimination.

[0112] In multi-classification tasks, such as Figure 6 As shown, the aforementioned performance improvement trends are consistent and significant. Compared to the single-channel model, the three-channel model improves the AUC from 0.9860 to 0.9924 and the F1 score from 0.8515 to 0.8869. Further applying regularization constraints to the three-channel model further enhances the overall performance, achieving an Accuracy of 0.9242 and an MCC of 0.9097, indicating that the proposed fusion and regularization strategy is also applicable to and significantly improves the performance of complex multi-class resistance recognition tasks.

[0113] In summary, the ablation experiment results clearly demonstrate the effectiveness and key contributions of the proposed multi-feature fusion framework and dual-constraint regularization method in the task of antibiotic resistance gene identification.

[0114] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A method for predicting antibiotic resistance genes based on multi-channel Transformer, characterized in that, The method includes: S1. Cluster the protein sequences of antibiotic resistance genes, and divide them into training, validation, and test sets based on the clusters; use the protein sequences to obtain the relative accessibility of their secondary structure sequences and residue surface accessibility continuous values; S2. The raw data is standardized, and the amino acid sequence and secondary structure are encoded into index vectors respectively. The RSA value is normalized to retain its surface exposure characteristics, and the sequence length is unified by zero padding. S3. A multi-channel Transformer architecture is used to extract features of each modality. Sequence and structural information are transformed by embedding layers and positional encoding. A multi-head self-attention mechanism is combined to capture global dependencies, and a feedforward network is used to further mine nonlinear features. S4. The feature fusion stage integrates the three-channel semantic information through average pooling and a fully connected layer to generate a comprehensive representation; S5. A dual-constraint regularization strategy is introduced. On the one hand, the attention is forced to focus on key residues by minimizing entropy, thus suppressing noise interference. On the other hand, the attention weights are clustered in local continuous regions by constraining the Gaussian kernel, which strengthens the modeling of the spatial regularity of α-helices and β-sheets. S6. The method performance was evaluated using cross-validation on an independent test set, and the impact of different loss weight combinations on the method performance was assessed through ablation experiments.

2. The method for predicting antibiotic resistance genes based on multi-channel Transformer according to claim 1, characterized in that, Step S1 includes setting a sequence identity threshold of 90% for unique antibiotic resistance genes and performing clustering. The data is then strictly divided according to the clusters: the clusters are randomly assigned to the training set, validation set, and test set in a ratio of 8:1:

1. The secondary structure information of the original protein sequence is obtained using SCRATCH-1D: H=α-helix, G=310-helix, I=π-helix, E=β-strand, B=β-bridge, T=turn, S=bend, C=coil and the relative accessibility sequence of continuous values ​​of residue surface accessibility, which is used for the encoding processing in step S2.

3. The method for predicting antibiotic resistance genes based on multi-channel Transformer according to claim 1, characterized in that, In step S2, the protein sequence and secondary structure sequence are digitized using an amino acid dictionary and a secondary structure dictionary, and relative accessibility is normalized. The amino acid dictionary contains 20 standard amino acids and terminators. The secondary structure dictionary contains 8 DSSP secondary structure types. .

4. The method for predicting antibiotic resistance genes based on multi-channel Transformer according to claim 1, characterized in that, In step S2, the RSA sequence is normalized by mini-max scaling to the [0,1] interval. The amino acid sequence, secondary structure sequence and RSA sequence are uniformly lengthened using a zero-filling strategy.

5. The method for predicting antibiotic resistance genes based on multi-channel Transformer according to claim 1, characterized in that, In step S3, feature extraction is performed using the multi-head attention mechanism and feedforward neural network in Transformer; the multi-head attention mechanism is used to calculate the relationship between each position in the normalized protein sequence and secondary structure sequence and other position representations to generate new feature representations that incorporate contextual information, thereby capturing the dependencies between different positions in the input sequence; then, the nonlinear transformation of the feedforward neural network is used to extract more complex features of the sequence, and the output of the self-attention mechanism is processed. By stacking these layers, deep features of the input protein sequence, secondary structure, and RSA information are gradually extracted.

6. The method for predicting antibiotic resistance genes based on multi-channel Transformer according to claim 1, characterized in that, In step S4, each channel is independently subjected to feature extraction and average pooling, and finally the features of the three channels are concatenated and fused.

7. The method for predicting antibiotic resistance genes based on multi-channel Transformer according to claim 1, characterized in that, In step S5, the proposed regularization method enhances the model's ability to capture key features through a dual constraint mechanism. First, entropy minimization is used, which will cause the attention distribution to exhibit a peak feature. According to information theory, entropy minimization corresponds to the maximum amount of information. When attention is concentrated in a few positions, the model can more clearly capture the interaction between key residues. Then, local continuity constraint is used, which is achieved by maximizing the similarity between attention weights and Gaussian kernels. The secondary structure of proteins is usually formed by continuous residues, which forces the attention weights to concentrate within a local window centered on the current position.

8. The method for predicting antibiotic resistance genes based on multi-channel Transformer according to claim 1, characterized in that, In step S5, a balance is achieved between attention focus and local smoothness by adjusting the weights of the entropy minimization loss function and the local continuity loss function. Entropy minimization forces the model to focus on a few key positions, while entropy regularization reduces attention to irrelevant positions and reduces noise interference. The α-helix structure depends on the synergistic effect of continuous residues, and continuity regularization makes it easier for the model to capture such patterns. Amino acid sequences, secondary structures, and RSA have different characteristics, and independent regularization allows each channel to retain its own characteristics. Finally, the information is integrated through the fusion layer.

9. The method for predicting antibiotic resistance genes based on multi-channel Transformer according to claim 1, characterized in that, In step S6, the method performance is evaluated using an independent test set, and an early stopping mechanism is set during training to prevent overfitting. The impact of different loss weight combinations on the method performance is evaluated through an ablation experiment system, and the cross-entropy loss weight is finally determined. The evaluation system covers six indicators: precision, recall, area under the AUC-ROC curve, accuracy, Matthews correlation coefficient, and F1 score. The evaluation is carried out in the binary and multi-classification tasks of antibiotic resistance. At the same time, in order to verify the effectiveness of multimodal feature fusion, four sets of control experiments are designed: single-channel, dual-channel, complete three-channel model, and three-channel + dual-constraint regularization model, which confirms the key role of multimodal information synergy in drug resistance prediction.

10. A prediction system for antibiotic resistance genes based on a multi-channel Transformer, characterized in that, The method for predicting antibiotic resistance genes based on a multichannel Transformer, as described in any one of claims 1 to 9, is employed.