Drug resistance gene and resistance class prediction method based on sequence features and protein language model
By combining amino acid features and protein language models, and utilizing XGBoost and CNN-BiLSTM networks, the accuracy problem of drug resistance gene identification in existing technologies has been solved, achieving more efficient and accurate prediction of drug resistance genes and resistance categories.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2026-03-24
AI Technical Summary
Current technologies rely on sequence similarity for accuracy in identifying drug resistance genes, resulting in a high false negative rate. This is especially true when there are gene variations or incomplete reference databases, making it difficult to accurately predict drug resistance genes and resistance categories.
A sequence feature-based and protein language model-based approach is adopted to predict drug resistance genes and resistance categories by calculating mutual information between amino acids, Fourier power spectrum features, and ESM2 matrix features, combined with XGBoost model and CNN-BiLSTM network. Multi-population particle swarm optimization and ensemble learning are used to improve prediction accuracy.
It significantly improves the accuracy and efficiency of drug resistance gene prediction, can identify potential novel drug resistance genes, reduces the bias of manual feature screening and the impact of noisy data, and improves the stability of the model and its generalization ability across databases.
Smart Images

Figure CN120954515B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of drug resistance prediction technology in bioinformatics, and involves technologies such as protein language models, machine learning models, and neural networks. Specifically, it is a method for predicting drug resistance genes and resistance categories based on sequence features and protein language models. Background Technology
[0002] Antibiotics are a crucial treatment for bacterial infections; however, the emergence of antibiotic resistance in bacteria poses a growing threat to human health and clinical treatment. Human activities, including the clinical misuse and overuse of antibiotics, are considered the main evolutionary drivers of antibiotic resistance. The emergence of resistance genes and their horizontal gene transfer among bacteria have led to the continuous emergence of multidrug-resistant bacteria and "superbugs," significantly reducing or even eliminating the efficacy of traditional antibiotics. Accurately assessing drug sensitivity and identifying resistance genes are essential foundations for addressing the bacterial resistance crisis. Timely and precise monitoring of bacterial resistance status can provide a scientific basis for relevant decision-making and response measures, thereby effectively curbing the development and spread of resistance.
[0003] Traditional methods for identifying drug resistance genes primarily rely on sequence alignment techniques. This involves using sequence similarity alignment algorithms (such as BLAST, Bowtie, and DIAMOND) to compare and annotate genomes, metagenomic contigs, or predicted open reading frames (ORFs) with drug resistance gene databases, employing fixed and typically high identity thresholds. The accuracy of this method depends heavily on the similarity between the query and reference sequences, leading to a high false negative rate, especially when there are gene variations or the reference database is incomplete.
[0004] Antibiotic resistance gene prediction technology has undergone a paradigm shift from rule-driven to data-driven approaches in recent years, with rapid advancements in deep learning and artificial intelligence bringing innovative breakthroughs to the field. Current technological development exhibits a clear progressive characteristic: from early feature engineering optimization to raw sequence parsing, and then to the application of pre-trained models and hybrid learning architectures, the focus of research is gradually shifting towards automated feature learning and improving cross-database generalization capabilities. Because amino acid sequences contain a wealth of biological information, simple sequence input often fails to fully capture their structural and functional features. The positions of amino acids in the sequence, the relationships between amino acids, and their physicochemical properties (such as hydrophobicity, polarity, and molecular size) can all significantly influence protein folding and function. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, this invention proposes a method for predicting drug resistance genes and resistance categories based on sequence features and protein language models. This method aims to improve the accuracy and efficiency of drug resistance gene prediction and also predict potential novel drug resistance genes, thereby providing a new solution for drug resistance gene type prediction.
[0006] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:
[0007] The present invention provides a method for predicting drug resistance genes and resistance categories based on sequence features and protein language models, characterized by the following steps:
[0008] Step 1: Construct the amino acid dataset U1 and the drug resistance gene dataset U2;
[0009] Step 2, Calculation mutual information Conditional mutual information Fourier power spectrum characteristics Dipeptide composition Composition of spacer amino acid pairs ESM2 matrix , mutual information Composition of spacer amino acid pairs Triple matrix Labeled ESM2 feature matrix Unlabeled ESM2 feature matrix ; and will , , , , spliced as amino acid sequence characteristics ;Will , , spliced as Drug resistance gene sequence characteristics Then, the multi-population particle swarm optimization feature selection algorithm is used to... Processing is performed to obtain Characteristics of drug resistance gene sequences after screening ;
[0010] Step 3: Construct a drug resistance gene prediction model, including: an XGBoost model and a first CNN-BiLSTM network, and respectively... and After processing, the optimal prediction score of drug resistance genes in the XGBoost model is obtained. The optimal prediction score of drug resistance genes from the first CNN-BiLSTM network And then and Weighted processing is performed to obtain the predicted score of drug resistance gene integration. ;
[0011] if middle corresponding value Greater than the preset threshold , then it means The predicted result is a drug resistance gene; if it is smaller than [a certain value]... ,but Predicted to be a non-drug-resistant gene;
[0012] Step 4: Construct a drug resistance gene resistance category prediction model, including: an ML-KNN model, a pre-trained CNN-BiLSTM network, and a fine-tuned CNN-BiLSTM network; and use the ML-KNN model to predict drug resistance gene categories. After processing, the optimal prediction score for the resistance category of the ML-KNN model is obtained. Using a pre-trained CNN-BiLSTM network to... The process is performed to obtain a CNN-BiLSTM model with optimal parameters; the output layer of the CNN-BiLSTM model with optimal parameters is then replaced with... The classification layer is used to obtain a fine-tuned CNN-BiLSTM network, and then... After processing, the optimal prediction score for the resistance category of the fine-tuned CNN-BiLSTM network is obtained. And then and After weighting, a comprehensive prediction score for the resistance category is obtained. ;
[0013] like middle The Value Greater than the preset threshold , then it means Belongs to the Number of categories; if less than , then it means Not belonging to the first There are several categories.
[0014] The method for predicting drug resistance genes and resistance categories based on sequence features and protein language models described in this invention is characterized in that step 1 is performed as follows:
[0015] Step 1.1: Define a finite set of amino acids. ,in Indicates the first One amino acid, Indicates the first There are 10 amino acids, where H represents the total number of amino acids, and , Amino acids in food can be classified according to their physicochemical properties. Functional categories;
[0016] Step 1.2: Obtain the amino acid dataset ,in, Indicates the first A single amino acid sequence, Indicates the number of amino acid sequences, and The sequence length is ;
[0017] make The tag is ,like If the amino acid sequence is a drug resistance gene, then let ,like If the amino acid sequence is a non-drug resistance gene, then let ;
[0018] Step 1.3, Take middle The amino acid sequences constitute the drug resistance gene dataset. ,in Indicates the first The amino acid sequence of a drug resistance gene. The number of amino acid sequences representing the drug resistance gene;
[0019] definition Multi-category labels are denoted as , Indicates the first Each category of labels, and , Indicates the total number of categories, when When, it means Not belonging to the first Each category, when When, it means Belongs to the There are several categories.
[0020] Furthermore, step 2 is performed as follows:
[0021] Step 2.1, The Middle amino acids and the amino acids The resulting biamino acid pair is denoted as And calculate using equation (1) middle Includes about mutual information Thus obtain The total amount of mutual information :
[0022] (1)
[0023] In equation (1), yes middle Frequency of occurrence and They are middle and Frequency of occurrence;
[0024] Step 2.2: Calculate according to the process in Step 2.1. The Middle amino acids Includes information about the first amino acids mutual information Thus obtain The total amount of mutual information ;
[0025] Step 2.3, and spliced together mutual information vector Thus obtain mutual information ;Will and spliced together mutual information vector Thus obtain mutual information ;
[0026] Step 2.4, In order from the first amino acids , No. amino acids , No. amino acids The three amino acids that make up the group are denoted as , by the amino acids and the amino acids The resulting biamino acid pair is denoted as , by the amino acids and the amino acids The biamino acid pair formed is denoted as Calculate using equation (2) Under the conditions, and Mutual information between Thus obtain The sum of conditional mutual information :
[0027] (2)
[0028] In equation (2), yes middle Frequency of occurrence and They are middle and Frequency of occurrence yes middle Frequency of occurrence;
[0029] Step 2.5, and spliced together Conditional mutual information vector Thus obtain Conditional mutual information ;
[0030] Step 2.6: Calculate based on the cumulative Fourier power spectrum. Moment vector and central moment vector and will and spliced together Fourier power spectrum characteristics Thus obtain Fourier power spectrum characteristics ;
[0031] Step 2.7: Obtain using dipeptide composition method dipeptide composition matrix :
[0032] Step 2.8: Obtain the compositional feature encoding method using spacer amino acid pairs. Composition of spacer amino acid pairs and Composition of spacer amino acid pairs ;
[0033] Step 2.9, Inputting into the ESM2 model based on the Transformer architecture generates... Context-aware matrix And then Perform average pooling to obtain ESM2 feature vector Thus obtain ESM2 feature matrix Where L is the length of the vector after ESM2 encoding and pooling;
[0034] Step 2.10: Obtain the multi-label category following the process in Step 2.9. of Labeled ESM2 feature vectors Thus obtain Labeled ESM2 feature matrix ;
[0035] Remove The tag information is then obtained by following step 2.9. Unlabeled ESM2 feature vectors Thus obtain Unlabeled ESM2 feature matrix ;
[0036] Step 2.11: Obtain the descriptor using the join triplet method. triple matrix .
[0037] Furthermore, step 3 is performed as follows:
[0038] Step 3.1, The input is fed into the XGBoost model for training, and the drug resistance gene prediction score of the XGBoost model is calculated using equation (3). :
[0039] (3)
[0040] In equation (3), This represents the total number of decision trees in the XGBoost model. Indicates the sequence number of the decision tree. Indicates the first A decision tree pair The predicted score vector of drug resistance genes;
[0041] Step 3.2: Construct the loss function of the XGBoost model using equation (4). :
[0042] (4)
[0043] In equation (4), express middle Corresponding drug resistance gene prediction score;
[0044] Step 3.3: Optimize the model using the XGBoost gradient boosting algorithm and calculate the loss function. The model parameters are updated until convergence is achieved, thus obtaining the optimal prediction score for drug resistance genes in the XGBoost model. ;
[0045] Step 3.4, The input is processed in the first CNN-BiLSTM network to obtain the flattened local features. Then The input is processed in the BiLSTM layer of the first CNN-BiLSTM network for drug resistance gene prediction, thus obtaining the final hidden state of the BiLSTM. ; Then input it into the fully connected layer, and use equation (5) to... Drug resistance gene prediction scores converted to the first CNN-BiLSTM network :
[0046] (5)
[0047] In equation (5), This represents the weight matrix of the fully connected layer. For the bias term of the fully connected layer, This represents the Sigmoid activation function;
[0048] Step 3.5: Construct the loss function of the first CNN-BiLSTM network using equation (6). :
[0049] (6)
[0050] In equation (6), express middle Corresponding drug resistance gene prediction score;
[0051] Step 3.6: Optimize the first CNN-BiLSTM network using the Adam optimization algorithm and calculate the loss function. The network parameters are updated until convergence is achieved, thus obtaining the optimal prediction score for drug resistance genes from the first CNN-BiLSTM network. ;
[0052] Step 3.7, for and A weighted integration is performed, and a comprehensive drug resistance gene prediction score is obtained using equation (7). :
[0053] (7)
[0054] In equation (7), and They are and The weight.
[0055] Furthermore, step 4 is performed as follows:
[0056] Step 4.1, The data is input into the ML-KNN model for processing, thereby obtaining the resistance category prediction score of the ML-KNN model. ;
[0057] Step 4.2, Defined as anchor sample, for Add noise to obtain Positive samples belonging to the same category Random selection and Samples not belonging to the same category are considered negative samples. ,Will The input is fed into a pre-trained CNN-BiLSTM network, and the loss function is constructed using equation (8). :
[0058] (8)
[0059] In equation (8), This refers to a pre-trained CNN-BiLSTM network. These are marginal hyperparameters used to control the minimum distance between positive and negative samples. Indicates Euclidean distance;
[0060] Step 4.3: Optimize the pre-trained CNN-BiLSTM network using the backpropagation algorithm and calculate the loss function. The network parameters are updated until convergence is achieved, thus obtaining the CNN-BiLSTM model with optimal parameters.
[0061] Step 4.4: Replace the output layer of the CNN-BiLSTM model with the optimal parameters. A classification layer is added to form a fine-tuned CNN-BiLSTM network; The input is processed by a fine-tuned CNN-BiLSTM network to obtain the resistance class prediction score of the fine-tuned CNN-BiLSTM network. ;
[0062] Step 4.5: Construct the loss function of the fine-tuned CNN-BiLSTM network using equation (9). :
[0063] (9)
[0064] In equation (9), Indicates the first Drug resistance gene sequence exist The Middle Resistance category prediction scores for each category;
[0065] Step 4.6: Optimize the fine-tuned CNN-BiLSTM network using the backpropagation algorithm and calculate the loss function. The network parameters are updated until convergence is achieved, thus obtaining the optimal parameters for the drug resistance gene resistance category classification model. Finally, the resistance category prediction score of the fine-tuned CNN-BiLSTM network is obtained. ;
[0066] Step 4.7, for and A weighted integration is performed, and a comprehensive resistance category prediction score is obtained using equation (10). :
[0067] (10)
[0068] In equation (10), and They are and The weight.
[0069] The present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the drug resistance gene and resistance category prediction method, and the processor is configured to execute the program stored in the memory.
[0070] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program, when executed by a processor, performs the steps of the method for predicting drug resistance genes and resistance categories.
[0071] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0072] 1. Based on information theory and discrete Fourier transform, this invention employs three amino acid encoding methods to extract key features. Information theory helps optimize the representation and processing efficiency of amino acid sequences by quantifying the amount of information between amino acids and compressing information, while Fourier transform can help extract periodic and spectral features from amino acid sequences, explain their implicit structural regularities, and effectively identify repetition patterns and variation information.
[0073] 2. This invention utilizes delabeled ESM2 encoded features, constructs a pre-trained model through a CNN-BiLSTM deep learning architecture, and then fine-tunes it using labeled features. By leveraging the self-supervised characteristics of unlabeled data, it automatically learns and extracts complex features from amino acid sequences through deep neural networks, eliminating the need for manually selected feature sets and avoiding bias and information loss from manual features. It also exhibits more stable performance on noisy data and sequences with low homology.
[0074] 3. In the process of predicting drug resistance genes and their resistance categories, this invention not only uses machine learning models but also neural network architectures. Furthermore, it utilizes ensemble learning to fuse the individual optimal results of the two models, thereby mining potential patterns in amino acid sequences from multiple perspectives and significantly improving the prediction accuracy of the models. Attached Figure Description
[0075] Figure 1 This is a framework diagram of the method for predicting drug resistance genes and resistance categories based on sequence features and protein language models according to the present invention.
[0076] Figure 2 This is a classification diagram of amino acids based on dipole and volume scale in this invention;
[0077] Figure 3 This is a schematic diagram of the drug resistance gene prediction model in this invention;
[0078] Figure 4 This is a schematic diagram of the drug resistance gene resistance category prediction model in this invention. Detailed Implementation
[0079] In this embodiment, as Figure 1 As shown, a method for predicting drug resistance genes and resistance categories based on sequence features and protein language models is performed according to the following steps:
[0080] Step 1: Construct the amino acid dataset U1 and the drug resistance gene dataset U2;
[0081] Step 1.1: Define a finite set of amino acids. ,in Indicates the first One amino acid, Indicates the first There are 10 amino acids, where H represents the total number of amino acids, and It includes 20 common amino acids and an unknown amino acid X. Amino acids in food can be classified according to their physicochemical properties. The functional categories are defined as follows: Since the physicochemical properties of X are unknown, it is treated as one of S amino acids that are any different from X during the classification process. Figure 2 As shown.
[0082] Step 1.2: Obtain the amino acid dataset ,in, Indicates the first A single amino acid sequence, Indicates the number of amino acid sequences, and The sequence length is In this embodiment, a high-quality antimicrobial resistance gene dataset was constructed by integrating multiple databases, including HMD-ARG-DB, Resfinder, MEGARes, AMRFinderPlus, PLM-ARG-DB, and CARD. Redundant sequences were eliminated using a global alignment algorithm, ultimately yielding 27,524 non-redundant ARG reference sequences. The negative dataset (non-resistant gene sequences) was constructed using a matching sampling strategy, systematically selecting 27,524 non-ARG sequences from the UniProt knowledge base.
[0083] make The tag is ,like If the amino acid sequence is a drug resistance gene, then let ,like If the amino acid sequence is a non-drug resistance gene, then let .
[0084] Step 1.3, Take middle The amino acid sequences constitute the drug resistance gene dataset. ,in Indicates the first The amino acid sequence of a drug resistance gene. The number of amino acid sequences representing the drug resistance gene;
[0085] definition Multi-category labels are denoted as , Indicates the first Labels for each category and , Indicates the total number of categories, when When, it means Not belonging to the first Each category, when When, it means Belongs to the There are 22 drug resistance gene categories in this embodiment. Categories with fewer than 30 samples are merged into the "other" category.
[0086] Step 2, Calculation mutual information Conditional mutual information Fourier power spectrum characteristics Dipeptide composition Composition of spacer amino acid pairs ESM2 matrix , mutual information Composition of spacer amino acid pairs Triple matrix Labeled ESM2 matrix Unlabeled ESM2 matrix ;
[0087] Step 2.1, The Middle amino acids and the amino acids The resulting biamino acid pair is denoted as And calculate using equation (1) middle Includes about mutual information Thus obtain The total amount of mutual information :
[0088] (1)
[0089] In equation (1), yes middle Frequency of occurrence and They are middle and Frequency of occurrence.
[0090] Step 2.2: Calculate according to the process in Step 2.1. The Middle amino acids Includes information about the first amino acids mutual information Thus obtain The total amount of mutual information :
[0091] Step 2.3, and spliced together mutual information vector Thus obtain mutual information ;Will and spliced together mutual information vector Thus obtain mutual information .
[0092] Step 2.4, In order from the first amino acids , No. amino acids , No. amino acids The three amino acids that make up the group are denoted as , by the amino acids and the amino acids The resulting biamino acid pair is denoted as , by the amino acids and the amino acids The biamino acid pair formed is denoted as Calculate using equation (2) Under the conditions, and Mutual information between Thus obtain The sum of conditional mutual information :
[0093] (2)
[0094] In equation (2), yes middle Frequency of occurrence and They are middle and Frequency of occurrence yes middle Frequency of occurrence.
[0095] Step 2.5, and spliced together Conditional mutual information vector Thus obtain Conditional mutual information ;
[0096] Step 2.6: Calculate based on the cumulative Fourier power spectrum. Moment vector and central moment vector and will and spliced together Fourier power spectrum characteristics Thus obtain Fourier power spectrum characteristics .
[0097] Step 2.7: Obtain using dipeptide composition method dipeptide composition matrix :
[0098] Step 2.8: Obtain the compositional feature encoding method using spacer amino acid pairs. Composition of spacer amino acid pairs and Composition of spacer amino acid pairs ;
[0099] Step 2.9, Inputting into the ESM2 model based on the Transformer architecture generates... Context-aware matrix And then Perform average pooling to obtain ESM2 feature vector Thus obtain ESM2 feature matrix , where L is the length of the vector after ESM2 encoding and pooling.
[0100] Step 2.10: Obtain the multi-label category following the process in Step 2.9. of Labeled ESM2 feature vectors Thus obtain Labeled ESM2 feature matrix ;
[0101] Remove The tag information is then obtained by following step 2.9. Unlabeled ESM2 feature vectors Thus obtain Unlabeled ESM2 feature matrix .
[0102] Step 2.11: Obtain the descriptor using the join triplet method. triple matrix ;
[0103] Step 2.12, , , , , spliced as amino acid sequence characteristics The feature matrices of different encoding methods are concatenated to construct a high-dimensional feature space; , , spliced as Drug resistance gene sequence characteristics ;
[0104] Step 2.13: Use the multi-population particle swarm optimization feature selection algorithm to... Processing is performed to obtain Characteristics of drug resistance gene sequences after screening In multi-label classification, the model needs to predict multiple labels simultaneously, which means that the features of the input data may have different effects on each label. Therefore, selecting effective features can reduce computational complexity and avoid overfitting.
[0105] Step 3: Construct a drug resistance gene prediction model, such as... Figure 3 As shown, it includes: the XGBoost model and the first CNN-BiLSTM network, and respectively... and After processing, the optimal prediction score of drug resistance genes in the XGBoost model is obtained. The optimal prediction score of drug resistance genes from the first CNN-BiLSTM network And then and Weighted processing is performed to obtain the predicted score of drug resistance gene integration. ;
[0106] Step 3.1, The input is fed into the XGBoost model for training, and the drug resistance gene prediction score of the XGBoost model is calculated using equation (3). :
[0107] (3)
[0108] In equation (3), This represents the total number of decision trees in the XGBoost model. Indicates the first A decision tree, Indicates the first A decision tree pair A score vector for predicting drug resistance genes based on amino acid sequences.
[0109] Step 3.2: Construct the loss function of the XGBoost model using equation (4). :
[0110] (4)
[0111] In equation (4), express middle Corresponding drug resistance gene prediction score;
[0112] Step 3.3: Optimize the model using the XGBoost gradient boosting algorithm and calculate the loss function. The model parameters are updated until convergence is achieved, thus obtaining the optimal prediction score for drug resistance genes in the XGBoost model. .
[0113] Step 3.4, The input is processed in the first CNN-BiLSTM network to obtain the flattened local features. Then The input is processed in the BiLSTM layer of the first CNN-BiLSTM network for drug resistance gene prediction, thus obtaining the final hidden state of the BiLSTM. ; Then input it into the fully connected layer, and use equation (5) to... Drug resistance gene prediction scores converted to the first CNN-BiLSTM network :
[0114] (5)
[0115] In equation (5), This represents the weight matrix of the fully connected layer. For the bias term of the fully connected layer, This represents the Sigmoid activation function.
[0116] Step 3.5: Construct the loss function of the first CNN-BiLSTM network using equation (6). :
[0117] (6)
[0118] In equation (6), express middle The corresponding drug resistance gene prediction score.
[0119] Step 3.6: Optimize the first CNN-BiLSTM network using the Adam optimization algorithm and calculate the loss function. The network parameters are updated until convergence is achieved, thus obtaining the optimal prediction score for drug resistance genes from the first CNN-BiLSTM network. .
[0120] Step 3.7, for and A weighted integration is performed, and a comprehensive drug resistance gene prediction score is obtained using equation (7). :
[0121] (7)
[0122] In equation (7), and They are and The weights are determined using a grid search in this embodiment, with the parameter range for the grid search set to 0.01. To obtain the best experimental results.
[0123] if middle corresponding value Greater than the preset threshold , then it means The predicted result is a drug resistance gene; if it is smaller than [a certain value]... ,but Predicted as a non-drug-resistant gene; in this embodiment, the threshold Set the value to 0.8. If the predicted score of the amino acid sequence is greater than 0.8, it means that the amino acid sequence belongs to the drug resistance gene. If it is less than 0.8, it means that the amino acid sequence does not belong to the drug resistance gene.
[0124] Step 4: Construct a predictive model for drug resistance gene categories, such as... Figure 4 As shown, it includes: an ML-KNN model, a pre-trained CNN-BiLSTM network, and a fine-tuned CNN-BiLSTM network; the ML-KNN model is used to... After processing, the optimal prediction score for the resistance category of the ML-KNN model is obtained. Using a pre-trained CNN-BiLSTM network to... The process is performed to obtain a CNN-BiLSTM model with optimal parameters; the output layer of the CNN-BiLSTM model with optimal parameters is then replaced with... The classification layer is used to obtain a fine-tuned CNN-BiLSTM network, and then... After processing, the optimal prediction score for the resistance category of the fine-tuned CNN-BiLSTM network is obtained. And then and After weighting, a comprehensive prediction score for the resistance category is obtained. .
[0125] Step 4.1, The data is input into the ML-KNN model for processing, thereby obtaining the resistance category prediction score of the ML-KNN model. ;
[0126] Step 4.2, Defined as anchor sample, for Add noise to obtain Positive samples belonging to the same category Random selection and Samples not belonging to the same category are considered negative samples. ,Will The input is fed into a pre-trained CNN-BiLSTM network, and the loss function is constructed using equation (8). :
[0127] (8)
[0128] In equation (8), This refers to a pre-trained CNN-BiLSTM network. These are marginal hyperparameters used to control the minimum distance between positive and negative samples. Indicates Euclidean distance;
[0129] Step 4.3: Optimize the pre-trained CNN-BiLSTM network using the backpropagation algorithm and calculate the loss function. The network parameters are updated until convergence, thus obtaining the CNN-BiLSTM model with optimal parameters.
[0130] Step 4.4: Replace the output layer of the CNN-BiLSTM model with the optimal parameters. A classification layer is added to form a fine-tuned CNN-BiLSTM network; The input is processed by a fine-tuned CNN-BiLSTM network to obtain the resistance class prediction score of the fine-tuned CNN-BiLSTM network. ;
[0131] Step 4.5: Construct the loss function of the fine-tuned CNN-BiLSTM network using equation (9). :
[0132] (9)
[0133] In equation (9), Indicates the first Drug resistance gene sequence exist The Middle Resistance category prediction scores for each category;
[0134] Step 4.6: Optimize the fine-tuned CNN-BiLSTM network using the backpropagation algorithm and calculate the loss function. The network parameters are updated until convergence is achieved, thus obtaining the optimal parameters for the drug resistance gene resistance category classification model. Finally, the resistance category prediction score of the fine-tuned CNN-BiLSTM network is obtained. .
[0135] Step 4.7, for and A weighted integration is performed, and a comprehensive resistance category prediction score is obtained using equation (10). :
[0136] (10)
[0137] In equation (10), and They are and The weights are determined using a grid search in this embodiment, with the parameter range for the grid search set to 0.01. To obtain the best experimental results.
[0138] like middle The Value Greater than the preset threshold , then it means Belongs to the Number of categories; if less than , then it means Not belonging to the first Categories; in this embodiment, the threshold Set to 0.8, if the drug resistance gene sequence is... If the predicted score is greater than 0.8, it indicates that the drug resistance gene sequence belongs to the [number missing]th [sequence missing]. If the value is less than 0.8, it indicates that the drug resistance gene sequence does not belong to the first category. There are several categories.
[0139] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.
[0140] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.
Claims
1. A method for predicting drug resistance genes and resistance categories based on sequence features and protein language models, characterized in that, The procedure is as follows: Step 1: Construct the amino acid dataset U1 and the drug resistance gene dataset U2; Step 2, Calculation mutual information Conditional mutual information Fourier power spectrum characteristics Dipeptide composition Composition of spacer amino acid pairs ESM2 matrix , mutual information Composition of spacer amino acid pairs Triple matrix Labeled ESM2 feature matrix Unlabeled ESM2 feature matrix ; and will , , , , spliced as amino acid sequence characteristics ;Will , , spliced as Drug resistance gene sequence characteristics Then, the multi-population particle swarm optimization feature selection algorithm is used to... Processing is performed to obtain Characteristics of drug resistance gene sequences after screening ; Step 3: Construct a drug resistance gene prediction model, including: an XGBoost model and a first CNN-BiLSTM network, and respectively... and After processing, the optimal prediction score of drug resistance genes in the XGBoost model is obtained. The optimal prediction score of drug resistance genes from the first CNN-BiLSTM network And then and Weighted processing is performed to obtain the predicted score of drug resistance gene integration. ; if The Middle amino acid sequence corresponding value Greater than the preset threshold , then it means The predicted result is a drug resistance gene; if it is smaller than [a certain value]... ,but Predicted to be a non-drug-resistant gene; Step 4: Construct a drug resistance gene resistance category prediction model, including: an ML-KNN model, a pre-trained CNN-BiLSTM network, and a fine-tuned CNN-BiLSTM network; and use the ML-KNN model to predict drug resistance gene categories. After processing, the optimal prediction score for the resistance category of the ML-KNN model is obtained. Using a pre-trained CNN-BiLSTM network to... The process is performed to obtain a CNN-BiLSTM model with optimal parameters; the output layer of the CNN-BiLSTM model with optimal parameters is then replaced with... The classification layer is used to obtain a fine-tuned CNN-BiLSTM network, and then... After processing, the optimal prediction score for the resistance category of the fine-tuned CNN-BiLSTM network is obtained. And then and After weighting, a comprehensive prediction score for the resistance category is obtained. ; like The Middle The amino acid sequence of a drug resistance gene The Value Greater than the preset threshold , then it means Belongs to the Number of categories; if less than , then it means Not belonging to the first There are several categories.
2. The method for predicting drug resistance genes and resistance categories based on sequence features and protein language models according to claim 1, characterized in that, Step 1 is performed as follows: Step 1.1: Define a finite set of amino acids. ,in Indicates the first One amino acid, Indicates the first There are 10 amino acids, where H represents the total number of amino acids, and , Amino acids in food can be classified according to their physicochemical properties. Functional categories; Step 1.2: Obtain the amino acid dataset ,in, Indicates the first A single amino acid sequence, Indicates the number of amino acid sequences, and The sequence length is ; make The tag is ,like If the amino acid sequence is a drug resistance gene, then let... ,like If the amino acid sequence is a non-drug resistance gene, then let ; Step 1.3, Take middle The amino acid sequences constitute the drug resistance gene dataset. ,in, Indicates the first The amino acid sequence of a drug resistance gene. The number of amino acid sequences representing the resistance gene; definition Multi-category labels are denoted as , Indicates the first Each category of labels, and , Indicates the total number of categories, when When, it means Not belonging to the first Each category, when When, it means Belongs to the There are several categories.
3. The method for predicting drug resistance genes and resistance categories based on sequence features and protein language models according to claim 2, characterized in that, Step 2 is performed as follows: Step 2.1, The Middle amino acids and the amino acids The resulting biamino acid pair is denoted as And calculate using equation (1) middle Includes about mutual information Thus obtain The total amount of mutual information : (1) In equation (1), yes middle Frequency of occurrence and They are middle and Frequency of occurrence; Step 2.2: Calculate according to the process in Step 2.
1. The Middle amino acids Includes information about the first amino acids mutual information Thus obtain The total amount of mutual information ; Step 2.3, and spliced together mutual information vector Thus obtain mutual information ;Will and spliced together mutual information vector Thus obtain mutual information ; Step 2.4, In order from the first amino acids , No. amino acids , No. amino acids The three amino acids that make up the group are denoted as , by the amino acids and the amino acids The resulting biamino acid pair is denoted as , by the amino acids and the amino acids The biamino acid pair formed is denoted as Calculate using equation (2) Under the conditions, and Mutual information between Thus obtain The sum of conditional mutual information : (2) In equation (2), yes middle Frequency of occurrence and They are middle and Frequency of occurrence yes middle Frequency of occurrence; Step 2.5, and spliced together Conditional mutual information vector Thus obtain Conditional mutual information ; Step 2.6: Calculate based on the cumulative Fourier power spectrum. Moment vector and central moment vector and will and spliced together Fourier power spectrum characteristics Thus obtain Fourier power spectrum characteristics ; Step 2.7: Obtain using dipeptide composition method dipeptide composition matrix : Step 2.8: Obtain the compositional feature encoding method using spacer amino acid pairs. Composition of spacer amino acid pairs and Composition of spacer amino acid pairs ; Step 2.9, Inputting into the ESM2 model based on the Transformer architecture generates... Context-aware matrix And then Perform average pooling to obtain ESM2 feature vector Thus obtain ESM2 feature matrix Where L is the length of the vector after ESM2 encoding and pooling; Step 2.10: Obtain the multi-label category following the process in Step 2.
9. of Labeled ESM2 feature vectors Thus obtain Labeled ESM2 feature matrix ; Remove The tag information is then obtained by following step 2.
9. Unlabeled ESM2 feature vectors Thus obtain Unlabeled ESM2 feature matrix ; Step 2.11: Obtain the descriptor using the join triplet method. triple matrix .
4. The method for predicting drug resistance genes and resistance categories based on sequence features and protein language models according to claim 3, characterized in that, Step 3 is performed as follows: Step 3.1, The input is fed into the XGBoost model for training, and the drug resistance gene prediction score of the XGBoost model is calculated using equation (3). : (3) In equation (3), This represents the total number of decision trees in the XGBoost model. Indicates the sequence number of the decision tree. Indicates the first A decision tree pair The predicted score vector of drug resistance genes; Step 3.2: Construct the loss function of the XGBoost model using equation (4). : (4) In equation (4), express middle Corresponding drug resistance gene prediction score; Step 3.3: Optimize the model using the XGBoost gradient boosting algorithm and calculate the loss function. The model parameters are updated until convergence is achieved, thus obtaining the optimal prediction score for drug resistance genes in the XGBoost model. ; Step 3.4, The input is processed in the first CNN-BiLSTM network to obtain the flattened local features. Then The input is processed in the BiLSTM layer of the first CNN-BiLSTM network for drug resistance gene prediction, thus obtaining the final hidden state of the BiLSTM. ; Then input it into the fully connected layer, and use equation (5) to... Drug resistance gene prediction scores converted to the first CNN-BiLSTM network : (5) In equation (5), This represents the weight matrix of the fully connected layer. For the bias term of the fully connected layer, This represents the Sigmoid activation function; Step 3.5: Construct the loss function of the first CNN-BiLSTM network using equation (6). : (6) In equation (6), express middle Corresponding drug resistance gene prediction score; Step 3.6: Optimize the first CNN-BiLSTM network using the Adam optimization algorithm and calculate the loss function. The network parameters are updated until convergence is achieved, thus obtaining the optimal prediction score for drug resistance genes from the first CNN-BiLSTM network. ; Step 3.7, for and A weighted integration is performed, and a comprehensive drug resistance gene prediction score is obtained using equation (7). : (7) In equation (7), and They are and The weight.
5. The method for predicting drug resistance genes and resistance categories based on sequence features and protein language models according to claim 4, characterized in that, Step 4 is performed as follows: Step 4.1, The data is input into the ML-KNN model for processing, thereby obtaining the resistance category prediction score of the ML-KNN model. ; Step 4.2, Defined as anchor sample, for Add noise to obtain Positive samples belonging to the same category Random selection and Samples not belonging to the same category are considered negative samples. ,Will The input is fed into a pre-trained CNN-BiLSTM network, and the loss function is constructed using equation (8). : (8) In equation (8), This refers to a pre-trained CNN-BiLSTM network. These are marginal hyperparameters used to control the minimum distance between positive and negative samples. Indicates Euclidean distance; Step 4.3: Optimize the pre-trained CNN-BiLSTM network using the backpropagation algorithm and calculate the loss function. The network parameters are updated until convergence is achieved, thus obtaining the CNN-BiLSTM model with optimal parameters. Step 4.4: Replace the output layer of the CNN-BiLSTM model with the optimal parameters. A classification layer is added, thus forming a fine-tuned CNN-BiLSTM network; The input is processed by a fine-tuned CNN-BiLSTM network to obtain the resistance class prediction score of the fine-tuned CNN-BiLSTM network. ; Step 4.5: Construct the loss function of the fine-tuned CNN-BiLSTM network using equation (9). : (9) In equation (9), Indicates the first Drug resistance gene sequence exist The Middle Resistance category prediction scores for each category; Step 4.6: Optimize the fine-tuned CNN-BiLSTM network using the backpropagation algorithm and calculate the loss function. The network parameters are updated until convergence is achieved, thus obtaining the optimal parameters for the drug resistance gene resistance category classification model. Finally, the resistance category prediction score of the fine-tuned CNN-BiLSTM network is obtained. ; Step 4.7, for and A weighted integration is performed, and a comprehensive resistance category prediction score is obtained using equation (10). : (10) In equation (10), and They are and The weight.
6. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing the drug resistance gene and resistance category prediction method according to any one of claims 1-5, and the processor is configured to execute the program stored in the memory.
7. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is run by the processor, it performs the steps of the method for predicting drug resistance genes and resistance categories as described in any one of claims 1-5.
Citation Information
Patent Citations
Hybrid prediction method for virulence factor and antibiotic resistance gene
CN115171792A
Antibacterial peptide recognition method and application of antibacterial peptide recognition method in inhibition of multi-drug-resistant bacteria
CN118136121A