Method and system for predicting antibacterial activity of antibacterial peptide sequence based on LSTM (Long Short Term Memory) model
Through the antibacterial peptide sequence prediction method based on the LSTM model, the problem of insufficient specificity of antibacterial activity prediction for specific bacterial genus is solved in the prior art, and high-accurate antibacterial activity prediction is achieved.
Patent Information
- Application Number
- CN202411917560.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-30
AI Technical Summary
The existing antibacterial activity prediction methods are insufficient to target specificity of specific bacterial genus, making it difficult to effectively capture complex dependencies of protein sequences, and the prediction accuracy is not high, making it difficult to meet the practical application needs.
Using the antibacterial activity prediction method based on the LSTM model, the training sample training set is constructed, and the high-dimensional embedding features are extracted using the pre-trained protein language model, and the initial LSTM model is trained to obtain the trained LSTM model to predict antibacterial activity.
It significantly improves the accuracy of antibacterial activity prediction, can effectively capture complex dependencies of protein sequences, improves the prediction accuracy of the model, and meets the practical application needs.
Smart Images

Figure CN120072028A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of bioinformatics and artificial intelligence, and particularly to a method and system for predicting the antibacterial activity of antibacterial peptide sequences based on an LSTM model. Background Art
[0002] Antibacterial peptides are a class of natural or synthetic peptide molecules that can effectively kill or inhibit pathogenic microorganisms and have broad application prospects in anti-infection treatment.
[0003] However, the existing antibacterial activity prediction methods have the following problems: the activity prediction model has insufficient target specificity for specific bacterial genera; there is a lack of effective capture of the complex dependence relationships in protein sequences during data processing and feature extraction; the model prediction accuracy is not high and it is difficult to meet the actual application requirements. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, system, computer device, and storage medium for predicting the antibacterial activity of antibacterial peptide sequences based on an LSTM model in view of the above technical problems.
[0005] In a first aspect, an embodiment of the present invention proposes a method for predicting the antibacterial activity of antibacterial peptide sequences based on an LSTM model, the method comprising:
[0006] Constructing a training sample training set, the sample training set including a positive sample set and a negative sample set, the positive sample set including antibacterial peptide sequences of a target bacterial genus after preprocessing, and the negative sample set including non-antibacterial peptide sequences after preprocessing;
[0007] Using a pre-trained protein language model to extract first high-dimensional embedding features of the sample training set;
[0008] Using the antibacterial activity of the sample training set as a label, training an initial LSTM model with the first high-dimensional embedding features to obtain a trained LSTM model;
[0009] Using the protein language model to extract second high-dimensional embedding features of the antibacterial peptide sequence to be predicted, inputting the second high-dimensional embedding features into the trained LSTM model, and outputting the antibacterial activity of the antibacterial peptide sequence to be predicted.
[0010] In some embodiments, the protein language model models the context information of the sample training set, captures the evolutionary information, the interaction information between amino acids, and the sequence features of the antibacterial peptide sequence, generates the first high-dimensional embedding features, and the dimension of the first high-dimensional embedding features is equal to the output dimension of the hidden state matrix of the protein language model.
[0011] In some embodiments, the trained LSTM model includes two LSTM layers, a Dropout layer, and an output layer constructed in sequence; the hidden state dimension of the LSTM layer is set to 256; the output layer is used to convert the hidden state dimension of the LSTM layer into a log MIC value for characterizing antibacterial activity.
[0012] In some embodiments, using the antibacterial activity of the sample training set as a label, and using the first high-dimensional embedding features to train an initial LSTM model, the obtained trained LSTM model includes:
[0013] Dividing the sample training set into a training set, a validation set, and a test set according to a ratio;
[0014] Using the antibacterial activity of the sample training set as a label, and using the training set, the validation set, and the test set, and using the first high-dimensional embedding features to train the initial LSTM model to obtain a trained LSTM model.
[0015] In some embodiments, using the antibacterial activity of the sample training set as a label, and using the first high-dimensional embedding features to train an initial LSTM model, the obtained trained LSTM model further includes:
[0016] During the process of training the initial LSTM model, the time backpropagation algorithm is used to obtain gradients, the mean squared error loss function is used to calculate errors, and the optimizer is used to optimize the model parameters based on the gradients and the errors.
[0017] In some embodiments, the preprocessing includes performing deduplication, filtering, and standardization processing on the antimicrobial peptide data in sequence.
[0018] In some embodiments, for the positive sample set, the standardization processing includes:
[0019] Deleting entries in the antimicrobial peptide sequence that contain ambiguous amino acids, less than 5 amino acids, or more than 65 amino acids;
[0020] For the same antimicrobial peptide sequence with multiple MIC values, calculating its average value;
[0021] Unifying the MIC values into micromoles and then performing a log10 conversion.
[0022] In a second aspect, an embodiment of the present invention proposes an antibacterial activity prediction system for antimicrobial peptide sequences based on an LSTM model, and the system includes:
[0023] A data processing module for constructing a training sample set, where the sample training set includes a positive sample set and a negative sample set. The positive sample set includes antibacterial peptide sequences of a target genus of bacteria after preprocessing, and the negative sample set includes non-antibacterial peptide sequences after preprocessing;
[0024] A feature extraction module for extracting first high-dimensional embedding features of the sample training set by using a pre-trained protein language model;
[0025] A model training module for training an initial LSTM model with the antibacterial activity of the sample training set as a label by using the first high-dimensional embedding features to obtain a trained LSTM model;
[0026] A prediction module for extracting second high-dimensional embedding features of the antibacterial peptide sequence to be predicted by using the protein language model, inputting the second high-dimensional embedding features into the trained LSTM model, and outputting the antibacterial activity of the antibacterial peptide sequence to be predicted.
[0027] In a third aspect, an embodiment of the present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the steps described in the first aspect.
[0028] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the processor executes the computer program, the steps described in the first aspect are implemented.
[0029] Compared with the prior art, the above method, system, computer device, and storage medium train a targeted LSTM model through the antibacterial peptide sequences of the target genus of bacteria, significantly improving the accuracy of antibacterial activity prediction; using a pre-trained protein language model to extract the first high-dimensional embedding features of the sample training set can effectively capture the complex dependence relationships of protein sequences. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a schematic flowchart of a method for predicting the antibacterial activity of an antibacterial peptide sequence based on an LSTM model in an embodiment of the present application;
[0031] Figure 2 It is a schematic structural diagram of an LSTM model in an embodiment of the present application;
[0032] Figure 3 It is a schematic flowchart of a method for training an LSTM model in an embodiment of the present application;
[0033] Figure 4 It is a schematic evaluation diagram of an LSTM model trained for an antibacterial peptide sequence against Escherichia coli;
[0034] Figure 5 Schematic diagram for evaluating the LSTM model trained with antibacterial peptide sequences against Staphylococcus aureus;
[0035] Figure 6 Schematic diagram of module connection of the antibacterial activity prediction system for antibacterial peptide sequences based on the LSTM model in an embodiment of the present application;
[0036] Figure 7 Schematic diagram of the structure of a computer device in an embodiment of the present application. Detailed implementation manners
[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, the present invention can also be applied to other similar scenarios based on these drawings. Unless obvious from the language context or otherwise stated, the same reference numerals in the figures represent the same structure or operation.
[0038] As shown in the present invention and the claims, unless the context clearly indicates an exceptional situation, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0039] Although the present invention makes various references to certain modules in the system according to the embodiments of the present invention, however, any number of different modules can be used and run on a computing device and / or a processor. The modules are only illustrative, and different aspects of the system and method can use different modules.
[0040] It should be understood that when a unit or module is described as "connected" or "coupled" to other units, modules, or blocks, it may mean direct connection or coupling, or communication with other units, modules, or blocks, or there may be intermediate units, modules, or blocks, unless the context clearly indicates otherwise. The term "and / or" used herein may include any and all combinations of one or more of the related listed items.
[0041] As Figure 1 shown, the embodiments of the present invention provide an antibacterial activity prediction method for antibacterial peptide sequences based on the LSTM model, including the following steps:
[0042] S102: Construct a training sample training set, where the sample training set includes a positive sample set and a negative sample set. The positive sample set includes the antibacterial peptide sequences of the target genus after preprocessing, and the negative sample set includes the non-antibacterial peptide sequences after preprocessing.
[0043] Retrieve antibacterial peptide sequences from six public antibacterial peptide (AMP) databases (APD, DADP, DBAASP, DRAMP, YADAMP, and dbAMP). These databases contain publicly available antibacterial peptide sequences and their corresponding antibacterial activities.
[0044] Retrieve non-antibacterial peptide sequences from the Uniport database.
[0045] S104: Use a pre-trained protein language model to extract the first high-dimensional embedding features of the sample training set.
[0046] The pre-trained protein language model is ESM2-t36-3B. This model is based on the Transformer architecture and can capture the complex structural dependencies of antibacterial peptide sequences.
[0047] S106: Use the antibacterial activities of the sample training set as labels, and use the first high-dimensional embedding features to train an initial LSTM model to obtain a trained LSTM model.
[0048] The Long Short-Term Memory (LSTM) network model is a special type of recurrent neural network. It solves the problem of gradient vanishing in ordinary neural networks when dealing with long sequence data by introducing memory units. The LSTM model can capture dependencies in long sequences and is a very suitable deep learning model for processing time series.
[0049] S108: Use the protein language model to extract the second high-dimensional embedding features of the antibacterial peptide sequence to be predicted, input the second high-dimensional embedding features into the trained LSTM model, and output the antibacterial activity of the antibacterial peptide sequence to be predicted.
[0050] Based on the above steps S102 - S108, training a targeted LSTM model with the antibacterial peptide sequences of the target genus significantly improves the accuracy of antibacterial activity prediction; using a pre-trained protein language model to extract the first high-dimensional embedding features of the sample training set can effectively capture the complex dependencies of protein sequences.
[0051] In step S102, the preprocessing includes performing deduplication, filtering, and normalization on the antibacterial peptide data in sequence.
[0052] Deduplicate the antibacterial peptide data to ensure the uniqueness of the sequences.
[0053] Filter and delete sequences containing incomplete information and ambiguous amino acids (such as "X").
[0054] Among them, for the positive sample set, the normalization process includes: deleting entries in the antimicrobial peptide sequence that contain ambiguous amino acids, less than 5 amino acids or more than 65 amino acids; for the same antimicrobial peptide sequence with multiple MIC values, calculating its average value; unifying the MIC values into micromoles and then performing a log10 transformation.
[0055] For the negative sample set, the normalization process includes: marking the log MIC value of the non-antimicrobial peptide sequence as 4.
[0056] By performing deduplication, filtering, and normalization processing on the data, the prediction accuracy of the LSTM model can be improved.
[0057] Taking Escherichia coli and Staphylococcus aureus as examples of the target bacterial genera, the sample construction process of the antimicrobial peptide sequence is as follows:
[0058] A. For Escherichia coli, 7,100 antimicrobial peptide sequences were collected;
[0059] B. For Staphylococcus aureus, 6,482 antimicrobial peptide sequences were collected;
[0060] C. For the same sequence with multiple MIC values, calculate the arithmetic mean of its minimum inhibitory concentration value (minimal inhibit concentration, MIC), unify the MIC values into micromoles and then perform a log10 transformation to obtain the positive sample set;
[0061] D. To maintain the balance between positive and negative samples, 7,193 non-antimicrobial peptide sequences were randomly selected, and the log MIC value was marked as 4 to obtain the negative sample set.
[0062] In step S104, using the protein language model to model the context information of the antimicrobial peptide sequence, capture the evolutionary information, the interaction information between amino acids, and the sequence features of the antimicrobial peptide sequence, and generate the first high-dimensional embedding feature, where the dimension of the first high-dimensional embedding feature is equal to the output dimension of the hidden state matrix of the protein language model.
[0063] Using the multi-head self-attention mechanism in the protein language model to calculate the dependencies between positions in the sequence, improve the feature expression ability, and make structurally related sequences closer in the high-dimensional embedding space.
[0064] In step S106, as Figure 2As shown, the trained LSTM model includes two LSTM layers, a Dropout layer, and an output layer constructed in sequence; the hidden state dimension of the LSTM layer is set to 256; the output layer is used to convert the hidden state dimension of the LSTM layer into a log MIC value for characterizing antibacterial activity. Adding a Dropout layer (Dropout rate is 0.7) after the LSTM layer can reduce the risk of model overfitting.
[0065] In step S106, as Figure 3 shown, using the antibacterial activity of the sample training set as a label, and using the first high-dimensional embedding feature to train the initial LSTM model, the trained LSTM model obtained includes:
[0066] S302: Divide the sample training set into a training set, a validation set, and a test set according to a ratio;
[0067] S304: Using the antibacterial activity of the sample training set as a label, and using the training set, the validation set, and the test set, and using the first high-dimensional embedding feature to train the initial LSTM model to obtain a trained LSTM model.
[0068] In an exemplary embodiment, the antibacterial peptide sequences are divided into a training set, a validation set, and a test set according to a ratio of 72:18:10.
[0069] The training set is used for weight optimization of the LSTM model; the validation set is used to monitor the performance of the model on non-training data; the test set is used to independently evaluate the generalization ability of the model.
[0070] In some embodiments, using the antibacterial activity of the sample training set as a label, and using the first high-dimensional embedding feature to train the initial LSTM model, the trained LSTM model further includes: during the process of training the initial LSTM model, the time backpropagation algorithm is used to obtain gradients, the mean squared error loss function is used to calculate errors, and the optimizer is used to optimize the model parameters based on the gradients and the errors.
[0071] In the specific training process, the L2 loss function is used to calculate the error between the predicted value and the true log MIC value, the time backpropagation algorithm is used to calculate gradients, and gradient clipping is used to avoid the problem of gradient explosion, and the Adam optimizer is used to update parameters based on the gradients and the errors.
[0072] In step S108, the protein language model is used to extract the second high-dimensional embedding feature of the antibacterial peptide sequence to be predicted, and the second high-dimensional embedding feature is input into the trained LSTM model to output the antibacterial activity of the antibacterial peptide sequence to be predicted.
[0073] According to the predicted log MIC value, the antibacterial activity of the antimicrobial peptide sequence in the target bacterial genus (such as Escherichia coli or Staphylococcus aureus) is evaluated. The lower the log MIC value, the stronger the antibacterial activity.
[0074] In some embodiments, the antibacterial activity of a batch of antimicrobial peptide sequences can also be predicted and sorted according to the log MIC value.
[0075] In some embodiments, the mean squared error (MSE) or the coefficient of determination (R 2 ) can also be used to evaluate the performance of the trained LSTM model.
[0076] In an exemplary embodiment, the target bacterial genus is Escherichia coli and Staphylococcus aureus, and two LSTM models are trained. Figure 4 It is a schematic diagram for evaluating the LSTM model trained for the antimicrobial peptide sequence against Escherichia coli. The R 2 value of the LSTM model on the Escherichia coli validation set is 0.89. Figure 5 It is a schematic diagram for evaluating the LSTM model trained for the antimicrobial peptide sequence against Staphylococcus aureus. The R 2 value of the LSTM model on the Escherichia coli validation set is 0.86.
[0077] It should be understood that although the steps in the above flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the above flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0078] In one embodiment, as Figure 6 shown, the present invention provides an antibacterial activity prediction system for antimicrobial peptide sequences based on an LSTM model. The system includes:
[0079] A data processing module 602, configured to construct a training sample training set, where the sample training set includes a positive sample set and a negative sample set. The positive sample set includes antimicrobial peptide sequences of the target bacterial genus after preprocessing, and the negative sample set includes non-antimicrobial peptide sequences after preprocessing;
[0080] A feature extraction module 604, configured to extract first high-dimensional embedding features of the sample training set by using a pre-trained protein language model;
[0081] A model training module 606, configured to use the antibacterial activity of the sample training set as a label, and train an initial LSTM model by using the first high-dimensional embedding features to obtain a trained LSTM model;
[0082] A prediction module 608, configured to use the protein language model to extract the second high-dimensional embedding features of the antibacterial peptide sequence to be predicted, input the second high-dimensional embedding features into the trained LSTM model, and output the antibacterial activity of the antibacterial peptide sequence to be predicted.
[0083] Training a targeted LSTM model with the antibacterial peptide sequences of the target genus significantly improves the accuracy of antibacterial activity prediction; using the pre-trained protein language model to extract the first high-dimensional embedding features of the sample training set can effectively capture the complex dependence relationships of protein sequences.
[0084] In some embodiments, the protein language model captures the evolutionary information, the interaction information between amino acids, and the sequence features of the antibacterial peptide sequence by modeling the context information of the sample training set, and generates the first high-dimensional embedding features, and the dimension of the first high-dimensional embedding features is equal to the output dimension of the hidden state matrix of the protein language model.
[0085] In some embodiments, the trained LSTM model includes two LSTM layers, a Dropout layer, and an output layer constructed in sequence; the hidden state dimension of the LSTM layer is set to 256; the output layer is configured to convert the hidden state dimension of the LSTM layer into a log MIC value for characterizing antibacterial activity.
[0086] In some embodiments, the method of using the antibacterial activity of the sample training set as a label and training an initial LSTM model by using the first high-dimensional embedding features to obtain a trained LSTM model includes:
[0087] Dividing the sample training set into a training set, a validation set, and a test set according to a ratio;
[0088] Using the antibacterial activity of the sample training set as a label, and using the training set, the validation set, and the test set, and training an initial LSTM model by using the first high-dimensional embedding features to obtain a trained LSTM model.
[0089] In some embodiments, the method of using the antibacterial activity of the sample training set as a label and training an initial LSTM model by using the first high-dimensional embedding features to obtain a trained LSTM model further includes:
[0090] During the training of the initial LSTM model, the time backpropagation algorithm is used to obtain gradients, the mean squared error loss function is used to calculate the error, and the optimizer is used to optimize the model parameters based on the gradients and the error.
[0091] In some embodiments, the preprocessing includes deduplication, filtering, and normalization of the antimicrobial peptide data in sequence.
[0092] In some embodiments, for the positive sample set, the normalization processing includes:
[0093] Deleting entries in the antimicrobial peptide sequence that contain ambiguous amino acids, less than 5 amino acids, or more than 65 amino acids;
[0094] For the same antimicrobial peptide sequence with multiple MIC values, calculating its average value;
[0095] Unifying the MIC values into micromoles and then performing a log10 transformation.
[0096] For the specific limitations of the antimicrobial activity prediction system for antimicrobial peptide sequences based on the LSTM model, reference can be made to the limitations on the tool automatic call method in the above text, which will not be elaborated here. Each module in the above antimicrobial activity prediction system for antimicrobial peptide sequences based on the LSTM model can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.
[0097] In one embodiment, the embodiment of the present invention provides a computer device, which can be a server, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store action detection data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes the steps in any one of the above embodiments of the antimicrobial activity prediction method for antimicrobial peptide sequences based on the LSTM model.
[0098] Those skilled in the art can understand, Figure 7The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0099] In one embodiment, the embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in any of the above-mentioned embodiments of the antibacterial activity prediction method for antibacterial peptide sequences based on the LSTM model are implemented.
[0100] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0101] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0102] The above embodiments only represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as limitations on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.
Claims
1. A method for predicting the antibacterial activity of antimicrobial peptide sequences based on an LSTM model, characterized in that: The method comprises: Constructing a training sample training set, wherein the sample training set includes a positive sample set and a negative sample set, wherein the positive sample set includes pre-processed antimicrobial peptide sequences of the target bacteria, and the negative sample set includes pre-processed non-antimicrobial peptide sequences; Extracting a first high-dimensional embedding feature of the sample training set using a pre-trained protein language model; Using the antibacterial activity of the sample training set as a label, and using the first high-dimensional embedding feature to train the initial LSTM model to obtain a trained LSTM model; The protein language model is used to extract a second high-dimensional embedding feature of the antimicrobial peptide sequence to be predicted, the second high-dimensional embedding feature is input into the trained LSTM model, and the antimicrobial activity of the antimicrobial peptide sequence to be predicted is output.
2. The method according to claim 1, characterized in that The protein language model captures the evolutionary information of the antimicrobial peptide sequence, the interaction information between amino acids and the sequence characteristics by modeling the context information of the sample training set, and generates the first high-dimensional embedding feature, wherein the dimension of the first high-dimensional embedding feature is equal to the output dimension of the hidden state matrix of the protein language model.
3. The method according to claim 1, characterized in that The trained LSTM model includes two LSTM layers, a Dropout layer and an output layer constructed in sequence; the hidden state dimension of the LSTM layer is set to 256; and the output layer is used to convert the hidden state dimension of the LSTM layer into a log MIC value for characterizing antibacterial activity.
4. The method according to claim 3, characterized in that The method of using the antibacterial activity of the sample training set as a label and using the first high-dimensional embedding feature to train the initial LSTM model to obtain a trained LSTM model includes: Dividing the sample training set into a training set, a validation set and a test set according to proportion; The antibacterial activity of the sample training set is used as a label, and the training set, the validation set and the test set are used to train the initial LSTM model using the first high-dimensional embedding feature to obtain a trained LSTM model.
5. The method according to claim 4, characterized in that The method of using the antibacterial activity of the sample training set as a label and using the first high-dimensional embedding feature to train the initial LSTM model to obtain a trained LSTM model further includes: During the training of the initial LSTM model, the time back propagation algorithm is used to obtain the gradient, the square error loss function is used to calculate the error, and the optimizer is used to optimize the model parameters based on the gradient and the error.
6. The method according to claim 1, characterized in that The preprocessing includes sequentially performing deduplication, filtering and standardization on the antimicrobial peptide data.
7. The method according to claim 1, characterized in that For the positive sample set, the standardization process includes: Delete the entries containing ambiguous amino acids, less than 5 amino acids, or more than 65 amino acids in the antimicrobial peptide sequence; For the same antimicrobial peptide sequence with multiple MIC values, the average value was calculated; MIC values were converted to micromolar and then log10 transformed.
8. A system for predicting the antimicrobial activity of antimicrobial peptide sequences based on an LSTM model, characterized in that: The system comprises: A data processing module, used to construct a training sample training set, wherein the sample training set includes a positive sample set and a negative sample set, wherein the positive sample set includes the pre-processed antimicrobial peptide sequences of the target bacteria, and the negative sample set includes the pre-processed non-antimicrobial peptide sequences; A feature extraction module, used for extracting a first high-dimensional embedding feature of the sample training set using a pre-trained protein language model; A model training module, used to train the initial LSTM model using the antibacterial activity of the sample training set as a label and the first high-dimensional embedding feature to obtain a trained LSTM model; The prediction module is used to extract the second high-dimensional embedding feature of the antimicrobial peptide sequence to be predicted by using the protein language model, input the second high-dimensional embedding feature into the trained LSTM model, and output the antimicrobial activity of the antimicrobial peptide sequence to be predicted.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Antibacterial peptide prediction method and device based on protein pre-training representation learning
CN112614538A
Antibacterial peptide prediction method based on BERT feature coding technology and deep learning combination model
CN117292749A
Model training method, antibacterial peptide prediction method and system
CN117875444A
Antibacterial peptide recognition and directed evolution method based on deep learning
CN118298907A