Peptide sequence list generation method, apparatus, device, and product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEKING UNIV SHENZHEN GRADUATE SCHOOL
- Filing Date
- 2026-07-13
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]本申请的主要目的在于提供一种肽序列列表生成方法、装置、设备以及产品,旨在解决现有抗菌肽预测与生成相互独立,导致流程繁琐效率低,难以快速得到高活性抗菌肽的技术问题
本申请实施例提出的一种肽序列列表生成方法、装置、设备以及产品,通过预先构建的抗菌肽预测核心模型对宏基因组数据进行预测,得到天然抗菌肽候选集合;通过抗菌肽生成模型进行AI肽序列生成,得到AI候选肽集合;对所述天然抗菌肽候选集合以及AI候选肽集合进行分层可解释标注,得到标注肽集合,并通过标注肽集合生成肽序列分类列表。由此,一方面从宏基因组数据挖掘天然抗菌肽候选,一方面利用模型生成全新肽序列,同时对两类序列统一开展分层标注分析,将预测、生成以及评价环节整合为一体,省去多模块拆分流转步骤,精简流程、提升效率,快速筛选出优质候选肽,解决了现有抗菌肽预测与生成相互独立,导致流程繁琐效率低,难以快速得到高活性抗菌肽的问题,提高了肽序列列表生成的效率。
Smart Images

Figure CN122531482A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biological sequence analysis technology, and in particular to a method, apparatus, device, and product for generating peptide sequence lists. Background Technology
[0002] Antimicrobial peptides (AMPs), a class of short peptides, typically no more than 100 amino acids long, are either natural or synthetic. Due to their broad-spectrum killing activity against pathogenic microorganisms such as bacteria, fungi, and viruses, they have been widely regarded as important candidate substances for combating drug-resistant bacteria and supplementing traditional small-molecule antibiotics. In addition to wet experiments, computational screening and de novo design of antimicrobial peptides have become the focus of research in this field. Currently, three main technical routes have been formed: one is the prediction method based on artificial features and shallow classifiers, which uses manually extracted sequence features as input to complete the prediction; the second is the end-to-end prediction method based on deep learning, which uses pre-trained protein language models to achieve high-precision prediction; and the third is the de novo design method based on generative models, which uses training sequences to generate models to sample new candidate peptide sequences.
[0003] However, existing technologies have core flaws. The three current technical routes are independent of each other, and the generation model can only complete the sampling of candidate sequences. It cannot directly evaluate the antibacterial activity of the sequences. Subsequently, it is still necessary to rely on independent prediction models to perform secondary screening of the sampling results. This results in redundancy and low efficiency in the entire antimicrobial peptide design process. Moreover, the effectiveness of the generated sequences is highly dependent on the accuracy of the prediction model, making it difficult to efficiently produce novel antimicrobial peptides with high antimicrobial activity. This cannot meet the current demand for rapid research and development of novel antimicrobial drugs under the current crisis of drug-resistant bacteria.
[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main objective of this application is to provide a method, apparatus, device, and product for generating peptide sequence lists, aiming to solve the technical problem that existing methods for predicting and generating antimicrobial peptides are independent, resulting in cumbersome processes, low efficiency, and difficulty in quickly obtaining highly active antimicrobial peptides.
[0006] To achieve the above objectives, this application proposes a method for generating a peptide sequence list, the method comprising: A set of natural antimicrobial peptide candidates is obtained by predicting metagenomic data using a pre-constructed core model for antimicrobial peptide prediction. AI peptide sequences were generated using an antimicrobial peptide generation model to obtain a set of AI candidate peptides. The candidate set of natural antimicrobial peptides and the candidate set of AI peptides are hierarchically interpretable labeled to obtain a labeled peptide set, and a peptide sequence classification list is generated from the labeled peptide set.
[0007] In one embodiment, before the step of predicting metagenomic data using a pre-constructed antimicrobial peptide prediction core model to obtain a candidate set of natural antimicrobial peptides, the method further includes: The sequences of the publicly available peptide database are summarized to obtain the first publicly available sequence. The substandard sequences of the first publicly available sequence are removed to obtain the positive sample set. Sequences were aggregated from a general protein resource database to obtain a second public sequence. The second public sequence was then removed using a keyword exclusion strategy to obtain a negative sample set. Stratified sampling is performed on the positive sample set and the negative sample set to obtain the training set and the test set; The core model for antimicrobial peptide prediction is obtained by training several sub-classifiers in the pre-trained antimicrobial peptide prediction core model using the training set and the test set.
[0008] In one embodiment, the step of predicting metagenomic data using a pre-constructed antimicrobial peptide prediction core model to obtain a candidate set of natural antimicrobial peptides includes: Extract metagenomic assembly fragments from the metagenomic data; Open reading frames (ORFs) are predicted from the assembled metagenomic fragments to obtain short open reading frames. The short open reading frames are deredundant to obtain non-redundant short peptide sequences; The antimicrobial peptide sequence is obtained by predicting the non-redundant short peptide sequence using the core model for antimicrobial peptide prediction. The antimicrobial peptide sequences are summarized and scored to obtain a candidate set of natural antimicrobial peptides.
[0009] In one embodiment, the step of generating an AI peptide sequence using an antimicrobial peptide generation model to obtain an AI candidate peptide set includes: The antimicrobial peptide samples were divided into an AI sequence training set and an AI sequence validation set. Based on the AI sequence training set, the antimicrobial peptide generation model is fine-tuned with all parameters through a preset optimizer and training parameters to obtain several sets of model parameters. The optimal model parameters are obtained by selecting parameters from the several sets of model parameters using the AI sequence validation set. Based on the optimal model parameters, the AI peptide sequence is obtained by controlled decoding and generation using the antimicrobial peptide generation model. The AI peptide sequences are screened to obtain a set of AI candidate peptides.
[0010] In one embodiment, the step of performing hierarchical interpretable annotation on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain an annotated peptide set includes: Consistency analysis was performed on several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set using the publicly available antimicrobial peptide classifier and the core model for antimicrobial peptide prediction, and the consistency analysis results were obtained. The compositional distribution of the aforementioned peptide sequences was analyzed to obtain the sequence diversity analysis results; The sequence differences between the aforementioned peptide sequences and the publicly disclosed antimicrobial peptide sequences, as well as the differences in the ESM2 high-dimensional latent vector space, were analyzed to obtain the sequence novelty analysis results. The degree of matching between the aforementioned peptide sequences and the distribution of real antimicrobial peptides was analyzed to obtain the distribution distance analysis results; Based on the consistency analysis results, sequence diversity analysis results, sequence novelty analysis results, and distribution distance analysis results, the natural antimicrobial peptide candidate set and the AI candidate peptide set are subjected to multidimensional quantitative annotation to obtain the annotated peptide set.
[0011] In one embodiment, the step of performing consistency analysis on several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set using a public antimicrobial peptide classifier and the antimicrobial peptide prediction core model to obtain consistency analysis results includes: The first prediction result is obtained by predicting several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set using the publicly available antimicrobial peptide classifier. The antimicrobial peptide prediction core model is used to predict several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain a second prediction result. The consistency analysis results are obtained by performing a consistency analysis on the several peptide sequences using the first prediction results and the second prediction results.
[0012] In one embodiment, the step of performing multidimensional quantitative annotation on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain an annotated peptide set includes: The disclosed antimicrobial peptide sequence is used to perform sequence clustering and alignment on the plurality of peptide sequences to obtain homology dimension features, and the plurality of peptide sequences are labeled using the homology dimension features to obtain homology labeling results; The protein three-dimensional structure of the several peptide sequences is predicted to obtain several sequence three-dimensional structures. The structural similarity of the several sequence three-dimensional structures is compared with the publicly disclosed antimicrobial peptide sequence to obtain structural dimension features. The several peptide sequences are then labeled using the structural dimension features to obtain structural labeling results. Distance filtering is performed on the disclosed antimicrobial peptide sequences to obtain characterization spatial dimension features, and the characterization spatial dimension features are used to label the peptide sequences to obtain characterization spatial labeling results. By summarizing the homology annotation results, structural annotation results, and characterization space annotation results, a set of labeled peptides is obtained.
[0013] Furthermore, to achieve the above objectives, this application also proposes a peptide sequence listing device, the peptide sequence listing device comprising: The prediction module is used to predict metagenomic data using a pre-built core model for predicting antimicrobial peptides, thereby obtaining a candidate set of natural antimicrobial peptides. The generation module is used to generate AI peptide sequences through an antimicrobial peptide generation model to obtain a set of AI candidate peptides. The annotation module is used to perform hierarchical interpretable annotation on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain an annotated peptide set, and generate a peptide sequence classification list through the annotated peptide set.
[0014] In addition, to achieve the above objectives, this application also proposes a peptide sequence listing generation device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the peptide sequence listing generation method as described above.
[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the peptide sequence listing method described above.
[0016] One or more technical solutions proposed in this application have at least the following technical effects: This application proposes a peptide sequence list generation method, apparatus, device, and product. It predicts metagenomic data using a pre-constructed antimicrobial peptide prediction core model to obtain a set of natural antimicrobial peptide candidates; generates AI peptide sequences using an antimicrobial peptide generation model to obtain an AI candidate peptide set; and performs hierarchical interpretable annotation on both the natural and AI candidate peptide sets to obtain an annotated peptide set, which is then used to generate a peptide sequence classification list. Thus, it mines natural antimicrobial peptide candidates from metagenomic data while simultaneously generating novel peptide sequences using a model. The unified hierarchical annotation analysis of both types of sequences integrates prediction, generation, and evaluation into a single process, eliminating the need for multiple module separations and streamlining the workflow. This simplifies the process, improves efficiency, and rapidly screens high-quality candidate peptides. It solves the problem of existing antimicrobial peptide prediction and generation processes being independent, leading to cumbersome processes, low efficiency, and difficulty in quickly obtaining highly active antimicrobial peptides, thereby improving the efficiency of peptide sequence list generation. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating the peptide sequence list generation method of this application (Example 1); Figure 2 This is a schematic diagram of the first backbone network architecture involved in the peptide sequence list generation method of this application; Figure 3 This is a schematic diagram of the second backbone network architecture involved in the peptide sequence list generation method of this application; Figure 4 This is a schematic diagram of the first architecture of the core model for predicting antimicrobial peptides involved in the peptide sequence listing method of this application; Figure 5 This is a schematic diagram of the second architecture of the core model for predicting antimicrobial peptides involved in the peptide sequence list generation method of this application; Figure 6 This is a schematic diagram of the ROC curve of the core model for predicting antimicrobial peptides involved in the peptide sequence list generation method of this application; Figure 7 This diagram illustrates the comparison of the AUROC of the core antimicrobial peptide prediction model with other classifiers in the peptide sequence list generation method of this application. Figure 8 This is a schematic diagram of the first architecture of the antimicrobial peptide generation model involved in the peptide sequence listing method of this application; Figure 9 This is a schematic diagram of the second architecture of the antimicrobial peptide generation model involved in the peptide sequence listing method of this application; Figure 10 This diagram illustrates the consistency distribution between the antimicrobial peptide generation model and other generation methods involved in the peptide sequence list generation method of this application. Figure 11 This is a schematic diagram illustrating the parameter fine-tuning of the antimicrobial peptide generation model involved in the peptide sequence list generation method of this application; Figure 12 This is a flowchart illustrating the second embodiment of the peptide sequence list generation method of this application. Figure 13 This is a schematic diagram illustrating the homology dimension annotation involved in the peptide sequence list generation method of this application; Figure 14 This is a schematic diagram illustrating the structural dimension annotation involved in the peptide sequence list generation method of this application; Figure 15 This is a schematic diagram illustrating the characterization space annotation involved in the peptide sequence list generation method of this application; Figure 16 This is a schematic diagram of the module structure of the peptide sequence listing generation device according to an embodiment of this application; Figure 17 This is a schematic diagram of the device structure of the hardware operating environment involved in the peptide sequence list generation method in the embodiments of this application.
[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0023] The main solution of this application embodiment is as follows: Sequences from a publicly available peptide database are summarized to obtain a first publicly available sequence; non-compliant sequences from the first publicly available sequence are removed to obtain a positive sample set; sequences from a general protein resource database are summarized to obtain a second publicly available sequence; sequences from the second publicly available sequence are removed using a keyword exclusion strategy to obtain a negative sample set; stratified sampling is performed on the positive and negative sample sets to obtain a training set and a test set; several sub-classifiers in the pre-trained antimicrobial peptide prediction core model are trained using the training and test sets to obtain the antimicrobial peptide prediction core model. Metagenomic assembly fragments are extracted from the metagenomic data; open reading frames are predicted from the metagenomic assembly fragments to obtain short open reading frames; redundancy is removed from the short open reading frames to obtain non-redundant short peptide sequences; the antimicrobial peptide prediction core model is used to predict the non-redundant short peptide sequences to obtain antimicrobial peptide sequences; positive sequences from the antimicrobial peptide sequences are summarized and scored to obtain a natural antimicrobial peptide candidate set. The antimicrobial peptide samples are divided into an AI sequence training set and an AI sequence validation set. Based on the AI sequence training set, the antimicrobial peptide generation model is fine-tuned using a preset optimizer and training parameters to obtain several sets of model parameters. The AI sequence validation set is used to select the optimal model parameters from these sets. Based on the optimal model parameters, the antimicrobial peptide generation model is used for controlled decoding to generate AI peptide sequences. The AI peptide sequences are then screened to obtain a set of AI candidate peptides. Consistency analysis is performed on several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set using a publicly available antimicrobial peptide classifier and the core antimicrobial peptide prediction model, yielding consistency analysis results. The compositional distribution of these peptide sequences is analyzed to obtain sequence diversity analysis results. Sequence differences and ESM2 high-dimensional latent vector space differences between these peptide sequences and the publicly available antimicrobial peptide sequences are analyzed to obtain sequence novelty analysis results. The matching degree between these peptide sequences and the distribution of real antimicrobial peptides is analyzed to obtain distribution distance analysis results. Based on the consistency analysis results, sequence diversity analysis results, sequence novelty analysis results, and distribution distance analysis results, multidimensional quantitative annotation is performed on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain an annotated peptide set. The publicly available antimicrobial peptide classifier is used to predict several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set, yielding a first prediction result. The core antimicrobial peptide prediction model is used to predict several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set, yielding a second prediction result. Consistency analysis is then performed on these peptide sequences using the first and second prediction results to obtain consistency analysis results.The method involves performing sequence clustering and alignment on several peptide sequences using the publicly available antimicrobial peptide sequences to obtain homology dimension features, and then labeling the peptide sequences using these homology dimension features to obtain homology annotation results. Next, the method involves predicting the three-dimensional structure of the peptide sequences to obtain three-dimensional structures, comparing the structural similarity of these structures with the publicly available antimicrobial peptide sequences to obtain structural dimension features, and then labeling the peptide sequences using these structural dimension features to obtain structural annotation results. Finally, the method involves performing distance filtering on the peptide sequences using the publicly available antimicrobial peptide sequences to obtain characterization space dimension features, and then labeling the peptide sequences using these characterization space dimension features to obtain characterization space annotation results. Finally, the method involves summarizing the homology annotation results, structural annotation results, and characterization space annotation results to obtain a labeled peptide set. This solves the problem of existing methods where antimicrobial peptide prediction and generation are independent, leading to cumbersome processes, low efficiency, and difficulty in quickly obtaining highly active antimicrobial peptides. It achieves the generation of peptide sequence lists and improves the efficiency of peptide sequence list generation. Based on the present invention, addressing the problem that in reality, there are three independent technical routes, where the generation model can only sample candidate sequences and cannot directly evaluate the antibacterial activity of the sequences, and subsequent secondary screening of the sampling results still requires independent prediction models, resulting in low efficiency, a peptide sequence list generation method is designed. The effectiveness of the peptide sequence list generation method of the present invention is verified when generating peptide sequences. Finally, the efficiency of peptide sequence list generation using the method of the present invention is significantly improved.
[0024] In this embodiment, for ease of description, the peptide sequence listing device will be used as the execution subject in the following description.
[0025] Antimicrobial peptides are important candidate substances for combating drug-resistant bacteria. Computational screening and de novo design have become mainstream research directions. Existing technologies are mainly divided into three routes: manual feature classification and prediction, deep learning end-to-end prediction, and generative model de novo design. However, due to the fragmentation of these three technical routes, it is difficult to improve the efficiency of peptide sequence development. First, there is the problem of functional fragmentation, as the generative model can only sample sequences and cannot simultaneously evaluate antimicrobial activity. Second, there is the problem of process redundancy, as the generated results require additional secondary screening, adding processing steps. Third, there is the problem of limited effectiveness, as sequence quality depends entirely on the accuracy of a single prediction model, making it difficult to produce highly active candidate peptides and meet the needs of rapid development of novel antimicrobial drugs.
[0026] This application provides a solution that mines natural antimicrobial peptide candidates from metagenomic data and generates novel peptide sequences using models. Simultaneously, it performs hierarchical annotation analysis on both types of sequences, integrating the prediction, generation, and evaluation processes into one. This eliminates the need for multiple module separation and transfer steps, streamlining the process, improving efficiency, and quickly screening high-quality candidate peptides. It solves the problem that existing antimicrobial peptide prediction and generation are independent, resulting in cumbersome processes, low efficiency, and difficulty in quickly obtaining highly active antimicrobial peptides, thus improving the efficiency of peptide sequence list generation.
[0027] Based on this, embodiments of this application provide a method for generating a peptide sequence list, referring to... Figure 1 , Figure 1 This is a schematic flowchart of the first embodiment of the peptide sequence list generation method of this application.
[0028] In this embodiment, the peptide sequence list generation method includes steps S01 to S03: Step S01: The metagenomic data is predicted using a pre-constructed core model for predicting antimicrobial peptides to obtain a candidate set of natural antimicrobial peptides. Before the implementation of this embodiment, it should be clear that the existing technology has a core defect in peptide sequence prediction and generation. The three current technical routes are independent of each other, and the generation model can only complete the sampling of candidate sequences and cannot directly evaluate the antibacterial activity of the sequences. The sampling results still need to be screened again by independent prediction models. This results in redundancy and low efficiency in the entire antimicrobial peptide design process. Moreover, the effectiveness of the generated sequence is highly dependent on the accuracy of the prediction model, making it difficult to efficiently produce new antimicrobial peptides with high antimicrobial activity. This cannot meet the rapid development needs of new antimicrobial drugs under the current crisis of drug-resistant bacteria.
[0029] Therefore, in order to solve the above problems, this embodiment uniformly collects metagenomic data from multiple types of differentiated habitats, covering metagenomic assembly fragments and MAGs data from mainstream habitats such as rivers, oceans, infant gut, soil, and inland salt lakes. It completes batch aggregation and preliminary standardization of the original data sources, and uses a pre-trained integrated antimicrobial peptide prediction core model to perform layer-by-layer analysis and screening of the standardized metagenomic raw data. It accurately identifies and screens effective sequences with antimicrobial peptide characteristics from massive metagenomic sequences, and obtains a highly reliable set of natural antimicrobial peptide candidates after aggregation and sorting.
[0030] Step S02: AI peptide sequence generation is performed using an antimicrobial peptide generation model to obtain an AI candidate peptide set; We pre-organize high-quality, validated antimicrobial peptide standard sample data and divide the overall sample dataset into an AI sequence training set for model iteration training and an AI sequence validation set for model performance verification and parameter optimization according to a preset ratio. Based on mainstream pre-trained protein generation models as the basic model structure, and with preset optimizer, learning rate, weight decay and iteration round parameters, we conduct full parameter fine-tuning training on the basic generation model. During the training process, we continuously update the model weight parameters and generate multiple sets of model parameter files corresponding to different iteration stages.
[0031] The AI sequence validation set is used to perform batch performance verification on the model parameters of each group, calculate the corresponding validation loss value, and select the parameter combination with the lowest validation loss and the best sequence generation effect as the final optimal model parameters.
[0032] After loading the optimal model parameters, the system autonomously generates novel peptide sequences through a controlled decoding strategy that limits the decoding temperature, truncation sampling threshold, and sequence length range. It outputs a large number of initial AI peptide sequences in batches. Then, by combining sequence length constraints, amino acid compliance, and activity screening conditions, the initial generated sequences are filtered step by step to remove abnormal, invalid, and low-quality sequences, ultimately resulting in a compliant and diverse set of AI candidate peptides.
[0033] Step S03: Perform hierarchical interpretable annotation on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain an annotated peptide set, and generate a peptide sequence classification list through the annotated peptide set.
[0034] The system simultaneously retrieves and acquires the selected natural antimicrobial peptide candidate set and the AI candidate peptide set, and integrates both types of candidate peptide sequences into the same evaluation and annotation system for batch processing. Sequence tasks include sequence prediction consistency verification, natural protein distribution fit analysis, amino acid composition diversity statistics, sequence novelty difference comparison, and overall distribution distance calculation. Based on the multi-dimensional quantitative analysis results, all peptide sequences are then subjected to multi-dimensional quantitative rating and annotation.
[0035] Based on this, further research was conducted on sequence homology clustering and alignment, protein three-dimensional structure prediction and structural similarity analysis, and high-dimensional characterization spatial distance screening. The refined and interpretable classification and annotation of sequence dimension, structural dimension and characterization spatial dimension were completed layer by layer. All quantitative data and hierarchical label information were integrated to form a complete set of labeled peptides. Finally, all labeled peptide sequence information was standardized, and after unification, classification and sorting, a standardized peptide sequence list was generated.
[0036] Specifically, prior to step S01 above, which involves predicting metagenomic data using a pre-constructed antimicrobial peptide prediction core model to obtain a candidate set of natural antimicrobial peptides, the method further includes: Step S0101: Summarize the sequences in the public peptide database to obtain the first public sequence; remove the substandard sequences from the first public sequence to obtain the positive sample set. Step S0102: Sequences are summarized from the general protein resource database to obtain the second public sequence. The second public sequence is then removed using a keyword exclusion strategy to obtain a negative sample set. Step S0103: Perform stratified sampling on the positive sample set and the negative sample set to obtain the training set and the test set; Step S0104: Train several sub-classifiers in the pre-trained antimicrobial peptide prediction core model using the training set and test set to obtain the antimicrobial peptide prediction core model.
[0037] In the process of constructing the sample dataset and training the core prediction model, the first step is to accurately construct the positive and negative sample datasets to completely avoid the shortcomings of traditional antimicrobial peptide datasets, such as mixed data, inconsistent sample quality, and low distinction between positive and negative samples.
[0038] Complete sequence data from mainstream publicly available antimicrobial peptide databases such as APD3, dbAMP 2.0, and DRAMP 3.0 were integrated in batches to obtain an initial first set of publicly available sequences. A refined screening and elimination process was then conducted to address various defects in the sequences. Sequences containing non-standard amino acids, exceeding 100 amino acid intervals, containing base deletions, incomplete sites, or exhibiting complete repetitions or high redundancy were uniformly removed. Only high-quality natural antimicrobial peptide sequences with complete sequence structures, compliant amino acid composition, and clear functional annotations were retained. After integration and standardization, a high-purity and highly representative positive sample set of antimicrobial peptides was formed.
[0039] Simultaneously, a massive amount of original protein sequences from the UniProt general protein resource database were retrieved in batches as second public sequences. A multi-keyword joint exclusion strategy was adopted, relying on functional keyword groups such as antibacterial, bacteriostatic, bactericidal, toxic, immune defense, and microbial inhibition to accurately search and screen all sequences in batches. All protein sequences with antibacterial and related antibacterial activity annotations were completely removed. The remaining sequences were combined with antimicrobial peptide length characteristics to complete secondary screening and deduplication, completely avoiding interference from active sequences, and obtaining a pure negative sample set with a balanced distribution and no active contamination.
[0040] After constructing the two types of sample sets, the positive sample set and the negative sample set are uniformly mixed. A stratified random sampling method is used to strictly preserve the original sample ratio and distribution characteristics. The samples are precisely divided into a training set for model iteration and optimization and a test set for model accuracy verification, hyperparameter tuning, and generalization ability detection according to the standard ratio of 8:2. This effectively avoids the sample distribution shift problem caused by random sampling.
[0041] like Figure 2 as well as Figure 3 As shown, this embodiment constructs a multi-classifier ensemble prediction framework based on the ESM2-650M pre-trained protein representation model. It abandons the shortcomings of traditional single machine learning models, such as low prediction accuracy and poor generalization. It uses the ESM2 bidirectional masked protein large model with excellent feature extraction capabilities as the feature extraction base, and combines it with three highly complementary sub-classifiers, namely random forest, extreme gradient boosting, and support vector machine, to form an ensemble prediction structure.
[0042] The three sub-classifiers were simultaneously iteratively trained using the high-quality training set that had been divided. The ESM2 model was used to extract 1280-dimensional high-dimensional protein features to provide refined feature inputs for each classifier. During the training process, the weight parameters and hyperparameters of each classifier were continuously optimized.
[0043] After training, multi-dimensional performance verification and parameter fine-tuning were carried out using an independent test set. The prediction weights and output results of the three sub-classifiers were integrated by combining the majority voting fusion strategy, which effectively made up for the prediction bias of a single classifier. Finally, a core model for antimicrobial peptide prediction with high accuracy, strong generalization ability and high robustness was constructed. Compared with the traditional single prediction model, the accuracy and reliability of antimicrobial peptide sequence identification were significantly improved.
[0044] More specifically, step S01 above, which involves predicting metagenomic data using a pre-constructed core model for antimicrobial peptide prediction to obtain a candidate set of natural antimicrobial peptides, includes: Step S011: Extract metagenomic assembly fragments from the metagenomic data; Step S012: Predict open reading frames for the assembled metagenomic fragments to obtain short open reading frames; Step S013: Redundancy removal is performed on the short open reading frame to obtain a non-redundant short peptide sequence; Step S014: The non-redundant short peptide sequence is predicted using the antimicrobial peptide prediction core model to obtain the antimicrobial peptide sequence. Step S015: The antimicrobial peptide sequences are summarized and scored to obtain a candidate set of natural antimicrobial peptides.
[0045] Based on the trained core model for predicting antimicrobial peptides, we will carry out large-scale mining of natural antimicrobial peptides in multiple habitats, breaking through the technical bottlenecks of traditional antimicrobial peptide mining, such as single sample sources, cumbersome screening process, and low detection rate of effective sequences.
[0046] First, the raw metagenomic data of various specific habitats, such as rivers, oceans, gut, soil, and salt lakes, were preprocessed to accurately extract complete metagenomic assembly fragments and metagenomic assembly genome data, preserving the microbial gene characteristics of different habitats and enriching the species and environmental diversity for subsequent antimicrobial peptide mining.
[0047] Subsequently, the Prodigal tool in meta mode was used to predict open reading frames across all metagenomic assembly fragments, accurately identifying gene coding regions. At the same time, strict screening was conducted based on the characteristics of antimicrobial peptide sequences, retaining only short open reading frames with an amino acid length of less than 100, no abnormal bases, and no non-standard amino acids, thus accurately locking potential short peptide coding regions and filtering out non-target sequences from the source.
[0048] To avoid repetitive predictions and data bias caused by sequence redundancy, the CD-HIT redundancy removal algorithm was adopted. Using 0.95 sequence identity as the threshold, batch deduplication and similarity filtering were performed on all short open reading frame sequences obtained by screening. Completely eliminating completely repetitive and highly similar redundant sequences, a set of non-redundant short peptide sequences with high sequence purity, excellent diversity and no redundancy interference was obtained.
[0049] like Figure 4 as well as Figure 5 As shown, all preprocessed non-redundant short peptide sequences are input into the fully trained integrated antimicrobial peptide prediction core model. Relying on the ESM2-650M model, 1280-dimensional high-dimensional features of each sequence are automatically extracted without manual feature engineering intervention, achieving end-to-end adaptive feature extraction.
[0050] Simultaneously, three sub-classifiers—random forest, extreme gradient boosting, and support vector machine—are used to perform parallel antimicrobial activity prediction. A three-to-two majority voting ensemble strategy is used to uniformly determine the antimicrobial peptide attributes of a single sequence, effectively reducing the prediction error of a single model and accurately screening out positive antimicrobial peptide sequences with potential antimicrobial activity.
[0051] The final result is as follows Figure 6 as well as Figure 7 As shown, the integrated prediction core model in this embodiment exhibits excellent ROC curve performance, with an AUROC value of 0.964, which is significantly better than traditional mainstream antimicrobial peptide prediction models such as cAMP, amp_CRE, iAMPCN, and amPEP, greatly improving prediction accuracy and stability.
[0052] Finally, all antimicrobial peptide sequences that were identified as positive by the model were summarized. Based on the predicted confidence probabilities output by the model, they were prioritized and sorted in descending order. High-confidence, high-quality sequences were selected and retained. Finally, a set of natural antimicrobial peptide candidates that covers multiple specific habitats, has high reliability, and is rich in diversity was obtained.
[0053] Further, step S02 above, which involves generating AI peptide sequences using an antimicrobial peptide generation model to obtain a set of AI candidate peptides, includes: Step S021: Divide the antimicrobial peptide samples into an AI sequence training set and an AI sequence validation set; Step S022: Based on the AI sequence training set, the antimicrobial peptide generation model is fine-tuned with full parameters using a preset optimizer and training parameters to obtain several sets of model parameters; Step S023: Select the optimal model parameters by using the AI sequence validation set to evaluate the several sets of model parameters; Step S024: Based on the optimal model parameters, the antimicrobial peptide generation model is used for controlled decoding to generate the AI peptide sequence. Step S025: Screen the AI peptide sequences to obtain a set of AI candidate peptides.
[0054] This embodiment overcomes the shortcomings of traditional antimicrobial peptide generation models, such as weak generalization, low activity of generated sequences, severe homogenization, and inability to adapt to the needs of specific habitat mining. It achieves high-quality and novel AI antimicrobial peptide sequence generation based on a finely tuned and optimized causal protein language model.
[0055] First, the massive, high-quality, publicly available antimicrobial peptide standard sample dataset was organized and randomly split into two parts according to the optimal ratio of 95:5. These parts were then used to construct an AI sequence training set for model iteration training and an AI sequence validation set for performance verification and parameter optimization. This approach balances the sufficiency of training data with the independence of validation data, ensuring the effectiveness of model training.
[0056] like Figure 8 as well as Figure 9 As shown, this embodiment uses the ProGen2-base 764M causal protein pre-trained language model as the basis for generating the framework. This model has excellent protein sequence autoregressive generation capabilities and is adapted to the regular learning of short peptide sequences and the sampling of novel sequences.
[0057] During the model training phase, the AdamW adaptive gradient optimizer is configured, along with refined training hyperparameters such as fixed learning rate, weight decay, and iteration rounds. Based on the AI sequence training set, the basic model is subjected to full parameter fine-tuning training. During the 10 complete iteration training rounds, all weight parameters of the model are continuously updated, and the amino acid arrangement, structural features, and physicochemical properties of the antimicrobial peptide sequence are continuously fitted. Multiple sets of model parameter files corresponding to different training stages are generated iteratively, realizing the precise adaptation of the model from general protein generation to antimicrobial peptide-specific generation.
[0058] This embodiment also includes Figure 11As shown, the changes in training loss and validation loss throughout the training process are recorded. With the increase of iterations, the model training loss and validation loss continue to decrease steadily and eventually converge. There is no overfitting or underfitting phenomenon, which proves that the model fine-tuning training effect is excellent and the parameter fitting degree is good.
[0059] After training, the parameters of each iterative model are loaded into the generated model in sequence. Batch inference verification is carried out through the AI sequence validation set. The validation loss value and sequence generation compliance rate corresponding to each set of parameters are accurately calculated. The parameter combination with the lowest validation loss, the best sequence generation quality, and the strongest generalization ability is selected as the final optimal model parameters.
[0060] The refined controlled decoding generation mechanism is initiated based on the optimal model parameters. It integrates three sampling strategies: temperature sampling, Top-K truncation sampling, and Top-P kernel sampling. This effectively balances the conservatism and innovation of sequence generation, while strictly limiting the peptide sequence generation length to the optimal range of 8 to 50 amino acids. It also generates novel AI peptide sequences in batches through autoregression.
[0061] Ultimately Figure 10 As shown, compared with traditional antimicrobial peptide generation models such as AMPGAN, PepDiffusion, HydrAMP, ampDiffusion, and PepcVAE, the generation model finely tuned in this embodiment has significantly better output sequence consistency, effectiveness, and quality level.
[0062] Finally, a multi-level fine screening was carried out on the initially generated massive AI peptide sequences to successively eliminate inferior sequences with abnormal length, unbalanced amino acid composition, substandard physicochemical properties, and no potential antibacterial activity. In the end, a set of AI candidate peptides with high novelty, sufficient diversity, excellent physicochemical properties, and reasonable distribution was obtained, realizing the efficient and high-quality de novo design of new antibacterial peptide sequences.
[0063] This embodiment, through the above-described scheme, specifically uses a pre-constructed core model for antimicrobial peptide prediction to predict metagenomic data, obtaining a set of natural antimicrobial peptide candidates; it then uses an antimicrobial peptide generation model to generate AI peptide sequences, obtaining a set of AI candidate peptides; finally, it performs hierarchical interpretable annotation on both the natural and AI candidate peptide sets to obtain an annotated peptide set, and generates a peptide sequence classification list from the annotated peptide set. Thus, it mines natural antimicrobial peptide candidates from metagenomic data on one hand, and generates novel peptide sequences using a model on the other, while simultaneously conducting hierarchical annotation analysis on both types of sequences. This integrates the prediction, generation, and evaluation processes into one, eliminating the need for multiple module separations and streamlining the process. This simplifies the workflow, improves efficiency, and rapidly screens high-quality candidate peptides. It solves the problem of existing antimicrobial peptide prediction and generation being independent, leading to cumbersome processes, low efficiency, and difficulty in quickly obtaining highly active antimicrobial peptides, thereby improving the efficiency of peptide sequence list generation.
[0064] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 12 In step S03, the method for generating a peptide sequence list further includes steps S031 to S036: hierarchical interpretable annotation is performed on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain the annotated peptide set. Step S031: By using the publicly disclosed antimicrobial peptide classifier and the core model for predicting antimicrobial peptides, a consistency analysis is performed on several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain the consistency analysis results. Step S032: Analyze the compositional distribution of the several peptide sequences to obtain the sequence diversity analysis results; Step S033: Analyze the sequence differences and ESM2 high-dimensional hidden vector space differences of the several peptide sequences and the publicly disclosed antimicrobial peptide sequences to obtain the sequence novelty analysis results; Step S034: Analyze the degree of matching between the several peptide sequences and the distribution of real antimicrobial peptides to obtain the distribution distance analysis results; Step S035: Based on the consistency analysis results, sequence diversity analysis results, sequence novelty analysis results, and distribution distance analysis results, multidimensional quantitative annotation is performed on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain the annotated peptide set.
[0065] First, a consistency evaluation of model predictions is conducted. Multiple mainstream publicly available antimicrobial peptide classification tools, such as Macrel, cAMP, and APEX, are simultaneously invoked along with the core antimicrobial peptide prediction model independently constructed in this embodiment to form a dual-pathway prediction verification mechanism. By cross-comparing the results of multiple models, the credibility of the antimicrobial activity prediction results of a single peptide sequence is accurately determined, avoiding errors caused by prediction failures of a single model.
[0066] Secondly, sequence quality evaluation is carried out. The pseudo-perplexity of a single peptide sequence is calculated based on the ESM2 pre-trained model. The degree of fit between each candidate peptide sequence and the distribution of natural proteins is accurately quantified. The lower the perplexity, the better the sequence fits the distribution pattern of natural antimicrobial peptides and the higher the sequence quality. This effectively identifies artificial pseudo sequences and inferior abnormal sequences.
[0067] Subsequently, sequence diversity evaluation was carried out. The overall frequency and arrangement characteristics of 20 standard amino acids in the candidate peptide set were statistically analyzed in batches. The uniformity of amino acid composition was quantified by Shannon entropy calculation formula. The higher the entropy value, the richer the amino acid arrangement of the set sequence and the better the overall diversity, effectively avoiding the problems of homogenization and uniformity of the generated sequences.
[0068] Simultaneously, sequence novelty evaluation is carried out, combining sequence-level and spatial-level dual comparison methods. Local sequence motif differences are compared through trimer Jaccard similarity, and global spatial distribution differences are compared through ESM2 high-dimensional characterization of spatial cosine distance. This comprehensively quantifies the degree of difference between candidate peptides and known published antimicrobial peptides, and accurately screens novel peptide sequences with innovative potential.
[0069] Finally, a distribution distance evaluation was conducted using the Fréchet ChemNet Distance index to measure the global distribution deviation between the candidate peptide set and the real antimicrobial peptide training dataset, thereby assessing the overall distribution rationality and sample reliability of the batch sequences.
[0070] After completing independent quantitative analysis across five dimensions, the quantitative scores and analysis results of all dimensions are integrated to perform comprehensive rating and refined quantitative annotation on each candidate peptide sequence. This distinguishes between high-confidence, high-quality sequences, medium-confidence, ordinary sequences, and low-confidence, low-quality sequences, achieving comprehensive multi-dimensional quantitative grading of the two types of candidate peptide sets. This provides accurate and comprehensive quantitative pre-labeling for subsequent three-layer interpretable fine annotation, significantly improving the accuracy and effectiveness of antimicrobial peptide screening.
[0071] Specifically, step S031 above, which involves performing a consistency analysis on several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set using a public antimicrobial peptide classifier and the core model for antimicrobial peptide prediction, to obtain the consistency analysis results, includes: Step S0311: Predict several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set using the publicly available antimicrobial peptide classifier to obtain a first prediction result; Step S0312: The antimicrobial peptide prediction core model is used to predict several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain a second prediction result. Step S0313: Perform consistency analysis on the plurality of peptide sequences using the first prediction result and the second prediction result to obtain consistency analysis results.
[0072] First, multiple mainstream, open-source, and stable antimicrobial peptide classifiers in the industry are called in batches. The model inference parameters and prediction thresholds are unified. All peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set are independently predicted in batches. Each sequence is determined to have antimicrobial activity. The positive and negative prediction results of a single sequence are accurately recorded. The prediction data of all public models are summarized and integrated to obtain the first prediction result corresponding to the external general model.
[0073] Meanwhile, all candidate peptide sequences from the same batch are input into the core antimicrobial peptide prediction model that was independently trained, validated in multiple rounds, and has the best performance in this application. Relying on the high-dimensional feature extraction capability of the ESM2 model and the multi-classifier integrated voting mechanism, the antimicrobial activity of each peptide sequence is accurately determined independently, and the second prediction result corresponding to the self-developed model of this application is output.
[0074] After completing the dual-pathway prediction, the first and second prediction results corresponding to a single peptide sequence are aligned and accurately compared one by one. The prediction overlap of the two models is comprehensively statistically analyzed, and the three judgment scenarios of complete consistency, partial consistency, and complete inconsistency are distinguished in detail. Sequences with different degrees of overlap are differentially labeled, and the prediction consistency classification of all candidate peptide sequences is completed in batches. Finally, complete and refined sequence consistency analysis results are generated, providing core and reliable basis for subsequent sequence quality rating and high-quality sequence screening.
[0075] More specifically, step S035 above, which involves multi-dimensional quantitative annotation of the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain the annotated peptide set, includes: Step S0351: Perform sequence clustering and comparison on the several peptide sequences using the disclosed antimicrobial peptide sequence to obtain homology dimension features, and label the several peptide sequences using the homology dimension features to obtain homology labeling results; Step S0352: Perform protein three-dimensional structure prediction on the several peptide sequences to obtain several sequence three-dimensional structures. Perform structural similarity comparison on the several sequence three-dimensional structures with the disclosed antimicrobial peptide sequence to obtain structural dimension features. Then, use the structural dimension features to annotate the several peptide sequences to obtain structural annotation results. Step S0353: The distance screening of the plurality of peptide sequences is performed through the disclosed antimicrobial peptide sequence to obtain the characterization spatial dimension features, and the plurality of peptide sequences are labeled through the characterization spatial dimension features to obtain the characterization spatial labeling results; Step S0354: Summarize the homology annotation results, structural annotation results, and characterization space annotation results to obtain the labeled peptide set.
[0076] like Figure 13As shown, in the sequence homology dimension annotation process, a publicly known standard sequence dataset of antimicrobial peptides is retrieved as a global comparison benchmark. The MMseqs high-speed sequence clustering and comparison tool is used to perform batch global homology retrieval and clustering analysis on all candidate peptide sequences. The preset e-value of less than 0.001 is used as the threshold for determining the significance of sequence alignment. Candidate sequences that have homologous conserved motifs with known antimicrobial peptides and whose sequence similarity meets the standard are accurately screened. The extended derivative sequences of known antimicrobial peptide families are accurately identified. Based on the degree of homology matching and clustering results, differentiated sequence homology labeling is completed, and finally, complete sequence homology annotation results are obtained, realizing the accurate mining of new sequences of known antimicrobial peptide families.
[0077] like Figure 14 As shown, in the process of structural dimension annotation, it breaks through the limitations of traditional methods that rely solely on sequence alignment, and achieves deep feature mining from the perspective of protein spatial structure. Through the ESMFold high-precision protein structure prediction tool, it performs fully automatic three-dimensional structure modeling for each candidate peptide sequence and accurately generates the corresponding three-dimensional spatial structure file.
[0078] Subsequently, the Foldseek professional structure comparison tool was used to perform global similarity matching between the three-dimensional structure of the candidate peptide and the publicly available standard three-dimensional structure of antimicrobial peptide. A TM-score greater than 0.5 was preset as the threshold for structural homology determination. Candidate sequences with highly similar spatial folding structures and conserved functional domains were accurately screened. Core structural dimensional features were extracted and structural labels were completed to obtain refined structural labeling results, ensuring the reliability of the antimicrobial activity of the candidate peptides from the functional structure level.
[0079] like Figure 15 As shown, in the process of characterization space dimension annotation, relying on the high-dimensional characterization capability of the ESM2-650M pre-trained large model, all candidate peptide sequences are uniformly mapped to a 1280-dimensional high-dimensional feature space, the global characterization features of each sequence are extracted, and the nearest neighbor cosine distance between each candidate sequence and the centroid of the known antimicrobial peptide positive set is accurately calculated.
[0080] By combining the 90th percentile distance threshold of the training set to complete the differential screening, the conventional sequences that are highly similar to the distribution of natural antimicrobial peptides and the novel sequences generated at the far end with significant spatial distribution differences are accurately identified. The feature determination and innovative labeling of the spatial dimension are completed, and the spatial labeling results are obtained.
[0081] Finally, by integrating the refined results from three dimensions—sequence homology annotation, structural homology annotation, and far-end innovation annotation in the characterization space—a hierarchical and interpretable classification of all candidate peptide sequences was completed. Ultimately, a standardized set of labeled peptides with multidimensional quantitative scoring and three layers of interpretable feature labels was obtained. This not only enables the accurate screening of high-quality antimicrobial peptides but also effectively distinguishes between conserved derived sequences and novel innovative sequences, significantly improving the innovation and effectiveness of novel antimicrobial peptide discovery.
[0082] This embodiment, through the above-described scheme, specifically uses a public antimicrobial peptide classifier and the core model for antimicrobial peptide prediction to perform consistency analysis on several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set, obtaining consistency analysis results; analyzes the compositional distribution of the several peptide sequences to obtain sequence diversity analysis results; analyzes the sequence differences and ESM2 high-dimensional latent vector space differences between the several peptide sequences and the public antimicrobial peptide sequences to obtain sequence novelty analysis results; analyzes the degree of matching between the several peptide sequences and the distribution of real antimicrobial peptides to obtain distribution distance analysis results; based on the consistency analysis results, sequence diversity analysis results, sequence novelty analysis results, and distribution distance analysis results, performs multidimensional quantitative annotation on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain an annotated peptide set. Therefore, this method mines natural antimicrobial peptide candidates from metagenomic data on the one hand, and generates novel peptide sequences using models on the other. Simultaneously, it conducts hierarchical annotation analysis on both types of sequences, integrating the prediction, generation, and evaluation processes into one, eliminating the need for multiple module splitting and transfer steps, streamlining the process, improving efficiency, and quickly screening high-quality candidate peptides. This solves the problem that existing antimicrobial peptide prediction and generation are independent, resulting in cumbersome processes, low efficiency, and difficulty in quickly obtaining highly active antimicrobial peptides, thus improving the efficiency of peptide sequence list generation.
[0083] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the peptide sequence list generation method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0084] This application also provides a peptide sequence listing device; please refer to... Figure 16 The peptide sequence listing device includes: Prediction module 10 is used to predict metagenomic data through a pre-built core model for predicting antimicrobial peptides to obtain a candidate set of natural antimicrobial peptides. Generation module 20 is used to generate AI peptide sequences through an antimicrobial peptide generation model to obtain a set of AI candidate peptides. The annotation module 30 is used to perform hierarchical interpretable annotation on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain an annotated peptide set, and generate a peptide sequence classification list through the annotated peptide set.
[0085] The peptide sequence listing device provided in this application, employing the peptide sequence listing method described in the above embodiments, solves the technical problem that existing methods for predicting and generating antimicrobial peptides are independent, resulting in cumbersome processes, low efficiency, and difficulty in rapidly obtaining highly active antimicrobial peptides. Compared with the prior art, the beneficial effects of the peptide sequence listing device provided in this application are the same as those of the peptide sequence listing method provided in the above embodiments, and other technical features in the peptide sequence listing device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0086] This application provides a peptide sequence listing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the peptide sequence listing method in Embodiment 1 above.
[0087] The following is for reference. Figure 17 The diagram illustrates a structural schematic suitable for implementing the peptide sequence listing generation device of the embodiments of this application. The peptide sequence listing generation device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 17 The peptide sequence listing device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this application.
[0088] like Figure 17As shown, the peptide sequence listing generation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the peptide sequence listing generation device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the peptide sequence listing device to communicate wirelessly or wiredly with other devices to exchange data. While the figure shows peptide sequence listing devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0089] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0090] The peptide sequence listing generation device provided in this application, employing the peptide sequence listing generation method described in the above embodiments, solves the technical problem that existing antimicrobial peptide prediction and generation are independent, resulting in cumbersome processes, low efficiency, and difficulty in quickly obtaining highly active antimicrobial peptides. Compared with the prior art, the beneficial effects of the peptide sequence listing generation device provided in this application are the same as those of the peptide sequence listing generation method provided in the above embodiments, and other technical features in this peptide sequence listing generation device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0091] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0092] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0093] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the peptide sequence list generation method in the above embodiments.
[0094] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0095] The aforementioned computer-readable storage medium may be included in the peptide sequence listing device; or it may exist independently and not assembled into the peptide sequence listing device.
[0096] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the peptide sequence listing device, the peptide sequence listing device performs the following actions: predicts metagenomic data using a pre-built antimicrobial peptide prediction core model to obtain a set of natural antimicrobial peptide candidates; generates AI peptide sequences using an antimicrobial peptide generation model to obtain a set of AI candidate peptides; performs hierarchical interpretable annotation on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain a set of annotated peptides; and generates a peptide sequence classification list using the annotated peptide set.
[0097] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0099] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0100] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described peptide sequence list generation method. This solves the technical problem that existing methods for predicting and generating antimicrobial peptides are independent, resulting in cumbersome processes, low efficiency, and difficulty in quickly obtaining highly active antimicrobial peptides. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the peptide sequence list generation method provided in the above embodiments, and will not be repeated here.
[0101] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the peptide sequence listing method described above.
[0102] The computer program product provided in this application solves the technical problem that existing methods for predicting and generating antimicrobial peptides are independent, resulting in cumbersome processes, low efficiency, and difficulty in quickly obtaining highly active antimicrobial peptides. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the peptide sequence listing generation method provided in the above embodiments, and will not be repeated here.
[0103] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for generating a peptide sequence list, characterized in that, The method for generating the peptide sequence list includes: A set of natural antimicrobial peptide candidates is obtained by predicting metagenomic data using a pre-constructed core model for antimicrobial peptide prediction. AI peptide sequences were generated using an antimicrobial peptide generation model to obtain a set of AI candidate peptides. The candidate set of natural antimicrobial peptides and the candidate set of AI peptides are hierarchically interpretable labeled to obtain a labeled peptide set, and a peptide sequence classification list is generated from the labeled peptide set.
2. The peptide sequence listing method as described in claim 1, characterized in that, Before the step of predicting metagenomic data using a pre-constructed antimicrobial peptide prediction core model to obtain a candidate set of natural antimicrobial peptides, the method further includes: The sequences of the publicly available peptide database are summarized to obtain the first publicly available sequence. The substandard sequences of the first publicly available sequence are removed to obtain the positive sample set. Sequences were aggregated from a general protein resource database to obtain a second public sequence. The second public sequence was then removed using a keyword exclusion strategy to obtain a negative sample set. Stratified sampling is performed on the positive sample set and the negative sample set to obtain the training set and the test set; The core model for antimicrobial peptide prediction is obtained by training several sub-classifiers in the pre-trained antimicrobial peptide prediction core model using the training set and the test set.
3. The peptide sequence listing method as described in claim 1, characterized in that, The step of predicting metagenomic data using a pre-constructed antimicrobial peptide prediction core model to obtain a candidate set of natural antimicrobial peptides includes: Extract metagenomic assembly fragments from the metagenomic data; Open reading frames (ORFs) are predicted from the assembled metagenomic fragments to obtain short open reading frames. The short open reading frames are deredundant to obtain non-redundant short peptide sequences; The antimicrobial peptide sequence is obtained by predicting the non-redundant short peptide sequence using the core model for antimicrobial peptide prediction. The antimicrobial peptide sequences are summarized and scored to obtain a candidate set of natural antimicrobial peptides.
4. The peptide sequence listing method as described in claim 1, characterized in that, The step of generating AI peptide sequences using an antimicrobial peptide generation model to obtain an AI candidate peptide set includes: The antimicrobial peptide samples were divided into an AI sequence training set and an AI sequence validation set. Based on the AI sequence training set, the antimicrobial peptide generation model is fine-tuned with all parameters through a preset optimizer and training parameters to obtain several sets of model parameters. The optimal model parameters are obtained by selecting parameters from the several sets of model parameters using the AI sequence validation set. Based on the optimal model parameters, the AI peptide sequence is obtained by controlled decoding and generation using the antimicrobial peptide generation model. The AI peptide sequences are screened to obtain a set of AI candidate peptides.
5. The peptide sequence listing method as described in claim 1, characterized in that, The step of performing hierarchical interpretable annotation on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain the annotated peptide set includes: Consistency analysis was performed on several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set using the publicly available antimicrobial peptide classifier and the core model for antimicrobial peptide prediction, and the consistency analysis results were obtained. The compositional distribution of the aforementioned peptide sequences was analyzed to obtain the sequence diversity analysis results; The sequence differences between the aforementioned peptide sequences and the publicly disclosed antimicrobial peptide sequences, as well as the differences in the ESM2 high-dimensional latent vector space, were analyzed to obtain the sequence novelty analysis results. The degree of matching between the aforementioned peptide sequences and the distribution of real antimicrobial peptides was analyzed to obtain the distribution distance analysis results; Based on the consistency analysis results, sequence diversity analysis results, sequence novelty analysis results, and distribution distance analysis results, the natural antimicrobial peptide candidate set and the AI candidate peptide set are subjected to multidimensional quantitative annotation to obtain the annotated peptide set.
6. The peptide sequence listing method as described in claim 5, characterized in that, The step of performing consistency analysis on several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set using a public antimicrobial peptide classifier and the antimicrobial peptide prediction core model to obtain the consistency analysis results includes: The first prediction result is obtained by predicting several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set using the publicly available antimicrobial peptide classifier. The antimicrobial peptide prediction core model is used to predict several peptide sequences in the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain a second prediction result. The consistency analysis results are obtained by performing a consistency analysis on the several peptide sequences using the first prediction results and the second prediction results.
7. The peptide sequence listing method as described in claim 5, characterized in that, The step of performing multi-dimensional quantitative annotation on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain the annotated peptide set includes: The disclosed antimicrobial peptide sequence is used to perform sequence clustering and alignment on the plurality of peptide sequences to obtain homology dimension features, and the plurality of peptide sequences are labeled using the homology dimension features to obtain homology labeling results; The protein three-dimensional structure of the several peptide sequences is predicted to obtain several sequence three-dimensional structures. The structural similarity of the several sequence three-dimensional structures is compared with the publicly disclosed antimicrobial peptide sequence to obtain structural dimension features. The several peptide sequences are then labeled using the structural dimension features to obtain structural labeling results. Distance filtering is performed on the disclosed antimicrobial peptide sequences to obtain characterization spatial dimension features, and the characterization spatial dimension features are used to label the peptide sequences to obtain characterization spatial labeling results. By summarizing the homology annotation results, structural annotation results, and characterization space annotation results, a set of labeled peptides is obtained.
8. A peptide sequence listing device, characterized in that, The peptide sequence list generation device includes: The prediction module is used to predict metagenomic data using a pre-built core model for predicting antimicrobial peptides, thereby obtaining a candidate set of natural antimicrobial peptides. The generation module is used to generate AI peptide sequences through an antimicrobial peptide generation model to obtain a set of AI candidate peptides. The annotation module is used to perform hierarchical interpretable annotation on the natural antimicrobial peptide candidate set and the AI candidate peptide set to obtain an annotated peptide set, and generate a peptide sequence classification list through the annotated peptide set.
9. A peptide sequence listing device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the peptide sequence listing method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the peptide sequence listing method as described in any one of claims 1 to 7.