A method for mining potential antioxidant collagen peptides based on artificial intelligence technology
By constructing an antioxidant peptide discrimination model using artificial intelligence technology and embedding features using ProtTrans technology, the problems of insufficient selectivity and high cost of traditional enzymatic hydrolysis methods are solved, enabling efficient identification and screening of novel antioxidant collagen peptides, which are applicable to drug development and cosmetics development.
Patent Information
- Application Number
- CN202310639003.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-31
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2043-05-31
AI Technical Summary
Traditional enzymatic hydrolysis methods for preparing antioxidant collagen peptides suffer from insufficient selectivity, risks of pathogen transmission, and sustainability issues, and are also costly, making it difficult to efficiently discover novel antioxidant collagen peptides.
Artificial intelligence technology is used to mine potential antioxidant peptides from collagen sequence databases. An antioxidant peptide discrimination model is constructed, and the ProtT5 protein language model in the ProtTrans open source package is used for feature embedding and training. Signal peptides and random peptides are introduced as negative samples to improve model performance.
It enables the efficient and low-cost identification of novel antioxidant peptides from previously unseen data, improving the accuracy and generalization ability of the model. It can directly screen potential antioxidant peptides from large peptide libraries and is suitable for practical applications.
Smart Images

Figure CN116631517B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to anti-oxidative collagen peptide mining technology, more particularly, it relates to a method for mining potential anti-oxidative collagen peptides based on artificial intelligence technology. BACKGROUND
[0002] Collagen is the most widely distributed protein in our body, which has the functions of supporting, connecting, moisturizing and protecting cells, and can also affect the growth, differentiation and metabolism of cells. Collagen peptide is a small molecule of collagen, and collagen peptide is a small molecule peptide formed by hydrolysis of collagen by collagenase. Compared with collagen which is difficult to be directly absorbed by the human body, collagen peptide with smaller molecular weight can be actively absorbed by the human body. Collagen peptide not only has rich sequence types, but also has many functions, including stabilizing calcium in the body, strengthening bones, enhancing the elasticity of the skin, delaying aging, improving the immune capacity of the body, and improving the sleep state. It is well known that collagen is rapidly lost with the growth of human age (the annual loss rate is as high as 1.5%), and the collagen content in the human body at the age of 40 is about half of that at the age of 18. It is worth noting that the reason for the loss of collagen not only comes from the growth of human age, but also is related to photoaging caused by ultraviolet radiation. Strong ultraviolet rays can penetrate the epidermis of the skin and reach the dermis, increase the free radicals in the dermis, and the oxidation of free radicals can destroy the collagen molecules in the body, resulting in problems such as relaxation, color spots and aging. In theory, supplementing anti-oxidative collagen peptides with the ability to scavenge free radicals can play a dual role: on the one hand, as active collagen peptides, it can directly penetrate the skin and help cells to produce more collagen; on the other hand, it can play an antioxidant role by scavenging free radicals to achieve the purpose of delaying aging.
[0003] The traditional method for identifying antioxidant collagen peptides is enzymatic hydrolysis: collagen is first extracted from animal tissues, then specific collagenases (such as collagenase, cysteine protease, etc.) are used to hydrolyze collagen under suitable conditions to form collagen peptides. Next, purification and separation techniques are used to obtain collagen peptides with specific molecular weight ranges. To evaluate the antioxidant properties of these collagen peptides, two experimental methods can be used: the ABTS method (2,2'-azino-bis(3-ethylbenzothiazoline-6-sulfonic acid) radical method) and the ORAC method (oxygen radical absorbance capacity assay). Through these two methods, the free radical scavenging ability of collagen peptides can be determined, thereby evaluating their antioxidant properties. Common oral collagen peptide health products, collagen peptide cosmetics, and medical collagen peptides are all prepared using this method. Although enzymatic hydrolysis is a common method for preparing collagen peptides, it still has some potential drawbacks: (1) selectivity of collagenase, collagenase has selectivity for specific peptide bonds, which can lead to uneven sampling of collagen peptide fragment sequences and incomplete acquisition of potential active peptide fragments; (2) risk of pathogen transmission, animal-derived collagen proteins may carry pathogens such as viruses, bacteria, and parasites, although the enzymatic hydrolysis and purification process can reduce this risk, there is still a small risk of pathogen transmission; (3) sustainability issues and production costs, the large-scale use of animal-derived collagen proteins can have negative environmental impacts, and the enzymatic hydrolysis method involves multiple steps, high-cost enzyme preparations, and equipment.
[0004] Another more intelligent and efficient approach is to use artificial intelligence technology to mine potential antioxidant collagen peptides from large databases of collagen protein sequences, that is, to establish an intelligent discrimination model to pre-evaluate the antioxidant capacity of polypeptide sequences. The advantages of this method are: (1) artificial intelligence models can directly perform virtual screening on proteome sequences, avoiding the traditional hydrolysis-determination-identification cycle, thereby reducing experimental testing workload and costs; (2) artificial intelligence models can theoretically evaluate the antioxidant activity of all possible short peptide sequences of collagen proteins, eliminating the dependence on the selectivity of specific collagenase preparations for peptide bonds, and have a greater probability of discovering new antioxidant collagen peptides that are not easily discovered by traditional enzymatic hydrolysis. In fact, there have been successful commercial cases of using artificial intelligence technology to mine natural active peptides. The AI biotechnology company Nuritas in Ireland uses artificial intelligence technology and genomics to mine a natural active peptide (HGPVEMPYTLLYPSSK, referred to as PeptiYouthTM) from pea genomics data. Through in vitro testing and clinical verification, it has been proven that this natural active peptide has multiple effects such as promoting cell proliferation and migration, stimulating elastin and collagen synthesis, and anti-wrinkle.
[0005] In the process of establishing an artificial intelligence model to determine whether a polypeptide sequence has antioxidant activity, a key technical challenge affecting the quality of model prediction is how to effectively perform feature engineering, i.e., sequence embedding, on the polypeptide sequence, so as to obtain a feature vector that can distinguish between antioxidant sequences and non-antioxidant sequences. Whether a polypeptide has antioxidant activity depends not only on the amino acid composition of the sequence, but also on the tertiary structure, hydrogen bond network, and hydrophilicity of the sequence. Antioxidant peptides can scavenge free radicals, donate electrons, or / and chelate metals. It has been proven that strong antioxidant peptides contain hydrophobic amino acids (such as Pro, Met, Trp, and Phe) and one or more His, Cys, and Tyr residues. Two aromatic amino acids, Tyr and Phe, are considered to be substances that can directly scavenge free radicals, and their phenolic groups have special ability as hydrogen donors. Cys can donate sulfhydryl, and the imidazole group of His can chelate and capture free radicals through proton donation ability. How to design more comprehensive and abstract sequence features to depict these properties requires the integration of more biological knowledge and artificial intelligence technology.
[0006] In recent years, large language models (LMs) have become a powerful paradigm for learning embeddings directly from large, unlabeled natural language datasets. Today, these state-of-the-art techniques in the natural language domain are being used for protein research, making exciting breakthroughs in protein sequence modeling. Protein language models (pLMs) have been shown to capture complex relationships between protein residues by simply pre-training using large, unlabeled protein sequence databases. Large protein databases represent a treasure trove of rich databases, and in recent years, researchers have developed various pLMs to extract complex abstract information from protein sequences. For example, a recently developed pLM called ProtTrans is a protein large language model package that applies six previously published transformer-based architectures (Transformer-XL, BERT, Albert, Xlnet, T5, and Electra) to the pre-training of large protein sequence libraries (including UniParc and BFD) and is completely community-developed (https: / / github.com / agemagician / ProtTrans). The study shows that the embedding of protein sequences (i.e., sequence feature vectors) can accurately predict the secondary structure and subcellular localization of each residue. SUMMARY
[0007] The technical problem to be solved by the present application is to provide a method for mining potential antioxidant collagen peptides based on artificial intelligence technology to overcome the shortcomings of the prior art.
[0008] The method for mining potential antioxidant collagen peptides based on artificial intelligence technology according to the present application comprises the following steps,
[0009] Step one, collect anti-oxidative collagen peptide sequences;
[0010] Step two, construct the positive sample polypeptide sequence set and the negative sample polypeptide sequence set required for training the anti-oxidative collagen peptide sequence binary classification discrimination model;
[0011] Step three, use the ProtT5 protein language model in the ProtTrans open source package to perform feature embedding on the positive sample polypeptide sequence set and the negative sample polypeptide sequence set to obtain positive and negative sample features;
[0012] Step four, train the anti-oxidative peptide discrimination model according to the positive and negative sample features.
[0013] Step one specifically includes,
[0014] Obtain anti-oxidative collagen peptide sequences from a polypeptide database, and format the anti-oxidative collagen peptide sequences and remove duplicate peptide segments in the anti-oxidative collagen peptide sequences.
[0015] The anti-oxidative collagen peptide sequences are stored in fasta format.
[0016] All the collected anti-oxidative collagen peptide sequences are used as the positive sample polypeptide sequence set; the negative sample polypeptide sequence set is composed of non-anti-oxidative peptides, signal peptides, and artificially generated random peptides.
[0017] Step three specifically includes,
[0018] First, load the ProtT5 protein language model in the ProtTrans open source package;
[0019] Second, read the positive sample polypeptide sequence set and the negative sample polypeptide sequence set, and convert the positive sample polypeptide sequence set and the negative sample polypeptide sequence set into a polypeptide dictionary containing protein sequences; replace the amino acids with a frequency less than a predetermined threshold in the protein sequences in the polypeptide dictionary with a uniform sequence symbol;
[0020] Third, input a set amount of polypeptide sequences into the ProtT5 protein language model to perform protein sequence coding, so as to convert the positive sample polypeptide sequence set and the negative sample polypeptide sequence set into Pytorch tensors, and obtain positive and negative sample features;
[0021] Fourth, store the positive and negative sample features in a uniform dictionary format in the polypeptide dictionary.
[0022] In the first step, the ProtT5 protein language model is loaded, and the vocabulary of the ProtT5 protein language model is obtained.
[0023] If the protein mean pooling embedding of each positive sample polypeptide sequence set and negative sample polypeptide sequence set is specified in the protein sequence encoding, a one-dimensional vector of a uniform length is output for each polypeptide sequence, otherwise a two-dimensional vector of a uniform dimension is output.
[0024] In step four, at least three binary classification discriminant models are trained as antioxidant peptide discriminant models according to the positive and negative sample features, and the binary classification discriminant models are trained by the Keras package.
[0025] Advantages
[0026] The advantages of the present application are:
[0027] 1. In constructing the polypeptide sequence set, the present application first introduces signal peptides and artificially generated random peptides as negative samples for training. These two peptide libraries as negative samples of the antioxidant peptide classification model can improve the performance of the model as a classifier, which is manifested as the ROC curve approaching 1. This method of constructing high-quality negative sample data sets can be widely applied to other polypeptide classification tasks to improve the accuracy and generalization ability of the classification model.
[0028] 2. In the feature embedding of the polypeptide sequence set, the present application first introduces the ProtTrans technology for feature embedding of the antioxidant peptide classification model. By pre-training the polypeptide sequence set through the ProtT5 protein language model, useful features can be learned from a large amount of protein sequence data, including physical and chemical properties, structure, conservation, etc. of the polypeptide sequence.
[0029] 3. The present application uses three trained binary classification discriminant models as antioxidant peptide discriminant models, which show high accuracy on a collagen hydrolysate antioxidant peptide data set independent of the training data set source, and can directly mine and screen new antioxidant peptides from unknown function short peptide databases. BRIEF DESCRIPTION OF DRAWINGS
[0030] Figure 1 The flowchart of the present application based on artificial intelligence technology for mining potential antioxidant collagen peptide method;
[0031] Figure 2 The flowchart of the present application for collecting antioxidant collagen peptide sequences;
[0032] Figure 3 The schematic diagram of the antioxidant peptide discriminant model training data set structure of the present application;
[0033] Figure 4A schematic diagram of a neural network architecture of the antioxidant collagen peptide sequence binary classification deep learning model of the present application;
[0034] Figure 5 A schematic diagram of the ROC curve of the antioxidant collagen peptide sequence binary classification model of the present application on the training set / test set;
[0035] Figure 6 A schematic diagram of the process of virtual screening of antioxidant collagen peptides from animal collagen sources of the present application. DETAILED DESCRIPTION
[0036] The present application will be further described below in conjunction with examples, but does not constitute any limitation on the present application, and any limited number of modifications made by anyone within the scope of the claims of the present application is still within the scope of the claims of the present application.
[0037] Reference Figures 1-6 The method for mining potential antioxidant collagen peptides based on artificial intelligence technology of the present application comprises the following steps.
[0038] Step one, collect antioxidant collagen peptide sequences. Antioxidant collagen peptide sequences are mainly collected from public data set sources and published literature. As shown in Figure 2 The present application collects a total of 1066 antioxidant collagen peptide sequences, and the channels for obtaining them can be searching for public polypeptide databases including APD3, PlantPepDB, biopep, and using crawler technology to grab polypeptide sequences annotated as having antioxidant function from these peptide databases; or downloading antioxidant collagen peptide sequences in the AnOxPePred training data set from the open source community.
[0039] After collecting antioxidant collagen peptide sequences, format the antioxidant collagen peptide sequences, which can be stored in fasta format, and remove duplicate peptide segments in the antioxidant collagen peptide sequences.
[0040] Step two, construct a positive sample polypeptide sequence set and a negative sample polypeptide sequence set.
[0041] As shown in Figure 3 The training data set sample size for training the antioxidant peptide discrimination model of the present application is 6990. The training set is divided into two parts, including a set of 1066 short peptide positive sample polypeptide sequence set and a set of negative sample polypeptide sequence set. Among them, all the antioxidant collagen peptide sequences that can be collected are used as the positive sample polypeptide sequence set. The negative sample polypeptide sequence set is composed of three subsets, namely non-antioxidant peptides, signal peptides and artificially generated random peptides.
[0042] As for the negative sample polypeptide sequence set, the first subset can be collected from the non-antioxidant peptides (a total of 242) collected by AnOxPePred and several published literatures, which do not exhibit activity in the characterization experiment of antioxidant peptide efficacy evaluation for free radical scavenging or / and metal ion chelation. The second subset can be collected from 4616 signal peptides in the training data of signalp-6.0 open source. As short peptides, signal peptides are mainly used to guide the newly synthesized proteins to the appropriate position in the cell. These peptides are usually cut off after the protein reaches the present location. Based on the main function of signal peptides to guide the protein transport process, we believe that they usually do not have antioxidant potential by themselves, so signal peptides are used as training negative samples for the antioxidant peptide binary classification discrimination model. The third subset is a total of 1066 artificially generated random peptides. This random peptide library is generated according to the 20 amino acid distribution of the protein sequence set of UniProtKB / Swiss-Prot, and the length of the random peptide library is 25 aa. Randomly generated polypeptide sequences generally do not have specific biological activity functions (i.e. antioxidant activity), and from the perspective of evolution, most of the sequences of bioactive peptides are highly optimized sequences, which are unlikely to be randomly distributed, i.e. random polypeptide sequences lack evolutionary optimization. In addition, the random peptide library can be used to construct a unified background model, which can enable the target model to be trained (i.e. the antioxidant peptide discrimination model) to learn the background patterns universally existing in random peptide sequences, filter out background patterns without biological activity and focus on detecting real biological activity patterns, which helps to reduce the false positive of the model, avoid overfitting and enhance the generalization ability of the model.
[0043] In this step, the format of the constructed positive sample polypeptide sequence set and negative sample polypeptide sequence set is unified to fasta format.
[0044] Step three, using ProtT5 protein language model in ProtTrans open source package to perform feature embedding on the positive sample polypeptide sequence set and the negative sample polypeptide sequence set to obtain positive and negative sample features.
[0045] The ProtT5 protein language model is a derivative model based on T5-3b, which is pre-trained on the UniRef50 protein sequence library in a self-supervised manner. In order to batch encode the collected positive / negative sample polypeptide sequence set into a vector, it is necessary to convert the positive / negative sample polypeptide sequence set from fasta format to vector and store it as h5 format, which is convenient for model training. After embedding in this way, a polypeptide containing N (N is a natural number) amino acid residues is encoded into an N*1024-dimensional vector.
[0046] The specific feature embedding process is as follows.
[0047] The first step is to load the ProtT5 protein language model in the ProtTrans open source package and obtain the vocabulary of the ProtT5 protein language model.
[0048] The second step is to read the positive sample polypeptide sequence set and the negative sample polypeptide sequence set in fasta format, and convert the positive sample polypeptide sequence set and the negative sample polypeptide sequence set into a polypeptide dictionary containing protein sequences. Replace the amino acids with a frequency less than a predetermined threshold in the protein sequences in the polypeptide dictionary with a uniform sequence symbol. For example, the infrequent amino acids (with a relative frequency less than 20%) can be replaced with "X" to optimize the polypeptide dictionary and reduce the data volume of the polypeptide dictionary.
[0049] The third step is to input the positive sample polypeptide sequence set and the negative sample polypeptide sequence set into the ProtT5 protein language model with a set amount of polypeptide sequences for protein sequence encoding, so as to convert the positive sample polypeptide sequence set and the negative sample polypeptide sequence set into Pytorch tensors to obtain positive and negative sample features.
[0050] The fourth step is to store the positive and negative sample features in a unified dictionary format in the polypeptide dictionary. If the protein mean pooling embedding of each positive sample polypeptide sequence set and negative sample polypeptide sequence set is specified in the protein sequence encoding, a one-dimensional vector with a uniform length of 1024 is output for each polypeptide sequence, otherwise a two-dimensional vector with a uniform dimension of (20, 1024) is output. And these positive and negative sample features are saved in an h5 format file.
[0051] Step four, training an antioxidant peptide discrimination model according to the positive and negative sample features.
[0052] The antioxidant peptide discrimination model uses a binary classification deep learning model in the prior art. As shown in Figure 4 , which shows the neural network architecture for training the antioxidant collagen peptide sequence binary classification deep learning model. According to the independently constructed positive / negative sample polypeptide sequence set, the present application has trained three binary classification discrimination models (see Table 1 for model summary). Each binary classification discrimination model uses the same neural network, and the model size is 110,028 parameters. The binary classification discrimination model is trained by the Keras package, and the training set / test set ratio used by each binary classification discrimination model is 8:2. Because the polypeptides in the training set (i.e. the polypeptide sequence set) are of different lengths, in the input part of the binary classification discrimination model, all polypeptide encodings can be unified into the format of 20*1024 by using the strategy of zero padding and truncating. As shown in Figure 5 , which shows the ROC curves of the three binary classification discrimination models on their respective training sets / test sets.
[0053] Table 1: Summary of antioxidant peptide binary classification discrimination model
[0054] Model Positive training set Negative training set Model size Test set ROC Test set MCC Model one Antioxidant peptides Non-antioxidant peptides 110,028 0.9151 0.6421 Model two Antioxidant peptides Signal peptides 110,028 1.0000 0.9942 Model three Antioxidant peptides Random peptides 110,028 0.9941 0.9766
[0055] Example One
[0056] The three binary classification deep learning models were tested on an independent set of positive polypeptide sequences to evaluate their accuracy in identifying antioxidant peptides. This independent test set contained 200 antioxidant collagen peptide sequences, of which 175 short peptides were from the literature collected on animal collagen hydrolysates. Notably, these antioxidant collagen peptide sequences were not present in the data sets used to train and test the binary classification deep learning models, so this independent test set can be considered a new challenge to test the generalization ability of these models in the real world. After testing the antioxidant sequences of the 200 independent data set samples, as shown in Table 2, it can be seen that the accuracy of the three binary classification deep learning models in identifying antioxidant peptides was as high as 99%, with only one short peptide KGEKGVPGSPGFPG having an average score below 0.5. This high accuracy indicates that the trained antioxidant peptide discrimination model can effectively identify polypeptides with antioxidant properties from unseen data, which is of great significance for predicting and developing new antioxidant peptides in practical applications.
[0057] Table 2: Performance of antioxidant peptide discrimination model in discriminating 200 antioxidant peptide sequences from an independent data set
[0058] sequence model1_score model2_score model3_score mean_score 1 VWYA 0.9251 0.9895 0.9391 0.9512 2 FFSGPNGFQ 0.9231 0.9894 0.9409 0.9511 3 HWYD 0.925 0.9894 0.9387 0.951 4 GPPGPPGPPGPPG 0.924 0.9868 0.9423 0.951 5 YHW 0.9247 0.9896 0.9387 0.951 6 GPRGPPGPVGP 0.9249 0.9877 0.9402 0.951
[0059]
[0060]
[0061]
[0062]
[0063]
[0064] Example Two
[0065] Based on the three trained antioxidant peptide discrimination models, an animal collagen short peptide library was constructed and virtual screening was performed using the trained antioxidant peptide discrimination model. Figure 6 The process of constructing the collagen short peptide library and screening potential antioxidant collagen peptides from the collagen short peptide library is shown. The virtual screening is divided into four steps:
[0066] Firstly, we collected 640 animal collagen sequences from NCBI RefSeq database, including pig, cow, chicken, Atlantic salmon, Indian tilapia, tuna, Nile tilapia, grass carp, and large yellow croaker.
[0067] Secondly, we constructed an animal collagen peptide library using the collagen sequence library. The peptide library contained 5,613,258 samples. We processed a FASTA file containing 640 protein sequences into a polypeptide sequence library. The length of each polypeptide in the output polypeptide library was between 5 and 10 amino acid residues. The specific processing process was as follows: First, we read all 640 protein sequences stored in a FASTA file through the SeqIO module in the open-source Bio Python library. Next, we defined an "extract_peptide" function, which took a protein sequence as input and extracted all peptides with a length of 5-10 aa by traversing all possible polypeptide start positions in the specified protein sequence. Finally, we generated all possible polypeptide sequences for each protein sequence by using the "extract_peptide" function to extract polypeptide sequences from each protein sequence. The program output the corresponding polypeptide sequence set for each protein sequence as 640 FASTA files.
[0068] Thirdly, we used the three trained antioxidant peptide discrimination models to virtually screen the peptide library and output the scores of each model for each short peptide as an antioxidant peptide.
[0069] Fourthly, we output 100 potential antioxidant collagen peptides based on the average scores of the three models (see Table III). It is important to note that the sequences of these potential peptides are not included in the model training set, the model test set, or the independent test set of Example I.
[0070] Table III: Potential antioxidant collagen peptides obtained by virtual screening using antioxidant peptide models
[0071]
[0072]
[0073]
[0074] The antioxidant peptide discrimination model was used to virtually screen a large library of collagen peptides, successfully identifying 100 peptide segments with potential antioxidant activity. These peptide segments were directly derived from animal collagen sequences, and the peptide segments could bypass enzymatic digestion and be directly synthesized by chemical methods and tested for antioxidant activity using the DPPH method. This strategy of directly using protein sequences to virtually screen antioxidant collagen peptides is helpful in finding new antioxidant collagen peptides for many practical application scenarios: drug development (such as developing anti-aging and anti-inflammatory drugs), health product and cosmetic development (such as enhancing the body's antioxidant capacity and improving skin firmness and moisture). This screening method has great application potential, as it is directly based on short peptide sequences to determine their antioxidant activity, and therefore can be used as an efficient and low-cost computational means for screening large peptide libraries.
[0075] The above only describes the preferred embodiments of the present application, and it should be noted that for those skilled in the art, without departing from the structure of the present application, a number of modifications and improvements can be made, which will not affect the effect of the implementation of the present application and the practicality of the patent.
Claims
1. A method for mining potential antioxidant collagen peptides based on artificial intelligence technology, characterized in that, Includes the following steps, Step 1: Collect antioxidant collagen peptide sequences; Step 2: Construct the positive sample peptide sequence set and the negative sample peptide sequence set required for training the antioxidant collagen peptide sequence binary classification discrimination model; Step 3: Use the ProtT5 protein language model in the ProtTrans open source package to embed features into the positive and negative sample polypeptide sequence sets to obtain positive and negative sample features; Step 4: Train the antioxidant peptide discrimination model based on the positive and negative sample features; Step three specifically includes, Step 1: Load the ProtT5 protein language model from the ProtTrans open-source package; The second step is to read the positive sample polypeptide sequence set and the negative sample polypeptide sequence set, and convert the positive sample polypeptide sequence set and the negative sample polypeptide sequence set into a polypeptide dictionary containing protein sequences; replace amino acids in the protein sequences of the polypeptide dictionary with a frequency probability of less than a prediction threshold with uniform sequence symbols. The third step involves inputting the positive and negative sample polypeptide sequence sets into the ProtT5 protein language model using a set amount of polypeptide sequences for protein sequence encoding, so as to convert the positive and negative sample polypeptide sequence sets into PyTorch tensors and obtain positive and negative sample features. The fourth step is to store the positive and negative sample features into the polypeptide dictionary in a unified dictionary format.
2. The method for mining potential antioxidant collagen peptides based on artificial intelligence technology according to claim 1, characterized in that, Step one specifically includes, Antioxidant collagen peptide sequences are obtained from a peptide database and then formatted to remove repetitive peptide segments.
3. The method for mining potential antioxidant collagen peptides based on artificial intelligence technology according to claim 2, characterized in that, The antioxidant collagen peptide sequence is stored in FASTA format.
4. The method for mining potential antioxidant collagen peptides based on artificial intelligence technology according to claim 1, characterized in that, All the collected antioxidant collagen peptide sequences were used as the positive sample peptide sequence set; the negative sample peptide sequence set consisted of non-antioxidant peptides, signal peptides, and artificially generated random peptides.
5. The method for mining potential antioxidant collagen peptides based on artificial intelligence technology according to claim 1, characterized in that, In the first step above, while loading the ProtT5 protein language model, the vocabulary of the ProtT5 protein language model is obtained.
6. The method for mining potential antioxidant collagen peptides based on artificial intelligence technology according to claim 4, characterized in that, The unified dictionary format is as follows: if the protein sequence encoding specifies the calculation of protein mean pooling embedding for each of the positive and negative sample polypeptide sequence sets, then a one-dimensional vector of uniform length is output for each polypeptide sequence; otherwise, a two-dimensional vector of uniform dimension is output.
7. The method for mining potential antioxidant collagen peptides based on artificial intelligence technology according to claim 1, characterized in that, In step four, at least three binary classification models are trained based on the positive and negative sample features to serve as antioxidant peptide discrimination models, and the binary classification models are trained using the Keras package.
Citation Information
Patent Citations
Method for predicting antibacterial peptides of lactic acid bacteria based on graph neural network
CN113571133A
Prediction method for mining protein interaction type based on deep learning
CN115588463A