Method for generating negative samples based on relationship between amino acids and source proteins
Patent Information
- Application Number
- TW114105123
- Authority / Receiving Office
- TW · TW
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2026-08-16
- Estimated Expiration
- 2045-02-11
AI Technical Summary
Current methods for generating negative samples for protease cleavage prediction models are limited by the lack of automated generation of high-quality negative samples, relying on manual creation and using only positive training data, which affects the accuracy of peptide linker design in personalized mRNA cancer vaccines.
A method is developed to generate negative samples based on the relationship between amino acids and source proteins, utilizing characteristic indicators like cleavage saturation, peptide coverage, and peptide density, integrating multiple databases to calculate amino acid cleavage scores, and selecting high-quality candidate sequences for negative samples.
Enhances the performance of protease cleavage prediction models by improving the accuracy of peptide linker design, ensuring precise cleavage of mRNA vaccines by proteases post-administration.
Smart Images

Figure TWG2TA001072238_001 
Figure TWG2TA001072238_002 
Figure TWG2TA001072238_003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for generating negative samples, and more particularly to a method for generating negative samples based on the relationship between amino acids and source proteins. [Previous Technology]
[0002] The design of personalized mRNA cancer vaccines includes aspects such as the prediction of tumor neoantigens, the selection of peptide linkers, and the combination of mRNA sequences. The purpose of peptide linker design is to ensure that the mRNA vaccine can be accurately cleaved by proteases after administration to the human body. Currently, peptide linkers are selected after probability ranking and conditional selection using a cleavage prediction model. Therefore, the quality of the cleavage prediction model indirectly affects the design of the peptide linkers. During the training process of the cleavage prediction model, since all training data are positive samples, negative samples need to be generated manually. Therefore, in addition to exploring the relationship between peptides and protein sequences, this invention further explores the cleavage relationship through the characteristics of amino acids and integrates the idea of multiple databases to design a method for generating negative samples. This enhances the performance of the cleavage prediction model and further optimizes the design of peptide linkers. [Summary of the Invention]
[0003] This invention provides a method for generating negative samples based on the relationship between amino acids and source proteins, in order to enhance the performance of protease cleavage prediction models and further optimize the design of peptide linkers.
[0004] A method for generating negative samples based on the relationship between amino acids and source proteins according to the present invention includes the following steps: In the protein sequence, peptides corresponding to the protein sequence in multiple databases are identified. Then, characteristic indicators are calculated based on the relationship between the peptides and the protein sequence, and high-quality protein sequences are screened out based on these characteristic indicators. Next, among the high-quality protein sequences, the non-cleavage positions of the protein sequence and peptides are used as candidate sequences for negative samples. Then, based on the cleavage positions of the peptides and protein sequences, an amino acid cleavage frequency table is generated. Finally, the amino acid cleavage scores of the candidate negative sample sequences are calculated, and multiple databases are integrated to generate negative samples.
[0005] In one embodiment of the present invention, the characteristic indicators include cleavage saturation, peptide coverage and peptide density.
[0006] In one embodiment of the present invention, the cleavage saturation is used to measure the consistency of cleavage on a protein sequence.
[0007] In one embodiment of the present invention, peptide coverage is used to measure the extent to which peptides cover protein sequences.
[0008] In one embodiment of the present invention, peptide density is used to measure the likelihood that a protein sequence will be cleaved.
[0009] In one embodiment of the present invention, the formula for calculating the amino acid cleavage score of the negative sample candidate sequence is as follows: where Sn is the score of the i-th position of the sequence in the cleavage frequency table, n is the sequence length, WDB_N is the weight of database N, and N is the Nth database.
[0010] Based on the above, the present invention provides a method for generating negative samples based on the relationship between amino acids and source proteins. This method not only considers the integration of amino acid cleavage frequency characteristics, but also the integration of multiple databases to find negative samples, in order to improve the actual proteasome mechanism of action and enhance the performance of protease cleavage prediction models.
Implementation Method
[0011] The following embodiments are described in detail with reference to the accompanying drawings, but the embodiments provided are not intended to limit the scope of this disclosure. In addition, the terms "comprising," "including," "having," etc., used in this document are open-ended terms, meaning "comprising but not limited to."
[0012] Figure 1 is a schematic diagram of a conventional method for generating negative samples. According to biological experimental results, the yellow area represents peptides, and the cleavage point is located at the C-terminus (in Figure 1, the cleavage point is S). The C-terminus is defined as a positive sample. Currently, there are two ways to generate negative samples: (1) Exclude the C-terminal cleavage point (the cleavage point is a positive sample) from the peptide (yellow area) region as a possible negative sample. (2) Screen for highly saturated protein sequences by defining saturation, and find negative samples in the regions of these protein sequences that are not covered by peptides. Furthermore, consider the distribution of peptides and source proteins to define representative regions, and find negative samples in these regions.
[0013] Furthermore, since peptides are formed through the degradation of the proteasome; for example, the proteasome participates in the MHC class I pathway and the lysosome participates in the MHC class II pathway. Current research has observed that the proteasome tends to cleave amino acids with specific properties, for example, the β5 unit of the 20S proteasome tends to cleave hydrophobic amino acids; however, the exact mechanisms and properties remain uncertain. Therefore, in addition to finding negative samples from the relationship between peptides and their source proteins, incorporating the amino acid hierarchy into the search for negative samples may provide an opportunity to simulate the actual proteasome mechanism of action, thereby enhancing the effectiveness of protease cleavage prediction models.
[0014] Currently, the sources of peptides (positive samples) are not entirely the same. Some come from in vitro cell experiments, while others come from in vivo experiments. In vivo experiments can also have species-specific differences. Therefore, this invention not only considers the integration of amino acid cleavage frequency characteristics, but also the integration of multiple databases to find negative samples, hoping to improve the actual proteasome mechanism of action and enhance the efficiency of protease cleavage prediction models.
[0015] Figure 2 is a schematic diagram of a method for generating negative samples based on the relationship between amino acids and source proteins according to an embodiment of the present invention. Figure 3 is a schematic diagram of generating candidate sequences for negative samples according to an embodiment of the present invention.
[0016] Referring to Figure 2, the present invention provides a method for generating negative samples based on the relationship between amino acids and source proteins, comprising the following steps: In the protein sequence, peptides corresponding to the protein sequence in multiple databases are identified. Then, based on the relationship between the peptide and the protein sequence, characteristic indicators are calculated, and high-quality protein sequences are screened out using these characteristic indicators. Next, among the high-quality protein sequences, the non-cleavage positions of the protein sequence and peptides are used as candidate sequences for negative samples. Then, based on the cleavage positions of the peptide and the protein sequence, an amino acid cleavage frequency table is generated. Finally, the amino acid cleavage scores of the candidate negative sample sequences are calculated, and multiple databases are integrated to generate negative samples.
[0017] In this embodiment, the characteristic indicators include cleavage saturation, peptide coverage, and peptide density. Cleavage saturation measures the consistency of cleavage on the protein sequence. Peptide coverage measures the degree to which peptides cover the protein sequence. Peptide density measures the likelihood that the protein sequence will be cleaved. The database refers to a protein reference protein, such as UniProt or SwissProt. Finding peptides that correspond to protein sequences in multiple databases means, given the existence of peptides, mapping them back to the protein sequence and using the gene name and accession name to match the source protein sequence (UniProt or SwissProt) in the database. High-quality protein sequences refer to those that, after screening by characteristic indicators, possess high reliability and relatively clear functional and structural information.
[0018] Refer to Figures 2 and 3. In high-quality protein sequences, the non-cleavage positions of the protein sequence and peptides are used as negative sample candidate sequences, as shown in the blue area of Figure 3. There are four negative sample candidate sequences, such as: {MSET}, {PAEA}, {PAPVEK}, and {KKK}.
[0019] In this embodiment, a cleavage frequency table for 20 amino acids is generated based on the cleavage positions of peptides and protein sequences. Taking Figure 3 as an example, there are five peptides corresponding to this protein sequence. The white background represents the cleavage positions, namely A, T, S, and K. For example, among the five peptides, including the N-terminus and C-terminus, there are a total of 10 cleavage positions. The cleavage position is A, which appears 5 times, with a score of 5 / 10. The occurrence frequency of all 20 amino acids is calculated, and the cleavage frequency table is shown in Table 1. Table 1 amino acids frequency Score (totaling 1) A 5 0.5 C 0 0 D 0 0 E 0 F 0 0 G 0 0 H 0 0 I 0 0 K 1 0.1 L 0 0 M 0 0 N 0 0 P 1 0.1 Q 0 0 R 0 0 S 1 0.1 T 2 0.2 V 0 0 W 0 0 Y 0 0
[0020] In this embodiment, the formula for calculating the amino acid cleavage score of the negative sample candidate sequence is as follows: where Sn is the score of the i-th position of the sequence in the cleavage frequency table, n is the sequence length, WDB_N is the weight of database N, the weight can be based on the amount of data or the same weight, and N is the Nth database.
[0021] Taking the negative sample candidate sequences generated in this embodiment as an example, in the case of only one database, the amino acid fragmentation score of {MSET} = (0 + 0.1 + 0 + 0.2) / 4 = 0.075, and the amino acid fragmentation score of {PAET} = (0 + 0.5 + 0 + 0.2) / 4 = 0.175. The lower the score, the more suitable it is as a negative sample.
[0022] If there are two databases, the databases may be of different ethnicities or different laboratories, etc. WDB_N can be weighted according to the amount of data. For example, if database 1 has 10,000 data entries and database 2 has 5,000 data entries, then WDB_1=10,000 / (10,000+5,000)=0.67, WDB_=5,000 / (10,000+5,000)=0.33.
[0023] According to the method of generating negative samples based on the relationship between amino acids and source proteins of the present invention, negative samples can be generated. The application of the generated negative samples is as follows: the generated negative samples are sorted according to the amino acid cleavage fraction, and then, according to the desired ratio of positive to negative samples, such as positive samples: negative samples = 1:1, the top N negative samples and N positive samples with higher quality are selected and added together to establish a protease cleavage prediction model to enhance the efficiency of the protease cleavage prediction model and further optimize the design of peptide linkers. In this way, it can be further used in the design of personalized mRNA cancer vaccines and the selection of peptide linkers, so that the mRNA vaccine can be accurately cleaved by proteases after being administered to the human body.
[0024] The following compares different processes that generate negative samples using an experimental protease cleavage model. Table 1 Workflow Sensitivity Specificity PPV (Accuracy) NPV (Recall rate) accuracy TP FP TN FN AUC pepsickle decoy generation 0.65 0.6 0.617 0.63 0.62 14950 9265 13806 8121 0.67 Invention process 0.71 0.59 0.632 0.67 0.65 16396 9539 13532 6675 0.71 pepCleave decoy generation 0.69 0.61 0.638 0.66 0.65 15998 9096 13975 7073 0.71
[0025] In summary, this invention proposes a method for generating negative samples based on the relationship between amino acids and source proteins. This method not only considers the integration of amino acid cleavage frequency characteristics but also the integration of multiple databases to find negative samples. The aim is to improve the actual proteasome mechanism of action, enhance the efficiency of protease cleavage prediction models, and further optimize peptide linker design. This can then be further applied to the design of personalized mRNA cancer vaccines and the selection of peptide linkers, ensuring that the mRNA vaccine is accurately cleaved by proteases after administration to the human body. [Simplified Explanation of the Diagram]
[0026] Figure 1 is a schematic diagram of a conventional method for generating negative samples. Figure 2 is a schematic diagram of a method for generating negative samples based on the relationship between amino acids and source proteins according to an embodiment of the present invention. Figure 3 is a schematic diagram of generating candidate sequences for negative samples according to an embodiment of the present invention.
Claims
1. A method for generating negative samples based on the relationship between amino acids and source proteins, comprising: In the protein sequence, identify the peptides that correspond to the protein sequence in multiple databases; By analyzing the relationship between the peptide and the protein sequence, characteristic indicators are calculated, and high-quality protein sequences are screened out based on these indicators. Among the high-quality protein sequences, the non-cleavage positions of the protein sequence and the peptide are used as negative sample candidate sequences. Based on the cleavage positions of the peptide and the protein sequence, an amino acid cleavage frequency table is generated. The amino acid cleavage scores of the negative sample candidate sequences are calculated, and multiple databases are integrated to generate negative samples.
2. The method for generating negative samples based on the relationship between amino acids and source proteins as described in claim 1, wherein the characteristic indicators include cleavage saturation, peptide coverage, and peptide density.
3. The method for generating negative samples based on the relationship between amino acids and source proteins as described in claim 2, wherein the cleavage saturation is used to measure the consistency of cleavage on the protein sequence.
4. The method for generating negative samples based on the relationship between amino acids and source proteins as described in claim 2, wherein the peptide coverage is used to measure the extent to which the peptide covers the protein sequence.
5. The method for generating negative samples based on the relationship between amino acids and source proteins as described in claim 2, wherein the peptide density is used to measure the likelihood that the protein sequence is cleaved.
6. The method for generating negative samples based on the relationship between amino acids and source proteins as described in claim 1, wherein the formula for calculating the amino acid cleavage fraction of the negative sample candidate sequence is as follows: where, Sn is the score of the i-th position of the sequence in the cutoff frequency table, n is the sequence length, WDB_N is the weight of database N, and N is the Nth database.