Method and device for reversely screening special protease based on target active peptide

By establishing a target and negative peptide database and combining virtual enzyme cleavage technology, suitable proteases were screened out, which solved the problem of insufficient screening efficiency and accuracy of proteases in the existing methods, and achieved efficient and accurate protease screening.

CN120356525APending Publication Date: 2025-07-22SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510351377.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing protease screening methods rely on experimental trial and error, which consumes a lot of time and resources. The virtual enzyme cutting technology cannot reversely screen adapted proteases from the target peptide, resulting in insufficient screening efficiency and accuracy.

Method used

Establish a target peptide database and a negative peptide database, obtain theoretical enzyme digestion peptide segments and peptide information through virtual enzyme cleavage technology, combine negative peptide information to screen candidate proteases, exclude enzymes that produce negative peptides, and give priority to the target peptide.

Benefits of technology

It improves the accuracy and efficiency of protease screening, ensures that the screened protease can produce target peptides and avoid adverse by-products, and improves the screening accuracy of proteases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356525A_ABST
    Figure CN120356525A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for reversely screening special protease based on a target active peptide, which are applied to the field of bioactive peptides, and are characterized in that a target peptide database and a negative peptide database are established based on task requirements, so that expected active peptides and peptides with adverse effects can be clearly distinguished in the screening process, thereby ensuring that the target is clear and the screening efficiency is high. And virtual enzyme digestion is carried out on the target protein data based on a preset enzyme digestion rule to obtain a first theoretical enzyme digestion peptide fragment and corresponding target peptide information, so that the applicability of each protease can be accurately evaluated directly according to the quantity and quality of the target peptide generated by each protease in the screening process, and the screening efficiency is improved. Therefore, the accuracy of the screening process is improved; the candidate protease set is further screened in combination with the negative peptide information, and the enzyme generating the negative peptide is excluded, so that the screened protease can preferentially generate the target peptide, adverse byproducts are avoided, and the screening precision of the protease is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioactive peptides, and particularly to a method and device for reverse screening of a dedicated protease based on a target bioactive peptide. Background Art

[0002] At present, with the wide application of bioactive peptides in fields such as food and medicine, their efficient preparation technology has become a research hotspot. The current mainstream preparation methods can be divided into two categories: chemical synthesis methods (such as solid-phase synthesis) can precisely control the peptide sequence, but have problems such as high cost and by-product pollution, and it is difficult to meet the large-scale industrial demand; biological preparation methods (such as enzymatic hydrolysis method, microbial fermentation method) have become the mainstream due to advantages such as low cost and mild conditions. Among them, the enzymatic hydrolysis method releases bioactive peptides by specifically cleaving proteins with proteases, which is the core process for the production of natural bioactive peptides. In the enzymatic hydrolysis process, the selection of proteases is the key link determining the peptide composition and functional activity of the final product.

[0003] Existing protease screening methods rely on experimental trial and error to screen proteases, which requires a large amount of time and resources. The virtual digestion technology provides a theoretical basis for protease optimization by simulating the digestion sites and products of target proteins by computer, and can reduce the R & D cost. However, the existing virtual digestion technology supports the one-way prediction of "from proteins and proteases to digestion sites" and cannot reverse-screen the suitable proteases and protein sources starting from the target peptide segments. Summary of the Invention

[0004] The present invention provides a method and device for reverse screening of a dedicated protease based on a target bioactive peptide to improve the accuracy and efficiency of protease screening.

[0005] To solve the above technical problems, an embodiment of the present invention provides a method for reverse screening of a dedicated protease based on a target bioactive peptide, including:

[0006] Establishing a target peptide database and a negative peptide database based on task requirements;

[0007] Obtaining the amino acid sequence of a protein source, screening in the amino acid sequence based on the target peptide database to obtain target protein data;

[0008] Performing virtual digestion on the target protein data based on the preset digestion rules of proteases to obtain the first theoretical digestion peptide segments of each protease and the corresponding target peptide information, and obtaining a candidate protease set based on the target peptide segment information;

[0009] Performing virtual digestion on the amino acid sequence of the protein source based on the candidate protease set and the digestion rules to obtain the second theoretical digestion peptides of the candidate protease set, and obtaining the negative peptide information of the second theoretical digestion peptides based on the negative peptide database;

[0010] Screening the candidate protease set based on the negative peptide information and the target peptide information to obtain the target protease.

[0011] In the present invention, by establishing a target peptide database and a negative peptide database based on task requirements, it is possible to clearly distinguish the active peptides expected to be obtained from the peptides with adverse effects during the screening process, thereby ensuring a clear goal. Furthermore, by performing virtual digestion on the target protein data based on the preset digestion rules, the first theoretical digestion peptides and their corresponding target peptide information are obtained, enabling the direct evaluation of the applicability of each protease based on the quantity and quality of the target peptides produced by each protease during the screening process, thus improving the accuracy of the screening process; further screening the candidate protease set in combination with the negative peptide information, excluding the enzymes that produce negative peptides, thereby ensuring that the screened proteases can preferentially produce target peptides, avoiding adverse by-products, and enhancing the screening accuracy of proteases.

[0012] Furthermore, the target peptide database includes target peptide segments; the method for obtaining the amino acid sequence of the protein source and screening the amino acid sequence based on the target peptide database to obtain the target protein data includes:

[0013] Obtaining the amino acid sequence of the protein source based on the protein database and screening the amino acid sequence based on the target peptide segments;

[0014] Obtaining the target protein data, where the target protein data includes the target protein and the corresponding target peptide data; the target protein is the amino acid sequence containing the target peptide segment.

[0015] In the present invention, by clarifying the target peptide segments and screening in the protein database, the protein sequence containing the target peptide segments can be accurately located. By taking the target protein and the target peptide data as the output, it is ensured that the screened protein data has a clear structure and is easy for subsequent digestion simulation and screening.

[0016] Furthermore, the target peptide information includes the quantity of the target peptide segments; the method for performing virtual digestion on the target protein data based on the digestion rules of the preset protease to obtain the first theoretical digestion peptides of each protease and the corresponding target peptide segment information, and obtaining the candidate protease set based on the target peptide segment information includes:

[0017] Performing virtual digestion on the target protein based on the digestion rules of each protease respectively to obtain the first theoretical digestion peptides of all proteases;

[0018] Obtain the number of target peptide segments in the first theoretical digested peptide segments, and screen out several candidate proteases based on a preset quantity threshold and the number of the target peptide segments to generate a candidate protease set.

[0019] In the present invention, the first theoretical digested peptide segments are obtained through virtual digestion, so that by screening the number of target peptide segments in the first theoretical digested peptide segments of each protease, it is possible to more precisely determine which proteases are most suitable for generating the target peptide, improving the efficiency and accuracy of the screening.

[0020] Further, the negative peptide database includes negative peptide segments; based on the candidate protease set and the digestion rules, the amino acid sequence of the protein source is virtually digested to obtain the second theoretical digested peptide segments of the candidate protease set, and negative peptide information of the second theoretical digested peptide segments is obtained based on the negative peptide database, including:

[0021] Based on the candidate protease and its digestion rules, the amino acid sequence of the protein source is virtually digested to obtain the second theoretical digested peptide segments of the candidate protease;

[0022] Count the number and types of negative peptide segments in the second theoretical digested peptide segments to obtain the negative peptide information.

[0023] In the present invention, by combining the negative peptide database, the candidate proteases are further screened, and those enzymes that may produce negative peptide segments can be effectively excluded. By counting the number and types of negative peptide segments in the second theoretical digested peptide segments, a more comprehensive negative effect assessment is provided, ensuring that the screened proteases not only have a high yield of the target peptide but also avoid more negative peptides.

[0024] Further, the screening of the candidate protease set based on the negative peptide information and the target peptide information to obtain the target protease includes:

[0025] Sort the candidate protease set in descending order based on the number of the target peptide segments to obtain a first protease list;

[0026] Sort the candidate protease set in ascending order based on the number of the negative peptide segments to generate a second protease list;

[0027] Integrate the first protease list and the second protease list to obtain the target protease.

[0028] By sorting the number of target peptides and the number of negative peptides respectively, the present invention can systematically evaluate the advantages and disadvantages of each protease. By sorting the target peptides in descending order and the negative peptides in ascending order, it can ensure that the finally selected protease maximally increases the production of target peptides while minimizing the generation of negative peptides, thereby obtaining a protease of higher quality.

[0029] In a second aspect, the present invention provides a device for reverse screening of a dedicated protease based on a target active peptide, comprising: a requirement confirmation module, a target protein screening module, a first digestion module, a second digestion module, and a screening module;

[0030] The requirement confirmation module is used to establish a target peptide database and a negative peptide database based on task requirements;

[0031] The target protein screening module is used to obtain the amino acid sequence of a protein source, screen the amino acid sequence based on the target peptide database, and obtain target protein data;

[0032] The first digestion module is used to perform virtual digestion on the target protein data based on the digestion rules of a preset protease, obtain the first theoretical digestion peptides of each protease and the corresponding target peptide information, and obtain a candidate protease set based on the target peptide information;

[0033] The second digestion module is used to perform virtual digestion on the amino acid sequence of the protein source based on the candidate protease set and the digestion rules, obtain the second theoretical digestion peptides of the candidate protease set, and obtain the negative peptide information of the second theoretical digestion peptides based on the negative peptide database;

[0034] The screening module is used to screen the candidate protease set based on the negative peptide information and the target peptide information to obtain a target protease.

[0035] Further, the target protein screening module is used for:

[0036] Obtain the amino acid sequence of the protein source based on a protein database, and screen the amino acid sequence based on the target peptide segment;

[0037] Obtain target protein data, where the target protein data includes a target protein and corresponding target peptide data; the target protein is an amino acid sequence containing a target peptide segment.

[0038] Further, the first digestion module is used for:

[0039] Perform virtual digestion on the target protein respectively based on the digestion rules of each protease to obtain the first theoretical digestion peptides of all proteases;

[0040] Obtain the number of target peptide segments in the first theoretical digested peptide segments, and screen out a number of candidate proteases based on a preset number threshold and the number of the target peptide segments to generate a candidate protease set.

[0041] Further, the second digestion module is configured to:

[0042] Perform virtual digestion on the amino acid sequence of the protein source based on the candidate protease and its digestion rule to obtain the second theoretical digested peptide segments of the candidate protease;

[0043] Count the number and types of negative peptide segments in the second theoretical digested peptide segments to obtain the negative peptide information.

[0044] Further, the screening module is configured to:

[0045] Sort the candidate protease set in descending order based on the number of the target peptide segments to obtain a first protease list;

[0046] Sort the candidate protease set in ascending order based on the number of the negative peptide segments to generate a second protease list;

[0047] Integrate the first protease list and the second protease list to obtain the target protease. Description of the Drawings

[0048] Figure 1 It is a schematic flow chart of a method for reverse screening of a dedicated protease based on a target active peptide provided in an embodiment of the present invention;

[0049] Figure 2 It is a schematic diagram of an enzyme-substrate complex provided in an embodiment of the present invention;

[0050] Figure 3 It is a schematic structural diagram of a device for reverse screening of a dedicated protease based on a target active peptide provided in an embodiment of the present invention. Detailed Description of the Invention

[0051] The following further describes in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0052] In the description, claims, and drawings of this application, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.

[0053] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0054] Embodiment 1

[0055] See Figure 1 , Figure 1 is a schematic flowchart of a method for reverse screening of a dedicated protease based on a target bioactive peptide provided in an embodiment of the present invention. An embodiment of the present invention provides a method for reverse screening of a dedicated protease based on a target bioactive peptide, including steps 101 to 105, as follows:

[0056] Step 101: Establish a target peptide database and a negative peptide database based on task requirements;

[0057] In this embodiment, the target peptide refers to the peptide that is desired to be obtained by enzymatic cleavage, and the negative peptide refers to the peptide that is not desired to be cleaved. According to the application requirements of bioactive peptides (such as in the food and pharmaceutical fields), the target peptides are usually those peptide segments with specific functions or biological activities (such as antioxidant, antibacterial, antihypertensive, etc. activities), while the negative peptides refer to those peptide segments that may cause adverse reactions or are not desired to be generated in the target application (such as toxic peptides, bitter peptides, etc.).

[0058] In this embodiment, by referring to relevant literature, public databases (such as BIOPEP, PeptideAtlas, etc.) and experimental data, known bioactive peptide segments are collected. These peptide segments may possess activities such as antioxidant, antibacterial, immunomodulatory, blood pressure lowering, anti-tumor, etc. Among the collected peptide segments, those relevant to specific task requirements are further screened. For example, in the pharmaceutical field, peptide segments with anti-inflammatory, anti-cancer or antiviral activities may need to be screened; while in the food field, peptide segments with improved flavor, nutritional value or health promotion may be the focus of screening. Information such as the amino acid sequences, functional characteristics, source proteins and application fields of the screened target peptide segments is sorted out and stored in a database, thus constructing a target peptide database. Each target peptide segment should include the following information: peptide segment sequence, peptide segment function, source protein, possible biological activity, experimental verification results, etc.

[0059] In this embodiment, the amino acid sequences, sources, experimental data and their potential negative impacts (such as toxicity, bitterness, etc.) of all known negative peptide segments (such as toxic peptides, allergenic peptides, bitter peptides, etc.) are sorted out and stored in a negative peptide database. Each negative peptide segment should include the following information: peptide segment sequence, negative effect, source protein and its impact type.

[0060] In this embodiment, target peptides and negative peptides are determined according to task requirements. For example, if you want to obtain antioxidant peptides by enzymatic hydrolysis, then you first need to access such a peptide database to obtain the sequences of antioxidant peptides, and these peptide sequences are the target peptides; then during the enzymatic digestion to obtain antioxidant peptides, we do not want to cut out toxic peptides, and we also access the database to obtain the sequences of toxic peptides, and these peptide sequences are the negative peptides. Another example is that in high-protein matrix foods, if you want to make them have good flavor by enzymatic hydrolysis, then sweet peptides, umami peptides and salty peptides that have a positive effect on the senses are the target peptides, and bitter peptides that have a negative effect on the senses are the negative peptides.

[0061] Step 102: Obtain the amino acid sequence of the protein source, screen the amino acid sequence based on the target peptide database, and obtain target protein data;

[0062] In this embodiment, the target peptide database includes target peptide segments; the obtaining of the amino acid sequence of the protein source, screening the amino acid sequence based on the target peptide database, and obtaining target protein data includes:

[0063] Obtain the amino acid sequence of the protein source based on a protein database, and screen the amino acid sequence based on the target peptide segments;

[0064] Obtain target protein data, where the target protein data includes target proteins and corresponding target peptide data; the target protein is the amino acid sequence containing the target peptide segment.

[0065] In this embodiment, based on the established target peptide database, protein data containing the target peptide segments are screened from the protein source. Thus, it is determined which proteins can generate the bioactive peptides we expect through enzymatic digestion, providing data support for subsequent virtual enzymatic digestion simulation and protease screening.

[0066] In this embodiment, according to the target peptide segments defined in the target peptide database, protein sequences containing these target peptide segments are screened from the amino acid sequences of the protein source.

[0067] In this embodiment, in the amino acid sequence of the protein source, a sequence matching algorithm (such as BLAST or other customized screening tools) is used to search for the presence of target peptide segments in the protein sequence; and each peptide segment in the target peptide database is aligned with the protein sequence to check whether the target peptide segment appears and record its position, thereby obtaining target protein data, where the target protein data includes those protein sequences containing the target peptide segments. For example: Suppose the target peptide database contains the antioxidant peptide segment "Gly-Pro-Arg". When screening the soybean protein sequence, if a certain protein sequence contains this peptide segment (for example, from amino acid positions 50 to 55), then this protein will be marked as containing the target peptide segment.

[0068] In this embodiment, based on the target peptide database, protein sequences containing the target peptide segments are screened from the protein source. This process obtains the amino acid sequence of the protein source from the protein database through a sequence alignment tool, and then screens the proteins according to the target peptide segments, finally obtaining the target protein data containing the target peptides. These data provide the necessary basis for subsequent virtual enzymatic digestion, protease screening, and optimization, ensuring that peptide segments with the desired biological activity can be selected.

[0069] Step 103: Perform virtual enzymatic digestion on the target protein data based on the enzymatic digestion rules of the preset protease, obtain the first theoretical enzymatic digestion peptide segments of each protease and the corresponding target peptide information, and obtain a candidate protease set based on the target peptide information;

[0070] In this embodiment, the enzymatic digestion rules of each protease are determined according to the enzymatic cleavage sites of the protease.

[0071] Please refer to Figure 2 , Figure 2 which is a schematic diagram of an enzyme-substrate complex provided by an embodiment of the present invention.

[0072] Figure 2The enzyme-substrate complex has 8 binding sites. The residues in the N-terminal direction of the protein substrate are designated as P1, P2, P3, P4, etc., and the residues in the C-terminal direction are designated as P1', P2', P3', P4', etc. The substrate undergoing enzymatic cleavage breaks at the position between P1 and P1' according to the cleavage specificities of the proteases in Table 1 and Table 2.

[0073] Table 1

[0074]

[0075] Table 2

[0076]

[0077] Among them, Table 1 and Table 2 are schematic tables of a cleavage rule provided by an embodiment of the present invention. Among them, K is lysine, and its full English name is Lysine; T is threonine, and its full English name is Threonine; L is leucine, and its full English name is Leucine; I is isoleucine, and its full English name is Isoleucine; V is valine, and its full English name is Valine; W is tryptophan, and its full English name is Tryptophan; M is methionine, and its full English name is Methionine; F is phenylalanine, and its full English name is Phenylalanine; R is arginine, and its full English name is Arginine; H is histidine, and its full English name is Histidine; A is alanine, and its full English name is Alanine; N is asparagine, and its full English name is Asparagine; D is aspartic acid, and its full English name is Aspartic acid; C is cysteine, and its full English name is Cysteine; Q is glutamine, and its full English name is Glutamine; E is glutamate, and its full English name is Glutamate; G is glycine, and its full English name is Glycine; Y is tyrosine, and its full English name is Tyrosine; S is serine, and its full English name is Serine; P is proline, and its full English name is Proline.

[0078] In Table 1 and Table 2, Arg-C proteinase is arginine carboxyl-terminal proteinase (clostripain); Asp-N endopeptidase is aspartic acid amino-terminal endopeptidase; BNPS-Skatole is BNPS-skatole; Caspase 1 is caspase 1; Caspase 2 is caspase 2; Caspase 3 is caspase 3; Caspase 4 is caspase 4; Caspase 5 is caspase 5; Caspase 6 is caspase 6; Caspase 7 is caspase 7; Caspase 8 is caspase 8; Caspase 9 is caspase 9; Caspase 10 is caspase 10; Chymotrypsin-high specificity (C-term to [FYW], not before P) is chymotrypsin with high specificity (C-terminal is phenylalanine, tyrosine, tryptophan, and does not cleave before P); Chymotrypsin-low specificity (C-term to [FYWML], not before P) is chymotrypsin with low specificity (C-terminal is phenylalanine, tyrosine, tryptophan, leucine, methionine, and does not cleave before P); Clostripain (Clostridiopeptidase B) is clostridiopeptidase B; CNBr is cyanogen bromide; Enterokinase is enterokinase; Factor Xa is factor Xa; Formic acid is formic acid; Glutamylendopeptidase is glutamyl endopeptidase; Granzyme B is granzyme B; Hydroxylamine is hydroxylamine; Iodosobenzoic acid is iodosobenzoic acid; LysC is lysine carboxyl-terminal proteinase; Neutrophil elastase is neutrophil elastase; NTCB (2-nitro-5-thiocyanobenzoic acid) is 2-nitro-5-thiocyanobenzoic acid; Pepsin (pH 1.3) is pepsin (pH 1.3); Pepsin (pH>2) is pepsin (pH>2); Proline-endopeptidase is proline endopeptidase; Proteinase K is proteinase K; Staphylococcal peptidase I is staphylococcal peptidase I; Thermolysin is thermolysin; Thrombin is thrombin; Trypsin is trypsin.

[0079] In this embodiment, the target peptide information includes the number of target peptide segments; performing virtual digestion on the target protein data according to the digestion rules of a preset protease to obtain the first theoretical digestion peptide segments of each protease and the corresponding target peptide information, and obtaining a candidate protease set based on the target peptide information, including:

[0080] Performing virtual digestion on the target protein according to the digestion rules of each protease respectively to obtain the first theoretical digestion peptide segments of all proteases;

[0081] Obtaining the number of target peptide segments in the first theoretical digestion peptide segments, and screening out several candidate proteases based on a preset quantity threshold and the number of target peptide segments to generate a candidate protease set.

[0082] In this embodiment, virtual digestion is performed on the target protein containing the target peptide segment in the target protein data according to the digestion rules of 35 different proteases, so as to obtain the first theoretical digestion peptide segments of each protease, and the number of target peptide segments digested by each protease is recorded and counted respectively.

[0083] In this embodiment, the quantity threshold is the quantity threshold of candidate proteases. The proteases are sorted in descending order according to the number of target peptide segments, so as to select the proteases with relatively high rankings as candidate proteases according to the quantity threshold.

[0084] In another embodiment, the quantity threshold is the number of target peptide segments, and proteases whose virtual digestion produces target peptide segments exceeding the quantity threshold can be selected as candidate proteases.

[0085] In this embodiment, if the proportion of each type of protein in a certain protein source is clear, weights are calculated based on the proportion of each protein in the protein matrix, and the number of target peptides digested is calculated by weighting based on the weights, so as to make the digestion prediction result more accurate.

[0086] In this embodiment, the weights are determined according to the proportion of each type of protein in the protein matrix.

[0087] In this embodiment, the first theoretical digestion peptide segments are obtained through virtual digestion, so that by screening the number of target peptide segments in the first theoretical digestion peptide segments of each protease, it is possible to more precisely judge which proteases are most suitable for generating target peptides, improving the efficiency and accuracy of the screening.

[0088] Step 104: Performing virtual digestion on the amino acid sequence of the protein source based on the candidate protease set and the digestion rules to obtain the second theoretical digestion peptide segments of the candidate protease set, and obtaining the negative peptide information of the second theoretical digestion peptide segments based on the negative peptide database;

[0089] In this embodiment, the negative peptide database includes negative peptide segments; virtual digestion is performed on the amino acid sequence of the protein source based on the candidate protease set and the digestion rules to obtain the second theoretical digestion peptides of the candidate protease set, and negative peptide information of the second theoretical digestion peptides is obtained based on the negative peptide database, including:

[0090] Virtual digestion is performed on the amino acid sequence of the protein source based on the candidate protease and its digestion rules to obtain the second theoretical digestion peptides of the candidate protease.

[0091] The number and types of negative peptide segments in the second theoretical digestion peptides are counted to obtain the negative peptide information.

[0092] In this embodiment, virtual digestion is performed on all amino acid sequences of the protein source through the candidate protease and its corresponding digestion rules, so as to obtain the second theoretical digestion peptides of the candidate protease for the amino acid sequences in the protein source. At this time, the number and types of negative peptides in the second theoretical digestion peptides of each candidate protease are counted to judge the negative effects of each candidate protease. The more the number and types of negative peptides, the greater the negative effect of the candidate protease. At this time, it should not be used as the dedicated protease for this protein source.

[0093] In this embodiment, by combining the negative peptide database, further screening of the candidate proteases can effectively exclude those enzymes that may produce negative peptide segments. By counting the number and types of negative peptide segments in the second theoretical digestion peptides, a more comprehensive negative effect assessment is provided to ensure that the screened proteases not only have a high yield of target peptides but also avoid more negative peptides.

[0094] Step 105: Screen the candidate protease set based on the negative peptide information and the target peptide information to obtain the target protease.

[0095] In this embodiment, the screening of the candidate protease set based on the negative peptide information and the target peptide information to obtain the target protease includes:

[0096] The candidate protease set is sorted in descending order based on the number of target peptide segments to obtain the first protease list; the candidate protease set is sorted in ascending order based on the number of negative peptide segments to generate the second protease list; the first protease list and the second protease list are integrated to obtain the target protease.

[0097] In this embodiment, the candidate protease set includes several proteases. For each candidate protease, the number of target peptide segments contained in the first theoretical protease cleavage peptide segments it generates is counted. The more the number of target peptide segments, the stronger the ability of the protease to release the target active peptide. Therefore, it should be preferentially selected. The candidate proteases in the candidate protease set are sorted in descending order according to the number of target peptide segments from high to low to form the first protease list.

[0098] In this embodiment, for each candidate protease, after digesting the amino acid sequence of the protein source, the number of negative peptide segments in the second theoretical protease cleavage peptide segments generated is counted. The number of negative peptide segments produced by each candidate protease is counted. The fewer the number of negative peptide segments, the less the generation of adverse by-products while the protease generates the target peptide segments. Therefore, it should be preferentially selected. The candidate proteases are sorted in ascending order according to the number of negative peptide segments from less to more to form the second protease list.

[0099] In this embodiment, since the candidate proteases are selected from all proteases and are proteases that generate a relatively large number of target peptides through virtual digestion, therefore, the proteases ranked at the front in both sorts are comprehensively considered as the most suitable proteases as the target proteases.

[0100] In this embodiment, the comprehensive scores of each candidate protease can also be calculated by combining the first protease list and the second protease list through methods such as comprehensive ranking or weighted scoring: for example, through the weighted average method (such as setting weight values) or other evaluation models, the number of target peptide segments and the number of negative peptide segments are quantitatively scored. Several proteases with the optimal comprehensive scores are selected as the final target proteases.

[0101] In this embodiment, by establishing a target peptide database and a negative peptide database based on the task requirements, it is possible to clearly distinguish the active peptides expected to be obtained from the peptides with adverse effects during the screening process, so as to ensure a clear goal. Furthermore, through virtual digestion of the target protein data based on the preset protease cleavage rules, the first theoretical protease cleavage peptide segments and their corresponding target peptide information are obtained, enabling the direct evaluation of the applicability of each protease based on the quantity and quality of the target peptides produced by each protease during the screening process, thereby improving the accuracy of the screening process; further screening the candidate protease set in combination with the negative peptide information, excluding the proteases that produce negative peptides, thus ensuring that the selected proteases can preferentially produce target peptides, avoiding adverse by-products, and improving the screening accuracy of proteases.

[0102] Table 3

[0103] Enzyme Target Peptide Count Chymotrypsin low specificity 15240 Pepsin pH1.3 15240 Pepsin pH>2 15240 Proteinase K 15240 Thermolysin 15240 Neutrophil elastase 15221 Trypsin 15209 Chymotrypsin high specificity 15171 CNBr 15165 Glutamyl endopeptidase 15152 Staphylococcal peptidase I 15152 Arg-C proteinase 15089 Clostripain 15089 LysC 15070 Asp-N endopeptidase 15058 Formic acid 15058 NTCB 13961 BNPS-Skatole 13392 Iodosobenzoic acid 13392 Proline-endopeptidase 12123 Hydroxylamine 9694 Caspase 1 5371 Thrombin 5201 Factor Xa 743 Enterokinase 463 Caspase 7 148 Caspase 10 131 Caspase 4 123 Granzyme B 104 Caspase 8 27

[0104] Table 4

[0105] Enzyme Negative Peptide Count Staphylococcal peptidase I 0 CNBr 6 Asp-N endopeptidase 144 Glutamyl endopeptidase 172 Arg-C proteinase 535 Clostripain 535 LysC 725 Chymotrypsin high specificity 1459 Neutrophil elastase 2867 Chymotrypsin low specificity 7739 Trypsin 17632 Pepsin pH1.3 23688 Pepsin pH>2 31419 Thermolysin 48043 Proteinase K 55893

[0106] As a specific example of an embodiment of the present invention, flavor peptides with fresh, sweet, salty, and rich flavors are released by enzymatically hydrolyzing soybeans to give it a delicious taste. 432 protein sequence information of soybean species is obtained from the Uniprot database; 138 target peptides with fresh, sweet, salty, and rich flavors and 345 negative peptides with bitter taste are obtained from the flavor peptide library (BIOPEP-UWM) and the literature; based on virtual digestion, the sequences and quantities of target peptides that can be digested by each protease, and the proteins where the target peptides are located are obtained; as shown in Table 3, Table 3 is the total quantity of target peptides that can be digested by each protease; furthermore, the first several proteases are subjected to virtual digestion to obtain the quantity of negative peptides obtained by the protease, as shown in Table 4.

[0107] In this embodiment, since it is required to digest a large number of target peptides while producing fewer negative peptides, Staphylococcal peptidase I can be determined as the most suitable protease after comprehensive consideration.

[0108] As a specific example of an embodiment of the present invention, the task requirement is to enzymatically release the sweetness of sugar-free milk, and the sweet peptides are confirmed as the target peptides based on the task requirement. First, the protein sequence information of milk is obtained from the Uniprot database, and the proportions of various proteins in milk are obtained by referring to the literature. The proportions of some proteins are shown in Table 5.

[0109] Table 5

[0110]

[0111] After obtaining the proportions of various proteins in milk, 67 sweet target peptides and 345 bitter negative peptides are obtained from the flavor peptide database (BIOPEP-UWM) and the literature; the virtual digestion program is run to obtain the sequences and weighted quantities of target peptides that can be digested by each protease, and the proteins where the target peptides are located; the total weighted quantity of target peptides that can be digested by each protease is shown in Table 6; and based on the first several proteases for virtual digestion, the weighted quantity of negative peptides digested is shown in Table 7.

[0112] Table 6

[0113] Enzyme Total_Adjusted_Peptide_Count Peptides Chymotrypsin low specificity 527.58 K,V,G,P,A,AA,EV,AAA,AM,LA,DL,VPY Neutrophil elastase 527.58 K,V,G,P,A,AA,EV,AAA,AM,LA,DL,VPY Pepsin pH1.3 527.58 K,V,G,P,A,AA,EV,AAA,AM,LA,DL,VPY Pepsin pH>2 527.58 K,V,G,P,A,AA,EV,AAA,AM,LA,DL,VPY Proteinase K 527.58 K,V,G,P,A,AA,EV,AAA,AM,LA,DL,VPY Thermolysin 527.58 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY LysC 520.68 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY Trypsin 520.68 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY CNBr 511.02 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY Arg-C proteinase 506.22 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY Clostripain 506.22 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY Glutamyl endopeptidase 502.08 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY NTCB 500.22 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY Asp-N endopeptidase 480.78 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY Chymotrypsin high specificity 480.78 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY Formic acid 480.78 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY Proline-endopeptidase 473.82 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY BNPS-Skatole 466.32 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY Iodosobenzoic acid 466.32 K, V, G, P, A, AA, EV, AAA, AM, LA, DL, VPY Caspase 1 187.92 K, V, G, P, A, EV, AM, LA, VPY Hydroxylamine 24.31 K, V, G, P, A, AA, EV, AAA, AM, LA, DL Thrombin 1 K, V, G, P, A, AA, EV, AAA, LA, DL

[0114] Table 7

[0115] Enzyme Total_Bitter_Peptide_Count Asp-N endopeptidase 286.14 Chymotrypsin high specificity 286.14 Glutamyl endopeptidase 301.08 NTCB 304.44 Arg-C proteinase 306.84 Clostripain 306.84 CNBr 310.44 LysC 317.34 Trypsin 317.34 Chymotrypsin low specificity 321.48 Neutrophil elastase 321.48 Pepsin pH1.3 321.48 Pepsin pH>2 321.48 Proteinase K 321.48 Thermolysin 321.48

[0116] In this embodiment, while obtaining a relatively large number of target peptides by enzymatic digestion, fewer negative peptides are produced. According to the statistics of the number of negative peptides, when using protease to digest milk protein system, a large number of bitter peptides will be produced (quantity: 286.14 - 321.48), and the highest number of sweet peptides produced is only 527.58. Considering comprehensively, the current proteases are not suitable for enzymatically hydrolyzing milk to produce sweetness.

[0117] Please refer to Figure 3 , Figure 3 FIG. is a schematic structural diagram of a device for reverse screening of a dedicated protease based on a target active peptide provided by an embodiment of the present invention, including: a requirement confirmation module 301, a target protein screening module 302, a first enzymatic digestion module 303, a second enzymatic digestion module 304, and a screening module 305;

[0118] The requirement confirmation module 301 is configured to establish a target peptide database and a negative peptide database based on task requirements;

[0119] The target protein screening module 302 is configured to obtain the amino acid sequence of a protein source, screen the amino acid sequence based on the target peptide database, and obtain target protein data;

[0120] The first enzymatic digestion module 303 is configured to perform virtual enzymatic digestion on the target protein data based on the enzymatic digestion rules of a preset protease, obtain the first theoretical enzymatic digestion peptides of each protease and the corresponding target peptide information, and obtain a candidate protease set based on the target peptide segment information;

[0121] The second enzymatic digestion module 304 is configured to perform virtual enzymatic digestion on the amino acid sequence of the protein source based on the candidate protease set and the enzymatic digestion rules, obtain the second theoretical enzymatic digestion peptides of the candidate protease set, and obtain the negative peptide information of the second theoretical enzymatic digestion peptides based on the negative peptide database;

[0122] The screening module 305 is configured to screen the candidate protease set based on the negative peptide information and the target peptide information to obtain a target protease.

[0123] In this embodiment, the target protein screening module is configured to:

[0124] Obtain the amino acid sequence of a protein source based on a protein database, and screen the amino acid sequence based on the target peptide segment;

[0125] Obtain target protein data, where the target protein data includes a target protein and corresponding target peptide data; the target protein is an amino acid sequence containing the target peptide segment.

[0126] In this embodiment, the first enzymatic digestion module is configured to:

[0127] Perform virtual digestion on the target protein according to the digestion rules of each protease to obtain the first theoretical digestion peptides of all proteases;

[0128] Obtain the number of target peptides in the first theoretical digestion peptides, and screen out several candidate proteases based on a preset quantity threshold and the number of target peptides to generate a candidate protease set.

[0129] In this embodiment, the second digestion module is used for:

[0130] Perform virtual digestion on the amino acid sequence of the protein source based on the candidate proteases and their digestion rules to obtain the second theoretical digestion peptides of the candidate proteases;

[0131] Count the number and types of negative peptides in the second theoretical digestion peptides to obtain the negative peptide information.

[0132] In this embodiment, the screening module is used for:

[0133] Sort the candidate protease set in descending order based on the number of target peptides to obtain a first protease list;

[0134] Sort the candidate protease set in ascending order based on the number of negative peptides to generate a second protease list;

[0135] Integrate the first protease list and the second protease list to obtain the target protease.

[0136] In an embodiment of the present invention, a terminal device is further provided, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method for reversely screening a dedicated protease based on a target active peptide is implemented.

[0137] In an embodiment of the present invention, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method for reversely screening a dedicated protease based on a target active peptide.

[0138] Exemplarily, the computer program can be divided into one or more modules. One or more modules are stored in the memory and executed by the processor to complete the present invention. One or more modules can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device.

[0139] The terminal device may be a computing device such as a desktop computer, notebook, palm computer, and cloud server. The terminal device may include, but is not limited to, a processor, a memory, and a display. Those skilled in the art can understand that the above components are only examples of the terminal device and do not constitute a limitation on the terminal device. It may include more or fewer components than those described, or combine certain components, or have different components. For example, the terminal device may also include input / output devices, network access devices, a bus, etc.

[0140] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device and connects various parts of the entire terminal device through various interfaces and lines.

[0141] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory and by invoking the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, a text conversion function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, text message data, etc.). In addition, the memory may include high-speed random access memory and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.

[0142] Among them, when the module for reverse screening of a dedicated protease based on a target active peptide is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0143] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. In particular, it is pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for reverse screening of a dedicated protease based on a target bioactive peptide, characterized in that, Comprising: Establishing a target peptide database and a negative peptide database based on task requirements; Obtaining the amino acid sequence of a protein source, screening the amino acid sequence based on the target peptide database to obtain target protein data; Performing virtual digestion on the target protein data based on the digestion rules of a preset protease to obtain the first theoretical digestion peptides of each protease and the corresponding target peptide information, and obtaining a candidate protease set based on the target peptide information; Performing virtual digestion on the amino acid sequence of the protein source based on the candidate protease set and the digestion rules to obtain the second theoretical digestion peptides of the candidate protease set, and obtaining the negative peptide information of the second theoretical digestion peptides based on the negative peptide database; Screening the candidate protease set based on the negative peptide information and the target peptide information to obtain a target protease.

2. The method for reverse screening of a dedicated protease based on a target bioactive peptide according to claim 1, characterized in that, The target peptide database includes target peptide segments; The obtaining the amino acid sequence of the protein source, screening the amino acid sequence based on the target peptide database to obtain target protein data includes: Obtaining the amino acid sequence of the protein source based on a protein database, and screening the amino acid sequence based on the target peptide segments; Obtaining target protein data, where the target protein data includes a target protein and corresponding target peptide data; the target protein is an amino acid sequence containing the target peptide segment.

3. The method for reverse screening of a dedicated protease based on a target bioactive peptide according to claim 2, characterized in that, The target peptide information includes the number of target peptide segments; the performing virtual digestion on the target protein data based on the digestion rules of a preset protease to obtain the first theoretical digestion peptides of each protease and the corresponding target peptide segment information, and obtaining a candidate protease set based on the target peptide segment information includes: Performing virtual digestion on the target protein respectively based on the digestion rules of each protease to obtain the first theoretical digestion peptides of all proteases; Obtaining the number of target peptide segments in the first theoretical digestion peptides, and screening out a number of candidate proteases based on a preset quantity threshold and the number of target peptide segments to generate a candidate protease set.

4. The method for reverse screening of a dedicated protease based on a target active peptide according to claim 3, wherein, The negative peptide database includes negative peptide segments; the performing virtual digestion on the amino acid sequence of the protein source based on the candidate protease set and the digestion rules to obtain the second theoretical digestion peptides of the candidate protease set, and obtaining the negative peptide information of the second theoretical digestion peptides based on the negative peptide database includes: Performing virtual digestion on the amino acid sequence of the protein source based on the candidate protease and its digestion rules to obtain the second theoretical digestion peptides of the candidate protease; Counting the number and types of negative peptide segments in the second theoretical digestion peptides to obtain the negative peptide information.

5. The method for reverse screening of a dedicated protease based on a target bioactive peptide according to claim 4, wherein, The screening the candidate protease set based on the negative peptide information and the target peptide information to obtain a target protease includes: Performing a descending sort on the candidate protease set based on the number of target peptide segments to obtain a first protease list; Performing an ascending sort on the candidate protease set based on the number of negative peptide segments to generate a second protease list; Integrating the first protease list and the second protease list to obtain a target protease.

6. An apparatus for reverse screening of a dedicated protease based on a target bioactive peptide, characterized in that, Comprising: A requirement confirmation module, a target protein screening module, a first digestion module, a second digestion module, and a screening module; The requirement confirmation module is used to establish a target peptide database and a negative peptide database based on task requirements; The target protein screening module is used to obtain the amino acid sequence of the protein source, screen the amino acid sequence based on the target peptide database, and obtain target protein data; The first digestion module is used to perform virtual digestion on the target protein data based on the digestion rules of a preset protease, obtain the first theoretical digestion peptides of each protease and the corresponding target peptide information, and obtain a candidate protease set based on the target peptide information; The second digestion module is used to perform virtual digestion on the amino acid sequence of the protein source based on the candidate protease set and the digestion rules, obtain the second theoretical digestion peptides of the candidate protease set, and obtain the negative peptide information of the second theoretical digestion peptides based on the negative peptide database; The screening module is used to screen the candidate protease set based on the negative peptide information and the target peptide information to obtain a target protease.

7. The device for reverse screening dedicated protease based on target active peptide according to claim 6, characterized in that, The target protein screening module is used for: Obtaining the amino acid sequence of the protein source based on a protein database, and screening the amino acid sequence based on the target peptide segment; Obtaining target protein data, where the target protein data includes a target protein and corresponding target peptide data; the target protein is an amino acid sequence containing the target peptide segment.

8. The device for reverse screening of a dedicated protease based on a target bioactive peptide according to claim 7, characterized in that, The first digestion module is used for: Performing virtual digestion on the target protein respectively based on the digestion rules of each protease to obtain the first theoretical digestion peptides of all proteases; Obtaining the number of target peptide segments in the first theoretical digestion peptides, and screening out several candidate proteases based on a preset number threshold and the number of target peptide segments to generate a candidate protease set.

9. The device for reverse screening of a dedicated protease based on a target active peptide according to claim 8, wherein The second digestion module is used for: Performing virtual digestion on the amino acid sequence of the protein source based on the candidate protease and its digestion rules to obtain the second theoretical digestion peptides of the candidate protease; Counting the number and types of negative peptide segments in the second theoretical digestion peptides to obtain the negative peptide information.

10. The device for reverse screening of a dedicated protease based on a target active peptide according to claim 9, wherein, The screening module is used for: Sorting the candidate protease set in descending order based on the number of target peptide segments to obtain a first protease list; Sorting the candidate protease set in ascending order based on the number of negative peptide segments to generate a second protease list; Integrating the first protease list and the second protease list to obtain a target protease.

Citation Information

Cited By

  • Method and system for discovering antioxidant peptide from food and agricultural by-products and outputting enzymolysis recommendation

    CN121122400A