A foodborne multifunctional active peptide recognition method and system

By constructing functional information and dependencies at the data unit level, end-to-end prediction from protein sequence to activity recognition results is achieved, solving the problem of low efficiency in active peptide recognition in existing technologies and meeting the industrial sector's demand for efficient and rapid identification of active peptides.

CN122436005APending Publication Date: 2026-07-21SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTH CHINA AGRICULTURAL UNIVERSITY
Filing Date
2026-04-03
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing active peptide identification technologies are inefficient and cannot quickly and on a large scale discover potential functional peptides, thus failing to meet the rapid identification needs of highly efficient active peptides in industrial fields such as electroplating and corrosion prevention.

Method used

By constructing a joint representation of functional information, dependencies, and location information at the data unit level, end-to-end prediction from protein sequence to activity identification results is achieved. The training set is labeled based on the activity identification results and the model is iteratively trained to improve the efficiency of active peptide identification.

Benefits of technology

It significantly improves the efficiency of active peptide recognition, meeting the large-scale practical needs of rapid and efficient active peptide recognition in industrial fields such as electroplating and corrosion prevention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122436005A_ABST
    Figure CN122436005A_ABST
Patent Text Reader

Abstract

The application discloses a food source multifunctional active peptide recognition method and system, and is applied to the active peptide recognition technical field. The method comprises the following steps: extracting data information with protein sequences from a protein sample training set, processing the data information, and obtaining each data segment with a protein peptide segment; determining the function information of each data unit with residue information in each data segment; determining the dependency relationship of each data unit in each data segment; obtaining the active recognition result of each data segment; performing label processing on the protein sample training set to obtain a target sample data set; training a to-be-trained active peptide recognition model by using the target sample data set to obtain a target active peptide recognition model; and inputting a to-be-recognized protein sequence into the target active peptide recognition model to obtain an active peptide recognition result. The food source multifunctional active peptide recognition method and system are beneficial to improving the efficiency of active peptide recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioactive peptide recognition technology, and in particular to a method and system for recognizing food-derived multifunctional bioactive peptides. Background Technology

[0002] Bioactive peptides can be used as metal chelating agents and corrosion inhibitors, and are widely applied in industrial fields such as electroplating and coatings. Food-derived bioactive peptides are renewable, biodegradable, and have low environmental toxicity, which can reduce the negative impact of industry on the ecology and meet the development needs of clean, safe, and sustainable materials.

[0003] Current methods for identifying bioactive peptides rely on sequence feature alignment. This requires segmenting the input protein sequence data into individual peptide segments, then comparing each segment individually with peptides in a known bioactive peptide sequence library to determine activity. Clearly, this method results in low efficiency for bioactive peptide identification, hindering the rapid and large-scale discovery of potential functional peptides. Ultimately, it fails to meet the practical needs of industries such as electroplating and corrosion prevention for efficient and rapid bioactive peptide identification. Summary of the Invention

[0004] This invention provides a method and system for identifying food-derived multifunctional bioactive peptides to solve the technical problem of low efficiency in existing bioactive peptide identification technologies, thereby improving the efficiency of bioactive peptide identification.

[0005] To address the aforementioned technical problems, this invention provides a method and system for recognizing food-derived multifunctional bioactive peptides, the method comprising: Data information containing protein sequences is extracted from the protein sample training set, and the data information is processed to obtain data segments containing protein peptides. Based on a pre-set database, the functional information of each data unit with residue information in each data segment is determined respectively; Based on the collaborative relationship of each data unit, the dependency relationship of each data unit in each data segment is determined; Based on the functional information, dependencies, and location information of each data segment, the activity identification result of each data segment is obtained; Based on at least all of the activity identification results, the protein sample training set is labeled to obtain the target sample dataset; The target sample dataset is used to train the active peptide recognition model to obtain the target active peptide recognition model; The protein sequence to be identified is input into the target active peptide identification model to obtain the active peptide identification result.

[0006] Preferably, the step of determining the functional information of each data unit with residue information in each data segment based on a preset database includes: Based on the protein peptide information of each data segment, the data segment is divided to obtain each data unit with residue information; Obtain a sequence of adjacent data units of a preset length; Based on the preset database, feature analysis is performed on the adjacent data unit sequence to obtain the functional information of each data unit.

[0007] Preferably, determining the dependency relationship of each data unit in each data segment based on the collaborative relationship of each data unit includes: Based on the functional information of each data unit, the functional similarity of any two data units is analyzed to obtain the synergistic relationship of each data unit. The synergistic relationship is quantified to obtain the synergistic strength value between any two data units; Based on the strength values ​​of each synergy, the dependencies of each data unit are obtained.

[0008] Preferably, obtaining the activity identification result for each data segment based on the functional information, dependencies, and location information of each data unit includes: Based on the functional information and location information of each data unit, a feature vector of each data unit is obtained; Based on the dependencies between the data units and the feature vectors, a segment-level representation vector is obtained; Based on the segment-level representation vector, the activity identification result of each data segment is obtained through a preset classification decision function.

[0009] Preferably, the method further includes: Based on the active peptide identification results, extract the data information of each active peptide information to generate an original active peptide record set. Each piece of data in the original record set of the active peptides is encoded to generate an index structure; Based on the original record set of active peptides and the index structure, an active peptide identification result database is obtained.

[0010] Another aspect of the present invention provides a food-derived multifunctional bioactive peptide recognition system, comprising: The extraction module is used to extract data information containing protein sequences from the protein sample training set, and to process the data information to obtain data segments containing protein peptides. A functional module is used to determine the functional information of each data unit with residue information in each data segment based on a preset database. A dependency module is used to determine the dependency relationship of each data unit in each data segment based on the collaborative relationship of each data unit. The results module is used to obtain the activity identification result of each data segment based on the functional information, dependency relationship and position information of each data unit of each data segment; The sample module is used to label the protein sample training set based on at least all of the activity identification results to obtain the target sample dataset. The training module is used to train the target active peptide recognition model with the target sample dataset to obtain the target active peptide recognition model. The identification module is used to input the protein sequence to be identified into the target active peptide identification model to obtain the active peptide identification result.

[0011] Preferably, the functional module includes: A segmentation unit is used to segment the data segment based on the protein peptide information of each data segment to obtain each data unit with residue information. Adjacency unit, used to obtain a sequence of adjacent data units of a preset length; The functional information unit is used to perform feature analysis on the adjacent data unit sequence based on the preset database to obtain the functional information of each data unit.

[0012] Preferably, the dependent module includes: The interaction relationship unit is used to analyze the functional similarity of any two data units based on the functional information of each data unit, and to obtain the synergistic interaction relationship of each data unit. An intensity value unit is used to quantify the synergistic relationship and obtain the synergistic intensity value between any two data units; A dependency relationship unit is used to obtain the dependency relationship of each of the data units based on the respective synergy strength values.

[0013] Preferably, the result module includes: A feature vector unit is used to obtain a feature vector for each data unit based on the functional information and the position information of each data unit. A segment-level unit is used to obtain a segment-level representation vector based on the dependencies between the data units and the feature vectors. An activity identification unit is used to obtain the activity identification result of each data segment based on the segment-level representation vector and through a preset classification decision function.

[0014] Preferably, the system further includes: The record set unit is used to extract the data information containing the active peptide information based on the active peptide identification result, and generate the original record set of active peptides. An indexing unit is used to encode each piece of data information in the original record set of the active peptides to generate an index structure; The database unit is used to obtain an active peptide identification result database based on the original record set of active peptides and the index structure.

[0015] Compared with the prior art, the beneficial effects of the present invention are at least one of the following: Compared with existing technologies, this invention no longer relies on one-to-one sequence alignment. Instead, it achieves end-to-end prediction from protein sequence to activity identification results by constructing a joint characterization of functional information, dependencies, and location information at the data unit level, significantly improving the efficiency of active peptide identification. At the same time, by labeling the training set based on the activity identification results and iteratively training the model, the model can continuously learn and mine potential functional peptides, thereby meeting the large-scale practical needs of rapid identification of efficient active peptides in industrial fields such as electroplating and corrosion prevention. Attached Figure Description

[0016] Figure 1 This is a schematic flowchart of a method for recognizing food-derived multifunctional bioactive peptides in one embodiment of the present invention; Figure 2 This is a schematic diagram of the data processing flow of the identification tool in one embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a food-derived multifunctional bioactive peptide recognition system in one embodiment of the present invention; Figure label: The module consists of: 11. Extraction module; 12. Functional module; 13. Dependency module; 14. Result module; 15. Sample module; 16. Training module; and 17. Recognition module. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0018] In the description of this invention, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0019] In the description of this invention, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to communication within two components. The terms "vertical," "horizontal," "left," "right," "upper," "lower," and similar expressions used herein are for illustrative purposes only and do not indicate or imply that the device or component referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as limiting the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0020] In the description of this invention, it should be noted that, unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this specification is for the purpose of describing specific embodiments only and is not intended to limit the invention. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0021] Existing methods for identifying bioactive peptides rely on sequence feature alignment, which requires cutting protein sequences into peptide segments and comparing them one by one with a library of known bioactive peptides to determine their activity. This results in low identification efficiency, making it difficult to quickly and on a large scale discover potential functional peptides and failing to meet the actual needs of industries such as electroplating and corrosion prevention for rapid identification of bioactive peptides.

[0022] One embodiment of the present invention provides a method for recognizing food-derived multifunctional bioactive peptides. For details, please refer to [link to documentation]. Figure 1 , Figure 1 The diagram shown is a flowchart illustrating a method for recognizing food-derived multifunctional bioactive peptides according to one embodiment of the present invention, comprising: S1. Extract data information containing protein sequences from the protein sample training set, and process the data information to obtain data segments containing protein peptides. S2. Based on a preset database, determine the functional information of each data unit with residue information in each data segment; S3. Based on the collaborative relationship of each data unit, determine the dependency relationship of each data unit in each data segment; S4. Based on the functional information, dependencies, and location information of each data segment, obtain the activity identification result of each data segment; S5. Based on at least all activity recognition results, label the protein sample training set to obtain the target sample dataset; S6. Train the target active peptide recognition model using the target sample dataset to obtain the target active peptide recognition model. S7. Input the protein sequence to be identified into the target active peptide identification model to obtain the active peptide identification result.

[0023] First, protein sequence data is extracted from the protein sample training set and processed to obtain data segments containing protein peptides. The protein sample training set refers to a batch of text-formatted data files stored on a local disk or distributed file system, each containing multiple string records beginning with a specific identifier. A protein sequence is a linear string composed of uppercase English letters, each letter representing a predefined amino acid code; the length of the entire string ranges from tens to thousands of characters. Data information refers to the original string content read from the training set files, along with metadata fields attached to each string. Metadata fields include a unique identifier and the path to the source file. Processing involves performing string slicing operations on the read long strings according to predefined rules, based on specific positional patterns or fixed-length parameters within the string. Protein peptides are the substrings obtained after slicing, with the length of each substring constrained between a predefined minimum and maximum character count. A data segment is a key-value pair structure composed of each substring and its associated metadata information. The key is the combination of the peptide's start and end indices in the original string, and the value is the peptide string itself and its calculated character length.

[0024] First, a protein sample training set is constructed. This training set is a dataset stored in text files, typically using FASTA format (a text-based sequence storage format) as the standard data exchange format. Each FASTA file contains several records, each consisting of an identifier line and several sequence lines. The identifier line begins with a greater-than sign ">" followed by the protein's unique identifier and optional descriptive information such as the species name. The sequence lines are strings composed of consecutive single-letter codes for 20 standard amino acids, such as "MVDQEKEGGQPQEW". In practical projects, the training set can be stored in a specified directory on the local file system, or multiple file paths can be specified through a configuration file. The construction process requires collecting protein sequence files from public databases, converting them to FASTA format, removing duplicate entries and empty sequences, and finally forming a raw dataset of manageable size.

[0025] Next, protein sequence data is extracted from the training set. The core of this step is reading the file content from the disk into computer memory and parsing it into structured data objects. In existing technology, each FASTA file can be scanned line by line using file reading functions provided by programming languages ​​such as Python. When a line starting with ">" is encountered, it is parsed as the header metadata of the current protein record. The portion from the first space after the ">" is extracted as a unique identifier, and the content from the space to the end of the line is used as descriptive information. Subsequently, subsequent lines not starting with ">" are read, and the strings of these lines are concatenated sequentially until the next ">" or the end of the file is encountered. The concatenated result is the complete protein sequence string. The input is a FASTA format text file, and the output is key-value pairs corresponding to each protein record in memory, where the key is the identifier and the value is a dictionary structure containing the sequence string and descriptive information. The advantage of this approach is that it can convert unstructured text data into discrete records that the program can iterate over, providing a unified data interface for subsequent segmentation operations.

[0026] After obtaining the protein sequence string, it needs to be processed to obtain data segments containing protein peptides. In existing technologies, the most common processing methods are sliding window-based sequence cutting or rule-based enzyme digestion simulation. Taking sliding window-based cutting as an example, the protein sequence string output from the previous step is received as input, and two key parameters are set: window length and step size. The window length determines the number of characters contained in each peptide segment, for example, set to 15 characters; the step size determines the interval between the starting positions of two adjacent windows, for example, set to 3 characters. Starting from the first character of the sequence string, each time 15 consecutive characters from the current starting position are extracted as a substring, and then the step size of 3 characters is increased at the starting position. This operation is repeated until the length of the remaining string is less than the window length. Each extraction operation simultaneously records the start and end indices of the substring in the original protein sequence, with indices counting from 1. The output consists of several peptide segment records, each containing the substring itself, the start index, the end index, and the identifier of the source protein. The advantage of this sliding window method is that it can generate a large number of peptide samples with controllable length and continuous position, providing sufficient coverage for subsequent model recognition. At the same time, the preservation of index information allows each peptide to be traced back to its precise position in the source protein, which is convenient for subsequent feature extraction and result verification.

[0027] If a rule-based virtual enzyme digestion method is used, a pre-constructed enzyme digestion rule library is built. Each rule in the library corresponds to a protease or chemical reagent and includes a pattern string for the cleavage site and a cleavage direction identifier. The input is a protein sequence string and a user-specified enzyme name. The algorithm matches the protein sequence according to the pattern string corresponding to that enzyme. For example, the pattern string "K|P" indicates a cleavage between lysine K and proline P, with the vertical line representing the cleavage position. The algorithm scans every two adjacent character combinations in the sequence; if a match is found, a cleavage marker is added between them. All cleavage markers divide the protein sequence into multiple consecutive substrings, each representing a peptide. For each generated peptide, its start and end indices in the original sequence and its theoretical molecular weight are recorded. The theoretical molecular weight is calculated by summing the nominal mass values ​​of each amino acid residue in the peptide. The advantage of this virtual enzyme digestion method is that it can simulate the enzymatic digestion process in actual experiments, generating peptides that are closer to real-world applications. It also avoids the high cost and long processing time of physical enzymatic digestion experiments, making large-scale, high-throughput identification possible.

[0028] After the above processing, all peptide records are organized into a list of data segments. Each data segment typically exists in computer memory as a dictionary or a custom object, containing the following fields: peptide sequence string, source protein identifier, start index, end index, theoretical molecular weight, sequence length, etc. This list of data segments serves as the final output of this step and is directly fed into the subsequent activity identification process. The advantage of this structured representation is that all relevant attributes of each peptide are fully preserved and interconnected, eliminating the need for subsequent modules to recalculate or backtrack to the original file, significantly improving the overall efficiency of the data processing pipeline. Simultaneously, converting complete protein sequences into a large number of short peptides achieves a data granularity transformation from protein-scale to peptide-scale, providing the necessary input format for deep learning-based residue-level feature extraction and dependency modeling.

[0029] Preferably, based on a preset database, the functional information of each data unit with residue information in each data segment is determined. Based on the protein peptide information of each data segment, the data segment is segmented to obtain each data unit with residue information; adjacent data unit sequences of a preset length are obtained; feature analysis is performed on the adjacent data unit sequences based on the preset database to obtain the functional information of each data unit.

[0030] A pre-built database refers to a key-value storage structure where each key corresponds to an amino acid character code, and each value corresponds to a pre-computed feature vector for that character. A data segment is a structured record object composed of a peptide sequence string, its start and end indices, and the source protein identifier. Residue information refers to a binary tuple consisting of the position number of a single character in a data segment and the character itself. A data unit is the smallest operable object obtained by encapsulating a character and its position number. An adjacent data unit sequence is a continuous character sequence truncated after extending a fixed length to the left and right of the current data unit. Functional information is a one-dimensional numerical vector obtained after database querying and feature aggregation, used to characterize the comprehensive attributes of the data unit.

[0031] First, based on the protein peptide information of each data segment, the data segment is divided into individual data units containing residue information. In existing technologies, protein peptide information is stored in computer memory as a string, for example, a peptide string might be "MVDQEK". The essence of the segmentation operation is to convert this string into a list or array of individual characters. Specifically, the data segment object output from the previous step is received as input. This object contains a peptide sequence string field, denoted as `peptide_seq`. Built-in string traversal functions of the programming language are called; for example, a `for` loop in Python is used to read `peptide_seq` character by character, while the `enumerate` function is used to obtain the index position of each character in the string, starting from 0 or 1. For each character read, a new data unit object is created, containing two attributes: the character itself and its index value. All data unit objects are stored in a list in index order. The input is a peptide string and its length attribute; the output is a list of data units, where each element corresponds to a character and its position number. The advantage of this approach is that it breaks down a continuous string into discrete atomic objects, allowing subsequent processing to perform feature extraction and dependency modeling independently for each residue position, while preserving the spatial order of each character in the original sequence.

[0032] Next, obtain a sequence of adjacent data units of a preset length. In existing technologies, the index value recorded in the previous step is obtained. The preset length is a hyperparameter defined by the user, usually an odd number such as 5, 7, or 9, to ensure the symmetry of the neighborhoods on both sides of the central unit. In specific implementation, the data unit list output from the previous step is received as input, and the preset length parameter is read and recorded as window_size. For the data unit with index i in the list, the size of the left half window is calculated as window_size minus 1, divided by 2, and rounded down. The size of the right half window is the same. A continuous sublist is extracted from the list from index i minus the left half window to index i plus the right half window. If the index exceeds the list boundary, a padding strategy is used. Common padding characters are special symbols such as "PAD" or "X". Each element in the extracted sublist is still a character plus index structure. These characters are concatenated into a new string in their original order, called the adjacent sequence. At the same time, the relative position of the central character in the adjacent sequence is preserved, usually located at the middle index. The input consists of a list of data units, the index of each unit, and a preset window length parameter. The output is the adjacent sequence string corresponding to each data unit. The advantage of this approach is that by introducing local context information, the functional features of each data unit are no longer isolated, but rather integrated with the sequence patterns of its preceding and following neighbors. This is crucial for identifying active features that depend on local sequence patterns, while also avoiding the computational burden of processing excessively long global sequences.

[0033] Finally, feature analysis is performed on the adjacent data unit sequences based on the pre-defined database to obtain the functional information of each data unit. In existing technologies, the pre-defined database is a pre-computed lookup table, usually stored in memory as a dictionary or hash map, or in efficient key-value storage systems such as Hierarchical Data Format 5 or Lightweight Memory-Mapped Database. Each key in this database corresponds to a single-letter amino acid character such as "A", "C", "D", etc., plus a special padding character such as "PAD". The value corresponding to each key is a fixed-length floating-point vector, with common vector lengths of 128, 256, or 512. These vectors can be pre-generated in various ways, such as using one-hot encoding, word embedding models trained on large-scale corpora such as the protein bidirectional encoder representation model, or embedding vectors extracted by evolutionary scale modeling version 2. In specific implementation, each adjacent sequence string output from the previous step is received as input. Each character in the adjacent sequence string is used as a query key to search in the pre-defined database and retrieve the corresponding floating-point vector. If the character is not in the database key set, the vector corresponding to the padding character is returned. All extracted vectors are stacked into a two-dimensional matrix according to the character's order in the adjacent sequence. The shape of the matrix is ​​the window length multiplied by the vector dimension. An aggregation operation is performed on this two-dimensional matrix along its row dimension. The aggregation method can be averaging, weighted averaging, or using an attention mechanism to calculate a weighted sum, where the character at the center position is usually assigned a higher weight. The output of the aggregation operation is a one-dimensional floating-point vector, which represents the functional information of the current data unit. This operation is performed on all data units sequentially, ultimately outputting the functional information vector corresponding to each data unit. The advantages of this approach are that it transforms discrete character sequences into continuous-space numerical vector representations, allowing subsequent machine learning models to process them directly; simultaneously, by introducing pre-trained embedding vectors, it can fully utilize the statistical regularities inherent in large-scale unsupervised data, improving the richness and generalization ability of feature representation; furthermore, by pre-processing the feature extraction process as a database query operation, only table lookups are needed during actual inference without recalculating the embeddings, significantly improving processing efficiency and meeting the real-time processing requirements of large-scale data segments.

[0034] Preferably, the dependencies between data units in each data segment are determined based on the synergistic relationships between the data units. Based on the functional information of each data unit, the functional similarity between any two data units is analyzed to obtain the synergistic relationships between them; the synergistic relationships are quantified to obtain the synergistic strength values ​​between any two data units; and the dependencies between the data units are obtained based on these synergistic strength values.

[0035] A data unit is the smallest processed object consisting of a single character and its index position within a peptide string. Functional information refers to a fixed-length one-dimensional floating-point vector corresponding to each data unit, used to characterize the numerical attributes of that unit. Functional similarity refers to the numerical calculation result of the cosine distance or Euclidean distance between the functional information vectors of two data units. Synergistic relationship refers to the original association matrix obtained after comprehensively evaluating the degree of association between any two data units based on functional similarity and spatial distance. Synergistic strength value refers to a scalar value between 0 and 1 obtained after normalizing the original association values ​​in the synergistic relationship. Dependency relationship refers to the final relationship matrix obtained after thresholding and direction determination based on all synergistic strength values, used to characterize the mutual influence weights between various data units within a data segment.

[0036] First, based on the functional information of each data unit, the functional similarity and spatial distance between any two data units are analyzed to obtain the synergistic relationship between them. In existing technologies, functional information refers to a fixed-length floating-point vector corresponding to each data unit, denoted as feat_i, with a vector length typically of 128, 256, or 512. For a peptide segment of length L, there are a total of L data units. The synergistic relationship between any two data units i and j needs to be calculated, resulting in L multiplied by L combinations. Specifically, functional similarity is calculated using cosine similarity, with the formula being the dot product divided by the product of the moduli. The input consists of two functional information vectors, feat_i and feat_j, where j represents the index of another data unit different from i, and the functional information vector is feat_j. The output is a floating-point number between -1 and +1; a larger value indicates that the two vectors are closer in direction, i.e., more similar in functional features. Simultaneously, the spatial distance is calculated, and the output is a non-negative integer. Functional similarity is used as the original representation of the synergistic relationship. Perform the above calculations on all data unit pairs, outputting an L-by-L cooperative relationship matrix, where each position stores the functional similarity of that data unit pair. The advantage of this approach is that by considering functional similarity, the cooperative relationships reflect the semantic connections between data units.

[0037] Secondly, the synergy relationship is quantified to obtain the synergy strength value between any two data units. In existing technologies, the synergy relationship matrix output in the previous step contains the original functional similarity data. These data have inconsistent numerical ranges and distribution characteristics, requiring conversion to a unified scalar strength value. Specifically, the L multiplied synergy relationship matrix output in the previous step is received as input. After calculating the functional similarity for all data unit pairs, these functional similarity values ​​are globally normalized, compressing the numerical range to between 0 and 1. Normalization can be achieved using min-max normalization, where the normalized strength value equals the original functional similarity value minus the global minimum divided by the global maximum minus the global minimum. Alternatively, the Sigmoid function can be used to map the functional similarity value to the 0-1 interval. The normalized output is an L multiplied synergy strength matrix, denoted as S, where Sij represents the synergy strength value of data unit i to data unit j, ranging from 0 to 1; a larger value indicates a stronger synergy between the two. Typically, Sij and Sji are not necessarily equal. Sji represents the synergistic effect strength of data unit j on data unit i. Because the vector directionality used in functional similarity calculation may lead to asymmetry, in practice, the average of the two is often taken to ensure matrix symmetry. The advantage of this approach is that it unifies functional similarity to the same numerical scale, making the synergistic effect strength between different data unit pairs comparable. At the same time, the normalization process eliminates the dimensional differences caused by different peptide lengths and different feature distributions, making it easier to set a uniform threshold for subsequent screening.

[0038] Finally, based on the strength values ​​of each synergy, the dependencies of each data unit are obtained. In existing technologies, dependencies are sparse relational structures obtained by filtering and directionality determination based on the synergy strength matrix. Specifically, the synergy strength matrix S (L multiplied by L) output from the previous step is received as input. First, a strength threshold is set, such as 0.5 or 0.7. This threshold can be specified by the user or automatically determined by statistical methods, such as taking the average of all non-zero strength values. All elements in the S matrix are traversed, and elements with strength values ​​below the threshold are directly set to zero, while elements with strength values ​​above or equal to the threshold are retained. After threshold filtering, the S matrix becomes a sparse matrix, with only a few strongly synergistic data unit pairs retaining non-zero values. The directionality of the retained connections is further determined. Since dependencies are usually not strictly symmetric, the direction of each dependency needs to be determined. A common method is to compare Sij and Sji. If Sij is greater than Sji, data unit i is determined to depend on data unit j; otherwise, data unit j is determined to depend on data unit i. If they are equal, it is determined to be a bidirectional dependency. Alternatively, the direction can be ignored, and the connections can be directly retained as undirected connections. The final output is a dependency matrix D, with the same dimension L multiplied by L, where Dij is a scalar representing the dependency strength of data unit i on data unit j, with zero indicating no direct dependency. Further, the out-degree of each data unit can be normalized by dividing all non-zero dependency strengths corresponding to each data unit i by their sum, ensuring that the sum of the dependency strengths of each data unit on other units is 1, thus obtaining a probabilistic dependency representation. The input is the cooperative action strength matrix S and a preset threshold parameter, and the output is the dependency matrix D. The advantage of this approach is that threshold filtering removes weak associations and noisy connections, making the dependencies more sparse and interpretable, reducing the time complexity and memory consumption of subsequent calculations; simultaneously, through directionality determination and normalization, the dependencies can be directly input into graph neural networks or attention mechanisms as weight coefficients for message passing, achieving feature aggregation based on dependencies.

[0039] Preferably, the activity identification result of each data segment is obtained based on the functional information, dependencies, and location information of each data unit. Based on the functional and location information of each data unit, a feature vector of each data unit is obtained; based on the dependencies between data units and each feature vector, a segment-level representation vector is obtained; based on the segment-level representation vector, the activity identification result of each data segment is obtained through a preset classification decision function.

[0040] A data segment is a structured record object composed of a peptide sequence string, its start and end indices, and a source protein identifier. Functional information refers to a fixed-length one-dimensional floating-point vector corresponding to each data unit. Location information refers to the sequential index number of the data unit within the peptide string. Feature vectors are the final numerical representation of a single data unit obtained by concatenating the encoded representations of functional and location information. Dependencies refer to the sparse connection matrix between data units within a data segment, where each element represents the mutual influence weight between two data units. Segment-level representation vectors are fixed-length one-dimensional floating-point vectors obtained by weighted aggregation of the feature vectors of all data units within a data segment according to their dependencies. A predefined classification decision function is a predefined mathematical transformation rule, typically the forward computation formula for a linear classifier or multilayer perceptron. Activity identification results refer to a probability distribution vector output by the classification decision function and the set of activity tags obtained after threshold determination.

[0041] First, based on the functional and positional information of each data unit, a feature vector for each data unit is obtained. In existing technologies, functional information is a fixed-length floating-point vector corresponding to each data unit, denoted as feat_i, with a vector length typically of 128, 256, or 512. Positional information is the integer index of each data unit within the peptide string, denoted as pos_i, ranging from 0 to L minus 1, where L is the length of the peptide string. These two parts of information need to be merged into a unified feature vector. In specific implementation, the positional information is first encoded. Since the positional information is a scalar integer, while the functional information is a high-dimensional floating-point vector, their dimensions do not match, and they cannot be directly concatenated. A common approach is to use positional encoding techniques to convert the integer index into a positional vector with the same dimension as the functional information. A classic positional encoding method uses sine and cosine functions, with the formula: for dimension index d, the sin function is used for even dimensions, and the cosine function is used for odd dimensions, with the position value pos_i included in the parameters. Another simpler approach is to use a learnable embedding layer. An embedding matrix with a shape equal to the maximum sequence length multiplied by the vector dimension is pre-initialized, and the vector representation for each position is automatically learned during training. After obtaining the position encoding vector `pos_vec_i`, it is fused with the function information vector `feat_i`. The fusion method can be direct concatenation, in which case the feature vector dimension becomes the sum of the function information dimension and the position encoding dimension; it can also be element-wise addition, requiring both to have the same dimension; or it can be a weighted summation. The output is the final feature vector corresponding to each data unit, denoted as `x_i`, whose dimension is usually consistent with the subsequent model input requirements. The advantage of this approach is that it transforms discrete sequence position information into a continuous vector representation, enabling the model to perceive the order of data units in the sequence. Furthermore, the sine and cosine forms of the position encoding have relative positional expression capabilities, maintaining a consistent encoding pattern even as the sequence length changes.

[0042] Secondly, based on the dependencies between data units and the feature vectors, a segment-level representation vector is obtained. In existing technologies, the dependency relationship is the sparse matrix output from the previous step, denoted as D, with a dimension of L multiplied by L, where Dij represents the dependency strength of data unit i on data unit j, and is usually normalized so that the sum of Dij for all j corresponding to each i is 1. The feature vector is x_i output from the previous step, with a dimension of d. The feature vectors of all data units need to be aggregated into a fixed-length segment-level vector according to the dependencies. In specific implementation, the message passing mechanism in graph neural networks is used. Each data unit is considered as a node in the graph, and the dependency matrix D serves as the weight of the directed edges between nodes. For each node i, the feature vectors x_j pointing to its neighboring nodes j are collected and weighted according to the dependency strength Dji to obtain the aggregated neighbor information. Alternatively, a multi-layer passing mechanism can be used; for example, the first layer aggregates the neighbor information to obtain an intermediate representation, and the second layer aggregates the information of higher-order neighbors. A common and efficient approach is to use the multi-head attention mechanism in graph attention networks, where the dependency matrix can serve as a priori or bias for the attention weights. After one or more layers of message passing, each node i receives an updated node representation h_i. These node representations are then aggregated into a segment-level vector. Aggregation methods can include global average pooling (summing all h_i and dividing by L), global max pooling (taking the maximum value in each dimension), or attention pooling (learning a queryable global vector and calculating the similarity between each h_i and this vector as weights for weighted summation). The output is a fixed-length segment-level representation vector, denoted as z, with the same dimensions as h_i. The advantages of this approach are that dependency-guided information aggregation allows the segment-level representation vector to selectively focus on data units that contribute more to liveness detection while suppressing interference from noisy units; the graph message passing mechanism captures high-order interaction information between data units, rather than simply averaging or concatenating them.

[0043] Finally, based on the segment-level representation vector, the activity identification result of each data segment is obtained through a preset classification decision function. In existing technologies, the preset classification decision function is a predefined differentiable mathematical transformation module, typically composed of one or more fully connected layers. The last layer uses the Sigmoid activation function for multi-label classification or the Softmax activation function for single-label multi-class classification. Specifically, the segment-level representation vector z output from the previous step is received as input, with its dimension denoted as d_z. First, z is input to the first fully connected layer, calculated as h1 equal to z multiplied by W1 plus b1, where W1 is a weight matrix of shape d_z multiplied by h1_dim, b1 is the bias vector, and h1_dim is the hidden layer dimension. Then, a non-linear activation function such as ReLU is applied, with the formula being to set all less than zero elements in h1 to zero. Multiple fully connected layers can be stacked to increase model capacity. The output dimension of the last fully connected layer is equal to the total number of active categories, denoted as C. For multi-label classification tasks, a Sigmoid activation function is applied after the last layer. The formula is: output value equals 1 divided by 1 plus e raised to the power of the negative input value. This compresses each output value to between 0 and 1, representing the probability that the data segment belongs to each active category. A decision threshold, typically 0.5, is set; categories with probabilities greater than the threshold are labeled positive, and the rest are labeled negative. The output is a probability distribution vector of length C and a binary set of active labels. For single-label multi-class classification tasks, a Softmax activation function is used. The formula is: probability for each category equals e raised to the power of the input value for that category divided by the sum of the powers of e for all categories. The output is a distribution vector with a probability sum of 1. The category with the highest probability is taken as the recognition result. The advantages of this approach are that the classification decision function is end-to-end differentiable, enabling joint training with the preceding feature extraction module and automatic optimization of all parameters through the backpropagation algorithm; the sigmoid thresholding method in multi-label classification naturally supports a data segment having multiple activity categories simultaneously, which aligns with the multifunctionality of peptides in practical applications; and the stacked structure of fully connected layers endows the model with the ability to non-linearly map from segment-level representations to activity labels, enabling it to fit complex data distributions.

[0044] Preferably, the protein sample training set is labeled based on at least all activity recognition results to obtain the target sample dataset.

[0045] The activity identification result refers to a probability distribution vector output by the classification decision function and a set of binary activity labels obtained after threshold determination. The protein sample training set refers to a batch of text-formatted data files stored on local disk or in distributed files. Each file contains multiple string records starting with a specific identifier and metadata fields attached to each string. Labeling processing refers to the process of writing or associating the set of binary activity labels from the activity identification result corresponding to each data segment as a new label field to the original record corresponding to that data segment in the training set. The target sample dataset refers to the new data set obtained after labeling processing. Each record in this set contains both the original protein sequence or peptide string and activity label annotation information predicted by the model.

[0046] The activity identification result refers to a probability distribution vector output by the classification decision function and a set of binary activity labels obtained after threshold determination. In actual engineering implementation, each data segment generates a floating-point array of length C as the probability distribution after forward computation by the model, for example, [0.92, 0.13, 0.78]. Each element in this array is compared with a preset threshold, such as 0.5; elements greater than or equal to the threshold are set to 1, otherwise to 0, resulting in a binary label array, for example, [1, 0, 1]. The protein sample training set refers to a batch of text-formatted data files stored on a local disk or in a distributed file system. Each file contains multiple string records starting with a specific identifier and metadata fields attached to each string. In existing technologies, these files are typically stored in formats such as CSV (Comma-Separated Values), with each record containing a unique identifier field, a peptide sequence string field, and possibly a raw label field. Labeling refers to the process of writing or associating the set of binary activity labels from the activity identification results corresponding to each data segment as a new label field to the original record corresponding to that data segment in the training set. Specifically, a mapping table is maintained first, where the key is the unique identifier of the data segment (e.g., a combination of the source protein identifier, start index, and end index), and the value is the array of binary activity labels corresponding to that data segment. Then, each record in the original protein sample training set is traversed, extracting the unique identifier from the record and searching for the corresponding activity label array in the mapping table. If the search is successful, the label array is appended as a new field to the original record. The updated record is then written to a new file or database table. The input is the original protein sample training set file and the activity identification result mapping table; the output is an expanded dataset with the added predicted label field for each original record. The target sample dataset refers to the new dataset obtained after labeling, where each record contains both the original protein sequence or peptide string and the activity label annotation information predicted by the model. In existing technologies, this dataset is usually stored in the same format as the original training set, but with the added label field, and can be directly used for the next round of model iteration training. The advantage of this approach is that by feeding the model's predictions back to the training set as pseudo-labels, a semi-supervised learning or self-training mechanism is achieved, enabling the model to continuously optimize itself using unlabeled data. This is especially suitable for scenarios where labeled samples are scarce in the identification of active peptides. At the same time, the labeling process retains all fields of the original data, ensuring data traceability and reproducibility, and providing a unified data format for subsequent model evaluation and iteration.

[0047] Preferably, the target sample dataset is used to train the active peptide recognition model to obtain the target active peptide recognition model.

[0048] The target sample dataset refers to the new data set obtained after labeling. Each record in this set contains both the original peptide string and a binary activity label array predicted by the model. The active peptide recognition model to be trained refers to a deep learning network structure whose parameters have not yet been optimized. This network contains an input layer, multiple hidden layers, and an output layer, connected by learnable weight matrices and bias vectors. Training refers to the process of feeding the peptide strings from the target sample dataset as input features into the model to be trained, using the corresponding binary activity label array as supervision signals, calculating the prediction error through forward propagation, and iteratively updating the model parameters through backpropagation. The target active peptide recognition model is the final model instance after training, where its parameters have converged and the prediction error is below a preset threshold. This instance can be saved to a disk file for subsequent inference tasks.

[0049] The target sample dataset refers to the new data set obtained after labeling. Each record in this set contains both the original peptide string and a binary activity label array predicted by the model. In existing technologies, this dataset is typically divided into three subsets: training set, validation set, and test set, with common ratios of 8:1:1 or 7:2:1. Before being input into the model, the peptide string of each record needs to be converted into a numerical tensor representation. The conversion method is consistent with the aforementioned feature extraction steps, i.e., first segmenting it into data units, then obtaining the functional information vector of each data unit through a pre-defined database query, ultimately forming a two-dimensional tensor with the shape of sequence length multiplied by the feature dimension. The binary activity label array corresponding to each record is a one-dimensional integer tensor of length C, where C is the total number of activity categories. Each element in the array takes a value of 0 or 1, where 1 indicates that the peptide has a corresponding activity category.

[0050] The active peptide recognition model to be trained refers to a deep learning network structure whose parameters have not yet been optimized. This network includes an input layer, multiple hidden layers, and an output layer, with each layer connected by learnable weight matrices and bias vectors. In existing technologies, the specific architecture of this model is consistent with the modules in the aforementioned active peptide recognition process: First, it includes a feature fusion layer, used to concatenate or add the functional information and positional encoding of each data unit to obtain a feature vector; second, it includes a dependency modeling layer, typically employing a graph neural network or multi-head attention mechanism, used for message passing and feature aggregation based on the dependencies between data units, outputting a segment-level representation vector; finally, it includes a classification decision layer, composed of one or more fully connected layers, with an output dimension of C, and the last layer uses the sigmoid activation function. All learnable parameters of the model to be trained include the positional encoding embedding matrix, the weight matrix in the graph neural network, the weight matrix of the fully connected layers, and the bias vector, etc., and these parameters need to be initialized before training begins.

[0051] Training refers to the process of feeding peptide strings from the target sample dataset as input features into the model to be trained, using the corresponding binarized activity label array as supervision signals, calculating the prediction error through forward propagation, and then iteratively updating the model parameters through backpropagation. The active peptide recognition model in this embodiment of the invention is a common deep learning model.

[0052] In practice, a batch of samples is first read from the target dataset. The batch size is typically set to 32, 64, or 128, depending on the available GPU memory. The peptide strings of this batch of samples are converted into an input tensor with a shape equal to the batch size multiplied by the sequence length multiplied by the feature dimension. This tensor is then fed into the model to be trained for forward propagation. The model outputs a probability prediction tensor with a shape equal to the batch size multiplied by C, where each element is between 0 and 1. The loss between the predicted output and the true label is then calculated. Since this is a multi-label classification task and labels may have high overlap and sample imbalance, existing techniques often employ binary cross-entropy loss, etc. The formula for binary cross-entropy loss is to calculate the logarithm of the true label multiplied by the predicted probability for each class, add 1 minus the logarithm of the true label multiplied by 1 minus the predicted probability, then take the negative sign and sum. Commonly used optimizers include stochastic gradient descent. The process of forward propagation, loss calculation, backpropagation, and parameter update is repeated. Each batch processed is called an iteration, and each complete traversal of the entire training set is called a round. Training typically involves multiple rounds, such as 50 rounds, 100 rounds, or until the loss on the validation set no longer decreases.

[0053] During training, the model's performance needs to be evaluated periodically on the validation set. The validation set does not participate in parameter updates; it is only used to monitor for overfitting and select the optimal model checkpoint. Evaluation metrics include precision, recall, F1 score, and the area under the precision-recall curve. For multi-label classification tasks, the metrics for each category are typically calculated and then averaged macro- or micro-averaged. Training can be terminated early to avoid overfitting when performance on the validation set no longer improves over several consecutive epochs. After training, the optimal model parameters from the validation set are saved; this saved model instance is the target active peptide recognition model.

[0054] Through end-to-end supervised learning, the model can automatically learn the complex mapping relationship from peptide sequences to active tags from data, without the need for manual feature rule design. Training using mini-batch stochastic gradient descent enables the model to handle ultra-large datasets exceeding the memory capacity of a single machine, while the randomness of gradients between batches helps the model escape local optima. Validation set monitoring and early stopping mechanisms effectively prevent overfitting and ensure the model's generalization ability. The trained target model can be saved in a standard model file format, facilitating subsequent model deployment, inference, and transfer learning, thus decoupling the training and inference phases.

[0055] Finally, the protein sequence to be identified is input into the target bioactive peptide identification model to obtain the bioactive peptide identification results. Based on the bioactive peptide identification results, data information containing bioactive peptide information is extracted to generate a raw bioactive peptide record set; each data information in the raw bioactive peptide record set is encoded to generate an index structure; based on the raw bioactive peptide record set and the index structure, a bioactive peptide identification result database is obtained.

[0056] The protein sequence to be identified refers to a linear string consisting of uppercase English letters, which has not undergone any segmentation or feature extraction processing and serves as the raw input data for the model inference stage. The target active peptide identification model refers to a deep learning network instance whose parameters have converged after training. This instance is loaded into computer memory and can receive input tensors and output prediction tensors. The active peptide identification result refers to a probability distribution vector output by the model and a set of binary active tags obtained after threshold determination. The active peptide raw record set refers to a list of records obtained by assembling the unique identifier of each data segment, the peptide string, the start and end indexes, the source protein identifier, and the corresponding active tag set into a structured record. The index structure refers to a lookup table that organizes the mapping relationship between the encoded processing values ​​and the physical addresses or offsets of the records in the storage file. The active peptide identification result database refers to a queryable data set formed by persistently storing the active peptide raw record set and the index structure on a disk file or in a database management system.

[0057] First, the protein sequence to be identified is input into the target active peptide identification model to obtain the active peptide identification results. The protein sequence to be identified is a linear string composed of uppercase English letters, which has not undergone any segmentation or feature extraction processing. In existing technologies, the data processing pipeline in the model inference stage is consistent with that in the training stage. Specifically, the protein sequence to be identified is first segmented into multiple peptide segments using the aforementioned sliding window or virtual enzyme digestion method, and each peptide segment is converted into a data segment structure. Then, for each data segment, data unit segmentation, adjacent sequence extraction, database query to obtain functional information, calculation of dependencies, and fusion of positional information are performed sequentially to obtain a segment-level representation vector. Finally, a probability distribution vector and a binary set of activity labels are output through a classification decision function. These operations reuse all the processing modules defined in the training stage. The input is a protein string, and the output is a list of activity identification results corresponding to all peptide segments derived from that protein.

[0058] Secondly, based on the bioactive peptide identification results, data information containing bioactive peptide information is extracted to generate a raw bioactive peptide record set. In existing technologies, each element in the bioactive peptide identification result list corresponds to the identification result of a peptide segment. For each identification result, the peptide sequence string, source protein identifier, start index, end index, theoretical molecular weight, and sequence length stored in its data segment are extracted. Then, the binarized bioactive tag array of the peptide segment is converted into a comma-separated string format. These fields are assembled into a record. All records are stored in a list, which is the raw bioactive peptide record set.

[0059] Next, each piece of data in the original bioactive peptide record set is encoded to generate an index structure. In existing technologies, an index is needed to quickly retrieve all bioactive peptides from a specific peptide segment or source protein. Specifically, each record in the original bioactive peptide record set is traversed, and its unique identifier is extracted. This identifier can be a concatenated string of the source protein identifier plus a start index and an end index, such as "P00738_112_126". A hash function, such as Python's built-in hash function, is applied to this identifier string to generate a fixed-length integer hash value. Then, two mapping relationships are established: one is a mapping from the hash value to the record's index position in the original record set, and the other is a mapping from the source protein identifier to a list of hash values ​​for all records derived from that protein. These two mapping relationships together constitute the index structure. In existing technologies, the index structure can be stored in a memory dictionary or serialized to a disk file.

[0060] Finally, based on the original record set and index structure of active peptides, a database of active peptide identification results is obtained. In existing technologies, the database is formed by persistently storing the original record set and index structure of active peptides. A simple implementation is to save the original record set as a CSV file and the index structure as a separate JSON (JavaScript Object Notation) file, placing the two files in the same directory to form a lightweight database.

[0061] Step S7 persistently stores the model inference results, avoiding the need to rerun the model for each query and significantly reducing computational costs. By establishing an index structure, the time complexity of linear lookups is reduced, enabling real-time retrieval of large-scale bioactive peptide data. The standardized database format facilitates data exchange and integration with other bioinformatics tools or industrial application systems. The design separating the original record set from the index structure allows for data updates by simply appending new records and incrementally updating the index, without rebuilding the entire database, thus exhibiting excellent scalability. This database supports various combined queries, such as filtering by activity category and by peptide length range, providing a data foundation for downstream applications such as enzymatic hydrolysis process optimization and functional food formulation development.

[0062] It should be noted that the above embodiments provide a complete method for identifying active peptides, from data preprocessing, feature extraction, dependency modeling to model training and inference. Its core lies in achieving end-to-end prediction from protein sequence to activity identification results through the joint representation of functional information, synergistic relationships, and location information at the data unit level. However, in practical applications, the specific implementation of this method can be further enhanced with informatics tools to improve feature expression and model interpretability. Therefore, the following embodiments, based on the method framework constructed in Embodiment 1, specifically provide an informatics tool implementation scheme called Foodpep-Hunter (Food-derived Peptide Hunter). This embodiment, based on the preferred steps of the above embodiments, introduces specific technical means such as semantic representation of the pre-trained model, terminal bias mechanism, three-way heterogeneous feature fusion architecture, and Focal-Dice composite loss function, while providing an engineered implementation for protease virtual screening and peptide activity backtracking.

[0063] Another embodiment of the present invention provides an informatics tool called Foodpep-Hunter (Food-derived PeptideHunter), which exhibits a high degree of intelligence and systematicity in data processing and model architecture. This tool first supports batch import of protein sequences from mainstream databases such as general protein resource libraries and protein databases, and automatically parses the FASTA format sequence header information using a set of regular expression rules, extracting and standardizing protein names, ensuring the consistency and structure of the input data, and laying a solid foundation for subsequent high-throughput analysis.

[0064] In the core feature engineering and modeling stages, Foodpep-Hunter abandons traditional one-hot encoding or simple average pooling methods, instead employing an ESM-2 (Evolutionary Scale Modeling 2) pre-trained large language model to extract high-dimensional semantic features of peptides. Its innovation lies in introducing a terminal bias mechanism, which, while preserving global semantic information, assigns greater attention to N-terminal and C-terminal residues through a learnable weight function, thus significantly enhancing the model's ability to represent functionally critical regions. Based on this, the system constructs a three-path heterogeneous feature deep coupling algorithm to process multi-dimensional information in parallel: the first path models the complex interactions between residues and functional labels through a sequence contribution map attention network; the second path utilizes a Transformer decoder architecture to target and scan features with learnable query vectors, generating activity-related residue contribution scores; and the third path captures local physicochemical motifs through a convolutional neural network module. The three feature paths are finally fused and then used for multi-label classification via a grouped linear layer.

[0065] To address the prevalent issues of highly overlapping labels and uneven sample distribution in real-world data, the model employs a Focal-Dice Loss composite loss function for optimization. This strategy effectively improves the model's ability to identify difficult-to-classify samples and enhances the modeling accuracy of label intersection regions. The entire prediction process not only outputs the multi-label probability distribution of each peptide corresponding to 11 biological activities, but also uses an attention reduction mechanism to backtrack and visualize the contribution of each amino acid residue to a specific function, achieving a leap from black-box prediction to interpretable decision-making.

[0066] Furthermore, Foodpep-Hunter's effectiveness is also reflected in its closed-loop intelligent recommendation system. Through large-scale virtual screening of a built-in library of 40 protease / chemical reagent cleavage rules, the model can quantitatively assess the efficiency of each enzyme in releasing the target active peptide, while simultaneously calculating the risk of generating negative peptides such as bitter peptides and sensitizing peptides. Based on this, the system can intelligently rank all candidate enzymes by combining these two indicators, prioritizing the optimal solution with high activity and low negative yield, thus transforming complex bioinformatics analysis into clear and actionable engineering guidance.

[0067] This invention provides a data processing flow for the Foodpep-Hunter identification tool. For details, please refer to [link / reference needed]. Figure 2 , Figure 2 The diagram illustrates the data processing flow of the identification tool in one embodiment of the present invention, including: First, in the protein sequence input stage, receiving FASTA format sequences and automatically parsing the header metadata to perform sequence naming standardization. Second, in the enzyme digestion method selection stage, based on user-specified activity type and peptide length constraints, dynamically calling the built-in 40-protease / chemical reagent cleavage rule library to achieve targeted virtual hydrolysis. Third, generating structured annotation information for each peptide, including its start and end coordinates in the source protein, enzyme digestion source, and theoretical molecular weight; Fourth, the core innovation layer, feature extraction, abandoning traditional mean pooling, using the 33rd layer hidden state of evolutionary scale modeling as the basic semantic representation, and introducing a learnable terminal bias function, with the N-terminal and C-terminal weights adaptively decreasing with sequence length to enhance the representation strength of key functional residues. The fifth stage is activity prediction, based on a three-way heterogeneous feature fusion architecture. A sequence contribution-aware map attention network models the interaction contributions between residues and tags, a Transformer decoder performs target-oriented attention scanning, and a convolutional neural network module captures local physicochemical motifs. The three outputs are concatenated and input into a linear classifier, employing a Focal-Dice composite loss function to address tag overlap and sample imbalance issues. The sixth stage outputs the multi-tag probability distribution of 11 biological activities and generates a residue-level contribution heatmap for each activity through an attention reduction mechanism, i.e., peptide activity output. Finally, in the seventh stage, the prediction results are back-mapped to entries in the original protein database, constructing a four-dimensional association record table of active peptides, source proteins, precise sites, and recommended enzymes. This ultimately supports downstream applications.

[0068] In practical data processing, Foodpep-Hunter demonstrated high-throughput recognition capabilities. The model performed full-library virtual enzyme digestion and activity prediction on 67,929 protein sequences from the Chinese pond turtle, identifying 548,859 peptides with single activity tags, validating its efficiency in processing large-scale sequence data and its high sensitivity to target activity categories.

[0069] In multifunctional peptide recognition, the model achieved high-confidence recognition by setting a prediction probability threshold. The four candidate peptides output by the model met the high-probability criteria in multi-label prediction, further validating the model's reliability in multi-task learning and label association modeling. Furthermore, the model provides actionable computational guidance for optimizing enzymatic digestion processes. Through virtual activity identification of 40 proteases, the model predicted proteinase K as the optimal enzyme for releasing the target active peptide. This prediction, based on the activity probability distribution and abundance statistics of the virtual digestion products of each enzyme, enables the rational design of enzymatic digestion strategies.

[0070] Another embodiment of the present invention provides a food-derived multifunctional bioactive peptide recognition system. For details, please refer to [link to relevant documentation]. Figure 3 , Figure 3 The diagram shown is a structural schematic of a food-derived multifunctional bioactive peptide recognition system according to one embodiment of the present invention, comprising: Extraction module 11 is used to extract data information containing protein sequences from the protein sample training set and process the data information to obtain data segments containing protein peptides. Functional module 12 is used to determine the functional information of each data unit with residue information in each data segment based on a preset database. Dependency module 13 is used to determine the dependency relationship of each data unit in each data segment based on the collaborative relationship of each data unit. Result module 14 is used to obtain the activity identification result of each data segment based on the functional information, dependency relationship and position information of each data unit of each data segment; Sample module 15 is used to label the protein sample training set based on at least all activity recognition results to obtain the target sample dataset; Training module 16 is used to train the target active peptide recognition model with the target sample dataset to obtain the target active peptide recognition model; The identification module 17 is used to input the protein sequence to be identified into the target active peptide identification model to obtain the active peptide identification result.

[0071] Preferably, functional module 12 includes: The segmentation unit is used to segment the data segment based on the protein peptide information of each data segment, so as to obtain each data unit with residue information. Adjacency unit, used to obtain a sequence of adjacent data units of a preset length; The functional information unit is used to perform feature analysis on adjacent data unit sequences based on a preset database to obtain the functional information of each data unit.

[0072] Preferably, dependent module 13 includes: The interaction relationship unit is used to analyze the functional similarity of any two data units based on the functional information of each data unit, and to obtain the synergistic interaction relationship of each data unit. The intensity value unit is used to quantify the synergistic relationship and obtain the synergistic strength value between any two data units; Dependency unit, used to obtain the dependency relationship of each data unit based on the strength value of each synergy.

[0073] Preferably, the result module 14 includes: Feature vector unit, used to obtain the feature vector of each data unit based on the functional and positional information of each data unit; Segment-level units are used to obtain segment-level representation vectors based on the dependencies between data units and each feature vector; The activity identification unit is used to obtain the activity identification result of each data segment based on the segment-level representation vector and through a preset classification decision function.

[0074] Preferably, the system also includes: The record set unit is used to extract data information containing active peptide information based on the active peptide identification results, and generate the original record set of active peptides. The index unit is used to encode each piece of data in the original record set of active peptides and generate an index structure. The database unit is used to obtain a database of active peptide identification results based on the original record set and index structure of active peptides.

[0075] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0076] Accordingly, embodiments of the present invention provide a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device containing the computer-readable storage medium to perform steps in the food-derived multifunctional bioactive peptide recognition method of the above embodiments, for example... Figure 1 Steps S1 to S7 as described above.

[0077] This invention first extracts protein sequence data from a protein sample training set and processes it to obtain data segments of various protein peptides. Then, based on a preset database, it determines the functional information of each residue data unit in each data segment, thereby mapping the original sequence into structured features. On this basis, it determines the dependency relationship of each data unit in each data segment based on the synergistic relationship of each data unit. Then, it combines the functional information, dependency relationship, and position information of each data unit in each data segment to generate an activity identification result. Then, based on at least all activity identification results, it labels the protein sample training set to obtain a target sample dataset. The target active peptide identification model is trained using this dataset. Finally, the active peptide identification result is obtained by inputting the protein sequence to be identified into the model. On the one hand, this invention introduces joint modeling of functional information, synergistic relationships, and location information at the data unit level, enabling the model to complete knowledge accumulation during the offline training phase and output results with only a single forward computation during the online recognition phase. This fundamentally eliminates the reliance on real-time comparison with known active peptide libraries, significantly improving recognition efficiency. On the other hand, by labeling and iteratively training the training set based on the activity recognition results, the model is equipped with the ability to autonomously mine potential functional peptides from a large-scale protein sample training set, thereby meeting the practical application needs of high-throughput and rapid-response active peptide recognition in industrial fields such as electroplating and corrosion prevention.

[0078] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A method for identifying food-derived multifunctional bioactive peptides, characterized in that, include: Data information containing protein sequences is extracted from the protein sample training set, and the data information is processed to obtain data segments containing protein peptides. Based on a pre-set database, the functional information of each data unit with residue information in each data segment is determined respectively; Based on the collaborative relationship of each data unit, the dependency relationship of each data unit in each data segment is determined; Based on the functional information, dependencies, and location information of each data segment, the activity identification result of each data segment is obtained; Based on at least all of the activity identification results, the protein sample training set is labeled to obtain the target sample dataset; The target sample dataset is used to train the active peptide recognition model to obtain the target active peptide recognition model; The protein sequence to be identified is input into the target active peptide identification model to obtain the active peptide identification result.

2. The method for recognizing food-derived multifunctional bioactive peptides as described in claim 1, characterized in that, The method, based on a preset database, determines the functional information of each data unit with residue information in each data segment, including: Based on the protein peptide information of each data segment, the data segment is divided to obtain each data unit with residue information; Obtain a sequence of adjacent data units of a preset length; Based on the preset database, feature analysis is performed on the adjacent data unit sequence to obtain the functional information of each data unit.

3. The method for recognizing food-derived multifunctional bioactive peptides as described in claim 1, characterized in that, The determination of the dependency relationship of each data unit in each data segment based on the collaborative relationship of each data unit includes: Based on the functional information of each data unit, the functional similarity of any two data units is analyzed to obtain the synergistic relationship of each data unit. The synergistic relationship is quantified to obtain the synergistic strength value between any two data units; Based on the strength values ​​of each synergy, the dependencies of each data unit are obtained.

4. The method for identifying food-derived multifunctional bioactive peptides as described in claim 1, characterized in that, The process of obtaining the activity identification result for each data segment based on the functional information, dependencies, and location information of each data unit includes: Based on the functional information and location information of each data unit, a feature vector of each data unit is obtained; Based on the dependencies between the data units and the feature vectors, a segment-level representation vector is obtained; Based on the segment-level representation vector, the activity identification result of each data segment is obtained through a preset classification decision function.

5. The method for recognizing food-derived multifunctional bioactive peptides as described in claim 1, characterized in that, The method further includes: Based on the active peptide identification results, extract the data information of each active peptide information to generate an original active peptide record set. Each piece of data in the original record set of the active peptides is encoded to generate an index structure; Based on the original record set of active peptides and the index structure, an active peptide identification result database is obtained.

6. A food-derived multifunctional bioactive peptide recognition system, characterized in that, include: The extraction module is used to extract data information containing protein sequences from the protein sample training set, and to process the data information to obtain data segments containing protein peptides. A functional module is used to determine the functional information of each data unit with residue information in each data segment based on a preset database. A dependency module is used to determine the dependency relationship of each data unit in each data segment based on the collaborative relationship of each data unit. The results module is used to obtain the activity identification result of each data segment based on the functional information, dependency relationship and position information of each data unit of each data segment; The sample module is used to label the protein sample training set based on at least all of the activity identification results to obtain the target sample dataset. The training module is used to train the target active peptide recognition model with the target sample dataset to obtain the target active peptide recognition model. The identification module is used to input the protein sequence to be identified into the target active peptide identification model to obtain the active peptide identification result.

7. The food-derived multifunctional bioactive peptide recognition system as described in claim 6, characterized in that, The functional modules include: A segmentation unit is used to segment the data segment based on the protein peptide information of each data segment to obtain each data unit with residue information. Adjacency unit, used to obtain a sequence of adjacent data units of a preset length; The functional information unit is used to perform feature analysis on the adjacent data unit sequence based on the preset database to obtain the functional information of each data unit.

8. The food-derived multifunctional bioactive peptide recognition system as described in claim 6, characterized in that, The dependent modules include: The interaction relationship unit is used to analyze the functional similarity of any two data units based on the functional information of each data unit, and to obtain the synergistic interaction relationship of each data unit. An intensity value unit is used to quantify the synergistic relationship and obtain the synergistic intensity value between any two data units; A dependency relationship unit is used to obtain the dependency relationship of each of the data units based on the respective synergy strength values.

9. The food-derived multifunctional bioactive peptide recognition system as described in claim 6, characterized in that, The result module includes: A feature vector unit is used to obtain a feature vector for each data unit based on the functional information and the position information of each data unit. A segment-level unit is used to obtain a segment-level representation vector based on the dependencies between the data units and the feature vectors. An activity identification unit is used to obtain the activity identification result of each data segment based on the segment-level representation vector and through a preset classification decision function.

10. The food-derived multifunctional bioactive peptide recognition system as described in claim 6, characterized in that, The system also includes: The record set unit is used to extract the data information containing the active peptide information based on the active peptide identification result, and generate the original record set of active peptides. An indexing unit is used to encode each piece of data information in the original record set of the active peptides to generate an index structure; The database unit is used to obtain an active peptide identification result database based on the original record set of active peptides and the index structure.