Oligopeptide toxicity prediction method based on ensemble learning

By establishing SpepToxPred short peptide toxicity prediction model, using multi-source data integration and integrated learning prediction system, the problem of time-consuming, labor-intensive and mutated traditional toxicity evaluation methods is solved, and high-precision and rapid toxicity prediction is achieved, reducing development costs and improving R&D efficiency.

CN119943157AInactive Publication Date: 2025-05-06HUNAN NORMAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510156953.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional short peptide toxicity assessment method is time-consuming and labor-intensive, the sample processing is complex, and the experimental variation is large, resulting in increased development costs and complexity, making it difficult to meet the rapid and accurate toxicity prediction needs of functional short peptides in the food and medicine fields.

Method used

A SpepToxPred short peptide toxicity prediction model was established, and high-precision toxicity prediction was achieved through multi-source data integration, sequence standardization processing, feature engineering system and integrated learning prediction system.

Benefits of technology

It improves the accuracy and efficiency of toxicity prediction of short peptides, reduces development costs, significantly improves the research and development efficiency of functional short peptides, and verifies the reliability and practical value of the model in the prediction of non-toxic short peptides.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention relates to an integrated learning-based oligopeptide toxicity prediction method, and belongs to the fields of bioinformatics and food safety. The TasToxPred prediction model based on multi-algorithm integration is provided for solving the technical problems that a traditional short peptide toxicity assessment method is long in consumed time and high in cost and an existing calculation method is insufficient in prediction precision of short peptides (less than or equal to 25 amino acids). According to the model, 20 kinds of sequence coding descriptors are integrated to construct a multi-dimensional feature representation system, and a dynamic weight configuration mechanism of 9 kinds of machine learning algorithms is adopted to carry out integrated prediction, so that accurate evaluation on the toxicity of the oligopeptide is realized. By applying the method, the Matthews correlation coefficient (MCC) of 0.7019 and the accuracy rate of 0.8445 are realized on an independent test set, 73 oligopeptides which are predicted to be non-toxic are verified through a CCK-8 cytotoxicity experiment and hemolytic activity determination, and the result shows that the cell survival rate is kept above 90% and the hemolysis rate is lower than 1.5% under the concentration of 100 mu M, so that the reliability of model prediction is fully verified, and the method can be applied to the field of clinical application. The method can be widely applied to safety evaluation, development and screening of functional oligopeptides.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and in particular to a short peptide toxicity prediction method based on ensemble learning. Background Art

[0002] As an important class of bioactive molecules, short peptides have broad application prospects in the fields of food and medicine. Especially in the food industry, functional short peptides such as taste peptides and antioxidant peptides (sequence length is usually ≤25 amino acids) have attracted much attention due to their unique biological functions. However, some short peptides may have potential toxicity risks, which not only affects their biosafety, but also limits their development process in practical applications. Traditional toxicity assessment methods mainly rely on biological verification methods such as cytotoxicity experiments and hemolytic activity assays. These methods are not only time-consuming and labor-intensive, but also have limitations such as complex sample processing and large experimental variation, which significantly increase development costs and complexity.

[0003] In recent years, computer-aided prediction methods have made significant progress in the field of peptide toxicity assessment, and multiple prediction models have been developed. With the increasing application of functional short peptides in food and medicine, the demand for accurate prediction of the toxicity of short peptides (sequence length ≤ 25 amino acids) is also growing. The development of a high-precision toxicity prediction model specifically for short peptides has important practical significance for promoting the research and development and application of functional short peptides. Especially in the development process of functional short peptides such as taste peptides, rapid and accurate toxicity prediction can significantly improve research and development efficiency and reduce development costs. Summary of the invention

[0004] The present invention establishes a systematic short peptide toxicity prediction model SpepToxPred, which specifically includes: Multi-source data integration system: Integrate professional databases, prediction model data sets and literature mining data to build a comprehensive training data set; Sequence standardization: Establish sequence screening and toxicity annotation mechanisms to ensure data quality and balance; Feature engineering system: construct a multidimensional feature analysis system including amino acid composition, physicochemical properties and sequence coding; Ensemble learning prediction system: adopts multi-classifier weighted integration strategy to achieve high-precision toxicity prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below.

[0006] Figure 1A flow chart was constructed for the dataset of the present invention, showing the process of acquiring and integrating short peptide sequence data from multiple data sources, including a unified integration strategy for professional databases, prediction model datasets, and literature mining data; Figure 2 This is a data set distribution feature analysis diagram of the present invention, showing the sequence length distribution and amino acid composition characteristics of toxic and non-toxic peptides, wherein the complete data set contains 6861 toxic peptides and 9183 non-toxic peptides, and the filtered data set (≤25 amino acids) contains 2821 toxic and non-toxic peptides of matching lengths; Figure 3 This is an amino acid frequency distribution analysis diagram of the present invention, showing the changing trend of the occurrence frequency of 20 amino acids in toxic peptides of different lengths (4-50 amino acids), and the threshold position of 25 amino acids is marked by a dotted line; Figure 4 The development and optimization framework of the TasToxPred model of the present invention is shown, which shows the integration strategy of the feature engineering system (20 sequence-encoded descriptors) and the machine learning algorithm (9 basic classifiers), as well as the performance comparison results with 17 existing toxicity prediction tools; Figure 5 This is a biosafety validation diagram of the present invention, showing the cytotoxicity (four cell lines: BEAS-2B, HEK293T, HPNE and HUVEC) and hemolytic activity evaluation results of 73 peptides at a concentration of 100 μM. DETAILED DESCRIPTION

[0007] In order to allow those skilled in the art to understand the present invention more clearly and intuitively, the present invention will be further described below in conjunction with the accompanying drawings.

[0008] like Figure 1 As shown, the data sources of the present invention include: Prediction model dataset: ToxGIN database; ToxinPred 3.0 database; ToxTeller database; ToxIBTL database.

[0009] Professional database: Conoserver database; DRAMP 3.0 database; CAMPR3 database.

[0010] like Figure 2 As shown in Figure 2, the data preprocessing process includes: Sequence Normalization: Sequence screening: Eliminate sequences containing non-standard amino acid residues; The sequence length was limited to no more than 25 amino acids.

[0011] Data balance: A random sampling strategy was used to ensure data representativeness; Equal positive and negative sample data sets were constructed. After processing, 2821 toxic and non-toxic peptides with matching lengths were screened out from the original 6861 toxic peptides and 9183 non-toxic peptides.

[0012] like Figure 3 As shown in the figure, the frequency distribution analysis of amino acids was used to study the frequency variation trend of 20 amino acids in the length range of 4-50 amino acids, and a threshold was set at 25 amino acid positions. Figure 4 The SpepToxPred model development framework shown in the figure includes the following feature engineering system: Sequence feature extraction: Amino acid composition characteristics: Amino acid composition (AAC): calculate the frequency of occurrence of each amino acid in the sequence; Dipeptide composition (DPC): analysis of the distribution of adjacent amino acid pairs; Tripeptide composition (TPC): studies the combination pattern of three consecutive amino acids; Grouped amino acid composition (GAAC): statistics of amino acid groups based on physicochemical properties; Grouped dipeptide composition (GDPC): distribution characteristics of grouped amino acid pairs; Grouped tripeptide composition (GTPC): distribution characteristics of grouped amino acid triplets.

[0013] Physicochemical characteristics: Composition-transformation-distribution descriptors (CTDC, CTDT, CTDD): describe the distribution of physicochemical properties of the sequence; Combined triplet descriptor (Ctriad): Analyze the physicochemical characteristics of amino acid triplets; Enhanced Amino Acid Composition (EAAC): an improved method for describing amino acid composition; Enhanced Grouped Amino Acid Composition (EGAAC): Optimized grouped amino acid statistics.

[0014] Sequence encoding features: K-spaced amino acid pair composition (CKSAAP): analyzes the distribution of non-adjacent amino acid pairs; BLOSUM62 encoding: sequence encoding using substitution matrix; Dipeptide Deviation Encoding (DDE): Calculates the deviation characteristics of dipeptide composition; Pseudo amino acid composition (PAAC): compositional features that integrate sequence order information; Amphipathic pseudo amino acid composition (APAAC): sequence features that take into account hydrophobicity; Z-scale descriptors: Multidimensional description of amino acid properties.

[0015] Integrated learning system: Model construction: Basic classifier configuration: Random Forest (RF): weight 0.3, mainly responsible for feature selection; LightGBM: weight 0.1, providing fast training capabilities; XGBoost: weight 0.2, providing prediction stability; K nearest neighbor (KNN): weight 0.2, processing local feature relationships; Logistic regression (LR): weight 0.2, providing a linear classification benchmark.

[0016] Training strategy: Grid search hyperparameter optimization: systematically explore the optimal parameter combination; Feature standardization and selection: ensuring feature scale consistency; 10-fold cross validation: evaluate the generalization ability of the model.

[0017] like Figure 5 As shown, the performance evaluation includes: Prediction performance indicators: Matthews correlation coefficient (MCC=0.7019): evaluates classification performance; Accuracy (Accuracy=0.8445): measures the overall prediction accuracy; Precision (Precision=0.9258): evaluates the accuracy of positive example prediction; Specificity (Specificity=0.9399): Evaluates the ability to identify negative examples.

[0018] Biological Validation: Cytotoxicity evaluation: The experiment used 73 non-toxic short peptides predicted by the SpepToxPred model as the research objects. The cytotoxicity evaluation was performed using the CCK-8 colorimetric method. The cells were cultured at 37°C and 5% CO2 for 24 hours, and the peptides were used at a concentration of 100 μM to treat human bronchial epithelial cell line (BEAS-2B), human embryonic kidney cell line (HEK293T), human pancreatic ductal epithelial cell line (HPNE), and human umbilical vein endothelial cell line (HUVEC). The experimental results showed that the survival rates of the four cell lines treated with all the tested peptides remained above 90%, indicating that these peptides had no obvious cytotoxicity to normal human cells.

[0019] Hemolytic activity evaluation: The standard hemolytic activity assay was used, and the cells were incubated at 37°C for 1 hour. Fresh mouse red blood cells were used as the test object, phosphate buffered saline (PBS, pH 7.4) was used as the negative control, and 1% Triton X-100 was used as the positive control. The results showed that the hemolytic rate of all tested peptides at a concentration of 100 μM was less than 1.5%, which was significantly lower than the safety threshold of 5%, confirming that these peptides have good blood compatibility.

[0020] The above biological validation experimental results strongly support the reliability and practical value of the SpepToxPred model in predicting non-toxic short peptides.

Claims

1. A short peptide toxicity prediction method based on ensemble learning, characterized in that: The method comprises the following steps: Build a multi-source data integration system: Obtain training data from prediction model datasets (ToxGIN, ToxinPred 3.0, ToxTeller, ToxIBTL); Obtain supplementary data from professional databases (Conoserver, DRAMP 3.0, CAMPR3); Data preprocessing: Eliminate sequences containing non-standard amino acid residues; The sequence length is limited to no more than 25 amino acids; A random sampling strategy was used to construct a balanced data set; Build a feature engineering system; Sequence toxicity prediction using an ensemble learning system.

2. The method according to claim 1, characterized in that The sequence feature engineering system includes: Amino acid composition characteristics: Amino acid composition (AAC); dipeptide composition (DPC); tripeptide composition (TPC); Grouped amino acid composition (GAAC); Grouped dipeptide composition (GDPC); Grouped tripeptide composition (GTPC); Physicochemical characteristics: Composition-Transformation-Distribution Descriptors (CTDC, CTDT, CTDD); Combined triplet descriptors (Ctriad); Enhanced amino acid composition (EAAC); Enhanced Grouped Amino Acid Composition (EGAAC); Sequence encoding features: K-spacer amino acid pair composition (CKSAAP); BLOSUM62 encoding; dipeptide deviation encoding (DDE); pseudo amino acid composition (PAAC); amphipathic pseudoamino acid composition (APAAC); Z-scale descriptor.

3. The method according to claim 1, characterized in that The integrated learning system includes: Basic classifier and its optimization strategy: Random Forest (RF, weight 0.3); LightGBM (weight 0.1); XGBoost (weight 0.2); K nearest neighbor (KNN, weight 0.2); Logistic regression (LR, weight 0.2); Training strategy: Grid search hyperparameter optimization; feature standardization and selection; 10-fold cross validation.

4. The method according to claim 1, characterized in that The performance evaluation system includes: Prediction performance evaluation Matthews correlation coefficient (MCC=0.7019); Accuracy (Accuracy=0.8445); Precision (Precision=0.9258); Specificity (Specificity=0.9399).

5. A system for implementing the method according to any one of claims 1 to 4, characterized in that: include: Data preprocessing module: realize sequence standardization and data set balance; Feature extraction module: realizes multi-dimensional sequence feature calculation; Model training module: realizes the construction of integrated learning system; Prediction and evaluation module: realizes toxicity prediction probability output.

Citation Information

Patent Citations

  • Method for predicting the toxicity of polypeptides

    CN111128295A

  • Toxicity prediction model construction method, prediction model, prediction method and device

    CN115662538A

  • Toxicity prediction method and system based on deep integrated machine learning model

    CN116541785A