Angiotensin converting enzyme inhibitory peptide screening method and system based on machine learning
By constructing a multidimensional feature-based XGBoost model and using SHAP analysis, the efficiency and accuracy issues of existing ACE inhibitory peptide screening methods were resolved. This enabled high-throughput screening and interpretability of ultrashort peptides, improved peptide drug design capabilities, and provided natural drug leads for the prevention and treatment of hypertension and cardiovascular diseases.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DALIAN POLYTECHNIC UNIVERSITY
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-21
AI Technical Summary
Existing methods for screening ACE inhibitory peptides suffer from low throughput, long cycles, and high costs. Furthermore, there is a lack of efficient identification tools for ultrashort peptides, and existing prediction models lack accuracy and interpretability, making it difficult to achieve efficient screening and guide the design of bioactive peptides.
We constructed an XGBoost machine learning model based on multidimensional features and combined it with SHAP interpretability analysis to achieve high-throughput virtual screening of 2-5 peptides. By extracting the physicochemical and sequence features of peptide sequences, we optimized the model to improve prediction accuracy and interpretability.
This study enabled efficient screening of ultrashort ACE inhibitory peptides, reduced experimental costs and time, improved peptide drug design capabilities, and provided more potential natural drug lead compounds for the prevention and treatment of hypertension and cardiovascular diseases.
Smart Images

Figure CN121905282A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a machine learning-based method and system for screening angiotensin-converting enzyme inhibitory peptides, belonging to the field of bioinformatics technology. Background Technology
[0002] Hypertension is a chronic cardiovascular disease characterized by persistently elevated arterial blood pressure and is one of the leading risk factors for cardiovascular and cerebrovascular events and death worldwide. The renin-angiotensin-aldosterone system (RAAS) plays a central role in blood pressure regulation, with angiotensin-converting enzyme (ACE) catalyzing the conversion of angiotensin I into angiotensin II, which has a potent vasoconstrictive effect, and serving as a key target for blood pressure regulation. Therefore, inhibiting ACE activity has become an important strategy for the clinical treatment of hypertension.
[0003] While commonly used ACE inhibitors (such as captopril and enalapril) have proven efficacy, they can cause side effects such as cough and angioedema. This has prompted the research community to focus on discovering safer ACE-inhibiting peptides from natural food proteins. These bioactive peptides typically consist of 2-20 amino acids and have the potential advantages of good absorption and low side effects. However, traditional screening methods (such as biochemical separation and in vitro activity verification) have low throughput, long cycles, and high costs, and lack efficient means for discovering ultrashort peptides (especially 2-5 peptides) with small molecular weights and easy absorption.
[0004] In recent years, machine learning techniques have been applied to the prediction and screening of bioactive peptides to improve discovery efficiency. Existing studies mostly rely on features such as amino acid composition and physicochemical properties, using models like support vector machines, random forests, or deep learning for prediction. However, most existing models or databases are not specifically optimized for ultrashort peptides, resulting in insufficient comprehensiveness and specificity in feature characterization. This leads to limited prediction accuracy, weak generalization ability, and a lack of reasonable explanation for the key physicochemical driving mechanisms of peptide-target interactions, thus restricting their effectiveness in practical screening.
[0005] In summary, the current screening and research of ACE inhibitory peptides mainly faces the following technical bottlenecks: (1) Screening efficiency and cost issues: Traditional experimental screening methods have low throughput, long cycle and high resource consumption, making it difficult to achieve high-throughput mining of large-scale peptide libraries.
[0006] (2) Insufficient targeting of ultrashort peptide screening: Existing prediction models and databases are mostly designed for longer peptide segments, and there is a lack of efficient identification tools specifically for ultrashort active peptides with 2-5 amino acids that are easily absorbed, resulting in blind spots in the discovery of this important category of peptides.
[0007] (3) The accuracy and interpretability of the prediction model are limited: the feature representation of the existing machine learning prediction method has failed to fully integrate the key physicochemical and structural information that determines the ACE inhibitory activity, and the prediction accuracy and generalization ability of the model need to be improved. At the same time, the model is mostly a "black box", its decision-making mechanism is not transparent, it is difficult to reveal the key factors affecting the activity, and it cannot effectively guide the rational design of active peptides.
[0008] Therefore, developing a novel screening system that combines computation and experimentation to efficiently and accurately screen ultrashort ACE inhibitory peptides and explain their mechanistic origins is of significant scientific research value and practical application significance. Summary of the Invention
[0009] To address the shortcomings of existing technologies, this invention provides a machine learning-based method and system for screening angiotensin-converting enzyme (ACE) inhibitory peptides. By extracting multidimensional physicochemical and sequence features of peptide sequences, an XGBoost machine learning prediction model is constructed and optimized, enabling high-throughput virtual screening of 2-5 peptides and efficiently discovering highly active ultrashort ACE inhibitory peptides.
[0010] Specifically, the following technical solutions are included: In a first aspect, the present invention provides a machine learning-based method for screening angiotensin-converting enzyme inhibitory peptides, comprising the following steps: Step 1: Construct a peptide sequence dataset, which includes known ACE-repressing peptides and non-repressing peptides; Step 2: Extract multidimensional features from the peptide sequence, including physicochemical properties and sequence composition features; Step 3: Use the multidimensional features to train and optimize the machine learning model to obtain the ACE inhibitory peptide prediction model; Step 4: Use the ACE inhibitory peptide prediction model to predict the activity of the peptide sequences to be screened and perform virtual screening to obtain candidate ACE inhibitory peptides.
[0011] In one embodiment of the present invention, in step 2, the physicochemical properties are obtained by computational chemistry toolkit and screened by low variance filtering and Euclidean distance-based separability verification; the sequence composition features are selected from at least one of amino acid composition, dipeptide composition, and pseudo-amino acid composition.
[0012] In one embodiment of the present invention, in step 3, the machine learning model is an XGBoost model, and the model has an accuracy of not less than 95% and a recall of not less than 97%; in step 4, the peptide sequence to be screened is an ultrashort peptide with a length of 2 to 5 amino acid residues.
[0013] In one embodiment of the present invention, the method further includes: step 5, using the SHAP interpretability analysis method to interpret the ACE inhibitory peptide prediction model in order to identify key features affecting ACE inhibitory activity.
[0014] In one embodiment of the present invention, step 1 specifically includes: Data sources: BIOPEP-UWM, Ferm FooDB, PlantPepDB, DFBP public database; Sample labeling: Experimentally validated ACE inhibitory peptide sequences were collected and defined as positive samples; to ensure the reliability of negative samples, peptides with other known biological functions but not reported to have ACE inhibitory activity were randomly selected from the DFBP database and defined as negative samples; all sequences were standardized using single-letter amino acid codes. Dataset partitioning: The merged dataset is randomly divided into a training set and an independent test set in an 8:2 ratio to ensure that no test set data is encountered during the training process.
[0015] In one embodiment of the present invention, step 2 specifically includes: For each peptide sequence, two main categories of features are extracted: A. Physicochemical properties: Preliminary descriptor calculation: RDKit was used to calculate 208 two-dimensional molecular descriptors for each peptide, covering multiple dimensions such as topology, charge distribution, and hydrophobicity; Redundant feature removal: First, a low variance filtering method was applied to remove descriptors with zero or close to zero variance, reducing the number of features to 141. Separability-driven feature selection: The inter-class discriminative power of each feature is evaluated based on Euclidean distance; the distance between each pair of positive and negative samples is calculated on a single feature dimension and sorted by average distance; the performance of simple classifiers under different feature subset sizes is observed using forward or recursive elimination strategies, and the optimal number of features is determined to be the following 15: fr_Ndealkylation2, fr_Nhpyrrole, fr_SH, fr_aldehyde, fr_amide, fr_benzene, fr_bicyclic, fr_guanido, fr_imidazole, fr_para_hydroxylation, fr_phenol, fr_phenol_noOrthoHbond, fr_priamide, fr_sulfide, fr_unbrch_alkane; Domain-knowledge-based feature optimization: Referring to literature on the structure-activity relationship of ACE inhibitory peptides, five features with weak correlation in known studies were manually removed: fr_phenol_noOrthoHbond, fr_priamide, fr_aldehyde, fr_Ndealkylation2, and fr_unbrch_alkane. Seven features widely considered to be key were added: MinEStateIndex, FpDensityMorgan2, Avglpc, MaxEStateIndex, FpDensityMorgan3, MaxABsEStateIndex, and FpDensityMorgan1. Finally, a set of 17 highly discriminative physicochemical features was formed. B. Sequence composition and evolutionary characteristics: Basic compositional characteristics: Calculate amino acid composition and dipeptide composition; Advanced sequence features: calculating pseudo-amino acid composition, dipeptide deviation from expected mean, adaptive jumping dinucleotide composition, and composition and transition portions in the composition-transition-distribution descriptor; Feature dimensionality reduction and fusion: For the four high-dimensional feature sets, namely amino acid composition, dipeptide composition, dipeptide deviation from expected mean, and adaptive jumping dinucleotide composition, the separability verification based on Euclidean distance is used to sort them, and the four features with the highest discriminative power are selected from each feature set. Finally, the optimized 17-dimensional physicochemical features and 16-dimensional sequence features are concatenated to form a 33-dimensional comprehensive feature vector, which serves as the digital representation of each peptide.
[0016] In one embodiment of the present invention, step 3 specifically includes: Model selection: Five algorithms that perform well on structured data were selected for comparison: XGBoost, LightGBM, CatBoost, GBDT, and MLP; Training and validation strategy: Five-fold cross-validation is used on the training set; for each model, grid search is used to optimize it within a predefined hyperparameter space. XGBoost / LightGBM / CatBoost / GBDT: Mainly adjust max_depth, learning_rate, n_estimators, and regularization parameters; MLP: Adjust the number of hidden layers, the number of neurons per layer, the activation function, the optimizer, and the dropout rate; Early stopping mechanism: Set up validation set performance monitoring during training. When the performance no longer improves within a certain number of consecutive rounds, terminate training early to prevent overfitting. The screening method also includes a comprehensive evaluation of model performance, specifically: evaluating the model comprehensively on an independent test set using accuracy, precision, recall, and F1 score.
[0017] Secondly, the present invention provides a machine learning-based angiotensin-converting enzyme (ACE) inhibitory peptide screening system for implementing the aforementioned machine learning-based ACE inhibitory peptide screening method, the system comprising: A peptide sequence dataset construction module is used to construct peptide sequence datasets containing known ACE-repressing and non-repressing peptides. A multidimensional feature extraction module is used to extract multidimensional features from the peptide sequence, the multidimensional features including physicochemical properties and sequence composition features; The model training and optimization module is used to train and optimize the machine learning model using the multidimensional features to obtain the ACE inhibitory peptide prediction model. The activity prediction and virtual screening module is used to perform activity prediction and virtual screening of the peptide sequences to be screened using the ACE inhibitory peptide prediction model, so as to obtain candidate ACE inhibitory peptides.
[0018] Thirdly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the machine learning-based angiotensin-converting enzyme inhibitory peptide screening method.
[0019] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the machine learning-based angiotensin-converting enzyme inhibitory peptide screening method.
[0020] The beneficial effects achieved by this invention are as follows: This invention provides a machine learning-based method and system for screening angiotensin-converting enzyme (ACE) inhibitory peptides. By extracting multidimensional physicochemical and sequence features of peptide sequences, and constructing and optimizing an XGBoost machine learning prediction model, it can achieve high-throughput virtual screening of 2-5 peptides, efficiently discovering highly active ultrashort ACE inhibitory peptides. By establishing an efficient machine learning model for rapid virtual screening of massive amounts of food protein peptide sequences, experimental costs and time can be significantly reduced, accelerating the discovery of safe, efficient, and easily absorbed ultrashort ACE inhibitory peptides. This method not only improves the rational design capability of peptide drugs but also provides methodological guidance for the discovery of other bioactive peptides, ultimately expanding the potential natural functional food ingredients and drug lead compounds for the prevention and treatment of hypertension and cardiovascular diseases. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the architecture of the ACE-XGBoost model of this invention.
[0022] Figure 2 This is a SHAP abstract diagram of the ACE inhibitory peptide prediction model of this invention.
[0023] Figure 3 The SHAP force maps used in this invention to interpret the prediction results of ACE inhibitory peptides are RY (A), QR (B), SRS (C), ARPP (D), GKGSW (E), and HYSSW (F).
[0024] Figure 4 The figure shows the ACE inhibitory activity assay results of the synthesized polypeptide of this invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Example 1 like Figure 1 As shown, this embodiment provides a machine learning-based method for screening angiotensin-converting enzyme inhibitory peptides, including the following steps: Step 1: Construct a peptide sequence dataset, which includes known ACE-repressing peptides and non-repressing peptides; Step 2: Extract multidimensional features from the peptide sequence, including physicochemical properties and sequence composition features; Step 3: Use the multidimensional features to train and optimize the machine learning model to obtain the ACE inhibitory peptide prediction model; Step 4: Use the ACE inhibitory peptide prediction model to predict the activity of the peptide sequences to be screened and perform virtual screening to obtain candidate ACE inhibitory peptides.
[0027] Optionally, in step 2, the physicochemical properties are obtained through a computational chemistry toolkit and screened by low variance filtering and Euclidean distance-based separability verification; the sequence composition features are selected from at least one of amino acid composition, dipeptide composition, and pseudo-amino acid composition.
[0028] Optionally, in step 3, the machine learning model is an XGBoost model, and the model's accuracy is not less than 95% and its recall is not less than 97%.
[0029] Optionally, in step 4, the peptide sequence to be screened is an ultrashort peptide with a length of 2 to 5 amino acid residues.
[0030] Optionally, the method further includes: step 5, using the SHAP interpretability analysis method to interpret the ACE inhibitory peptide prediction model in order to identify key features affecting ACE inhibitory activity.
[0031] Example 2 This embodiment provides a machine learning-based angiotensin-converting enzyme (ACE) inhibitory peptide screening system to implement the machine learning-based ACE inhibitory peptide screening method described in Embodiment 1. The system includes: A peptide sequence dataset construction module is used to construct peptide sequence datasets containing known ACE-repressing and non-repressing peptides. A multidimensional feature extraction module is used to extract multidimensional features from the peptide sequence, the multidimensional features including physicochemical properties and sequence composition features; The model training and optimization module is used to train and optimize the machine learning model using the multidimensional features to obtain the ACE inhibitory peptide prediction model. The activity prediction and virtual screening module is used to perform activity prediction and virtual screening of the peptide sequences to be screened using the ACE inhibitory peptide prediction model, so as to obtain candidate ACE inhibitory peptides.
[0032] The angiotensin-converting enzyme inhibitory peptide obtained by the method described in Example 1 above, or by the system screened by the method described in Example 2 above, can be used in drugs or health foods for the prevention or treatment of hypertension.
[0033] Example 3 Construction and Evaluation of an ACE Repressor Peptide Prediction Model Based on Multidimensional Feature Fusion and XGBoost. The core of this embodiment lies in constructing a high-precision, high-generalization machine learning model for predicting ACE repressor peptides from short peptides (2-5 peptides). The specific steps are as follows: S1. System construction and preprocessing of the dataset: Data source: Integration from public databases—BIOPEP-UWM, Ferm FooDB, PlantPepDB, DFBP—and relevant research literature up to November 2024.
[0034] Sample labeling: Experimentally validated ACE-inhibiting peptide sequences were collected and defined as positive samples, totaling 1376. To ensure the reliability of negative samples, 1021 peptides with other known biological functions (such as antibacterial and antioxidant properties) but not reported to have ACE-inhibiting activity were randomly selected from the DFBP database and defined as negative samples. All sequences were standardized using single-letter amino acid codes.
[0035] Dataset partitioning: The merged total dataset (2397 peptides) was randomly partitioned into a training set (1918 peptides) and an independent test set (479 peptides) in an 8:2 ratio to ensure that no test set data was encountered during the training process.
[0036] S2. Engineered extraction and optimized screening of multidimensional features: For each peptide sequence, the system extracts two main categories of features: A. Physicochemical properties: Preliminary descriptor calculation: RDKit (version 2022.09.5) was used to calculate 208 two-dimensional molecular descriptors for each peptide, covering multiple dimensions such as topology, charge distribution, and hydrophobicity. Redundant feature removal: First, a low-variance filtering method was applied to remove descriptors with zero or near-zero variance, reducing the number of features to 141.
[0037] Separability-driven feature selection: The inter-class discriminative power of each feature is evaluated based on Euclidean distance. The distance between each pair of positive and negative samples is calculated along a single feature dimension and sorted by average distance. Forward or recursive elimination strategies are employed to observe the performance of simple classifiers (such as logistic regression) at different feature subset sizes, determining the optimal number of features to be the following 15: fr_Ndealkylation2, fr_Nhpyrrole, fr_SH, fr_aldehyde, fr_amide, fr_benzene, fr_bicyclic, fr_guanido, fr_imidazole, fr_para_hydroxylation, fr_phenol, fr_phenol_noOrthoHbond, fr_priamide, fr_sulfide, and fr_unbrch_alkane.
[0038] Domain-knowledge-based feature optimization: Referring to literature on the structure-activity relationship of ACE inhibitory peptides, five features with weak correlation in known studies were manually removed: fr_phenol_noOrthoHbond, fr_priamide, fr_aldehyde, fr_Ndealkylation2, and fr_unbrch_alkane. Seven features widely considered to be key were added: MinEStateIndex, FpDensityMorgan2, Avglpc, MaxEStateIndex, FpDensityMorgan3, MaxABsEStateIndex, and FpDensityMorgan1. Finally, a set of 17 highly discriminative physicochemical features was formed.
[0039] B. Sequence composition and evolutionary characteristics: Basic composition characteristics: calculate amino acid composition and dipeptide composition.
[0040] Advanced sequence features: Calculate pseudo-amino acid composition (parameters λ=2, ω=0.05), dipeptide deviation from expected mean, adaptive jumping dinucleotide composition, and composition and transition portions in the composition-transition-distribution descriptor.
[0041] Feature dimensionality reduction and fusion: For the four high-dimensional feature sets—amino acid composition, dipeptide composition, dipeptide deviation from the expected mean, and adaptive jumping dinucleotide composition—separability verification based on Euclidean distance was used for sorting. The four features with the highest discriminative power from each feature set (a total of 16 features) were selected. Finally, the optimized 17-dimensional physicochemical features and 16-dimensional sequence features were concatenated to form a 33-dimensional comprehensive feature vector, which serves as the digital representation of each peptide.
[0042] S3. Comparison, training, and hyperparameter optimization of machine learning models: Model selection: Five algorithms that perform well on structured data were selected for comparison: XGBoost, LightGBM, CatBoost, GBDT (Gradient Boosting Decision Tree) and MLP (Multilayer Perceptron).
[0043] Training and validation strategy: Five-fold cross-validation is used on the training set. For each model, grid search is used to optimize it within a predefined hyperparameter space. XGBoost / LightGBM / CatBoost / GBDT: The main adjustments are to max_depth (3-10), learning_rate (0.01-0.3), n_estimators (100-500), and regularization parameters (such as reg_alpha, reg_lambda).
[0044] MLP: Adjust the number of hidden layers (1-3), number of neurons per layer (64-256), activation function (ReLU, tanh), optimizer (Adam, SGD), and dropout rate ( Figure 1 ).
[0045] Early stopping mechanism: Set up validation set performance monitoring during training. When the performance no longer improves within a certain number of consecutive rounds, training is terminated early to prevent overfitting.
[0046] S4. Comprehensive evaluation and interpretability analysis of model performance: Performance metrics: The model is comprehensively evaluated on a separate test set using accuracy, precision, recall, and F1 score.
[0047] Results: As shown in Table 1, the XGBoost model performed best among all models, achieving an accuracy of 95%, precision of 94.5%, and recall of 97.17%. High recall means the model has a very low risk of missing real ACE inhibitory peptides, which is crucial for screening applications.
[0048] Table 1 Predictive performance of different machine learning models on the test set.
[0049] Model interpretability (SHAP analysis): To understand the basis of the model's decisions, the SHAP framework is used for analysis. For example... Figure 2 and Figure 3 As shown, the most important features affecting model predictions are FpDensityMorgan1, which describes molecular topological complexity, and the pseudo-amino acid composition (PseAAC) feature, which characterizes sequence order information. Meanwhile, SHAP value analysis clearly shows a significant negative correlation between peptide chain length and the probability of being predicted as an ACE inhibitory peptide. That is, the model has "learned" and confirmed the rule that shorter peptide chains (2-5 peptides) are more likely to be highly efficient ACE inhibitory peptides, providing strong data-driven support for the focusing strategy of this invention.
[0050] Example 4 Virtual high-throughput screening and in vitro biochemical activity verification based on a predictive model. This embodiment aims to apply a trained model for large-scale virtual screening and verify the reliability of the screening results through standard in vitro experiments.
[0051] 1. Construction of a virtual screening library and high-throughput prediction: Peptide library generation: Theoretically, all peptide sequences of length 2, 3, 4, and 5 composed of 20 natural amino acids are enumerated to form a huge virtual library of ultrashort peptides (2 peptides: 400, 3 peptides: 8000, 4 peptides: 160,000, 5 peptides: 3,200,000).
[0052] Feature Calculation and Prediction: Using the same procedure as in Example 3, 33-dimensional features were calculated for each peptide in the virtual library. The feature matrix was then input into the saved optimal XGBoost model for batch prediction, outputting a "suppression peptide probability" score between 0 and 1 for each peptide sequence.
[0053] Identification of high-potential candidate peptides: A high probability threshold was set to rapidly screen hundreds of high-scoring candidate peptides from millions of sequences. Based on the score ranking, the top 6 peptides were selected for subsequent experiments, including RY, QR, SRS, ARPP, GKGSW, and HYSSW.
[0054] 2. Chemical synthesis and quality control of candidate peptides: Solid-phase synthesis: The above 6 candidate peptides were synthesized by a professional peptide synthesis company using the solid-phase synthesis method.
[0055] 3. Preliminary screening and dose-response study of in vitro ACE inhibitory activity: Initial activity screening (single concentration test). The reaction system consisted of: 30 μL of 5 mM substrate hippuryl-histyl-leucine (HHL, dissolved in 0.1 M HEPES buffer containing 0.3 M NaCl, pH 8.3), 20 μL of the test peptide solution (final concentration 1 mM), and 20 μL of ACE enzyme solution (0.1 U). After shaking in a 37 ℃ water bath for 40 min, the reaction was terminated by adding 60 μL of 1 M HCl. The peak area of the product hippuric acid was quantitatively determined by HPLC (C18 column, mobile phase: water / acetonitrile containing 0.05% trifluoroacetic acid = 75 / 25, flow rate: 0.5 mL / min, detection wavelength: 228 nm).
[0056] The results are as follows Figure 4 As shown, at a concentration of 1 mM, QR is 60.72% and PY is 55.38%.
[0057] In summary, for specific protein resources, the aim is to systematically discover short ACE inhibitory peptides with the potential to lower blood pressure. Traditional methods rely on large-scale enzymatic digestion and random screening, which are inefficient and difficult to focus on easily absorbed 2-5 peptides. This invention provides a machine learning-based method and system for screening angiotensin-converting enzyme inhibitory peptides. By extracting multidimensional physicochemical and sequence features of peptide sequences and constructing and optimizing an XGBoost machine learning prediction model, high-throughput virtual screening of 2-5 peptides can be achieved, efficiently discovering highly active ultrashort ACE inhibitory peptides. By establishing an efficient machine learning model for rapid virtual screening of massive amounts of food protein peptide sequences, experimental costs and time can be significantly reduced, accelerating the discovery of safe, efficient, and easily absorbed ultrashort ACE inhibitory peptides. This method not only improves the rational design capability of peptide drugs but also provides methodological guidance for the discovery of other bioactive peptides, ultimately expanding the potential natural functional food ingredients and drug lead compounds for the prevention and treatment of hypertension and cardiovascular diseases.
[0058] Furthermore, the present invention provides a computer device that may include a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it causes the processor to perform the steps of the machine learning-based angiotensin-converting enzyme inhibitory peptide screening method as described in any of the above embodiments.
[0059] The working process, working details, and technical effects of the computer device provided in this embodiment can be found in the embodiment above regarding the screening method for angiotensin-converting enzyme inhibitory peptides based on machine learning, and will not be repeated here.
[0060] Furthermore, the present invention provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the machine learning-based angiotensin-converting enzyme inhibitory peptide screening method as described in any of the above embodiments. The computer-readable storage medium refers to a data storage medium, which may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives, and / or memory sticks, etc. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0061] The working process, working details, and technical effects of the computer-readable storage medium provided in this embodiment can be found in the embodiment above regarding the screening method for angiotensin-converting enzyme inhibitory peptides based on machine learning, and will not be repeated here.
[0062] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).
[0063] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A machine learning-based method for screening angiotensin-converting enzyme inhibitory peptides, characterized in that, Includes the following steps: Step 1: Construct a peptide sequence dataset, which includes known ACE-repressing peptides and non-repressing peptides; Step 2: Extract multidimensional features from the peptide sequence, including physicochemical properties and sequence composition features; Step 3: Use the multidimensional features to train and optimize the machine learning model to obtain the ACE inhibitory peptide prediction model; Step 4: Use the ACE inhibitory peptide prediction model to predict the activity of the peptide sequences to be screened and perform virtual screening to obtain candidate ACE inhibitory peptides.
2. The machine learning-based method for screening angiotensin-converting enzyme inhibitory peptides according to claim 1, characterized in that, In step 2, the physicochemical properties are obtained through a computational chemistry toolkit and screened by low variance filtering and Euclidean distance-based separability verification; the sequence composition features are selected from at least one of amino acid composition, dipeptide composition, and pseudo-amino acid composition.
3. The machine learning-based method for screening angiotensin-converting enzyme inhibitory peptides according to claim 1, characterized in that, In step 3, the machine learning model is the XGBoost model, and the model's accuracy is not less than 95% and its recall is not less than 97%; in step 4, the peptide sequence to be screened is an ultrashort peptide with a length of 2 to 5 amino acid residues.
4. The machine learning-based method for screening angiotensin-converting enzyme inhibitory peptides according to claim 1, characterized in that, It also includes: Step 5, using the SHAP interpretability analysis method to interpret the ACE inhibitory peptide prediction model in order to identify key features affecting ACE inhibitory activity.
5. The machine learning-based method for screening angiotensin-converting enzyme inhibitory peptides according to claim 1, characterized in that, Step 1 specifically includes: Data sources: BIOPEP-UWM, Ferm FooDB, PlantPepDB, DFBP public database; Sample labeling: Experimentally validated ACE inhibitory peptide sequences were collected and defined as positive samples; to ensure the reliability of negative samples, peptides with other known biological functions but not reported to have ACE inhibitory activity were randomly selected from the DFBP database and defined as negative samples; all sequences were standardized using single-letter amino acid codes. Dataset partitioning: The merged dataset is randomly divided into a training set and an independent test set in an 8:2 ratio to ensure that no test set data is encountered during the training process.
6. The machine learning-based method for screening angiotensin-converting enzyme inhibitory peptides according to claim 5, characterized in that, Step 2 specifically includes: For each peptide sequence, two main categories of features are extracted: A. Physicochemical properties: Preliminary descriptor calculation: RDKit was used to calculate 208 two-dimensional molecular descriptors for each peptide, covering multiple dimensions such as topology, charge distribution, and hydrophobicity; Redundant feature removal: First, a low variance filtering method was applied to remove descriptors with zero or close to zero variance, reducing the number of features to 141. Separability-driven feature selection: The inter-class discriminative power of each feature is evaluated based on Euclidean distance; the distance between each pair of positive and negative samples is calculated on a single feature dimension and sorted by average distance; the performance of simple classifiers under different feature subset sizes is observed using forward or recursive elimination strategies, and the optimal number of features is determined to be the following 15: fr_Ndealkylation2, fr_Nhpyrrole, fr_SH, fr_aldehyde, fr_amide, fr_benzene, fr_bicyclic, fr_guanido, fr_imidazole, fr_para_hydroxylation, fr_phenol, fr_phenol_noOrthoHbond, fr_priamide, fr_sulfide, fr_unbrch_alkane; Domain-knowledge-based feature optimization: Referring to literature on the structure-activity relationship of ACE inhibitory peptides, five features with weak correlation in known studies were manually removed: fr_phenol_noOrthoHbond, fr_priamide, fr_aldehyde, fr_Ndealkylation2, and fr_unbrch_alkane. Seven features widely considered to be key were added: MinEStateIndex, FpDensityMorgan2, Avglpc, MaxEStateIndex, FpDensityMorgan3, MaxABsEStateIndex, and FpDensityMorgan1. Finally, a set of 17 highly discriminative physicochemical features was formed. B. Sequence composition and evolutionary characteristics: Basic compositional characteristics: Calculate amino acid composition and dipeptide composition; Advanced sequence features: calculating pseudo-amino acid composition, dipeptide deviation from expected mean, adaptive jumping dinucleotide composition, and composition and transition portions in the composition-transition-distribution descriptor; Feature dimensionality reduction and fusion: For the four high-dimensional feature sets, namely amino acid composition, dipeptide composition, dipeptide deviation from expected mean, and adaptive jumping dinucleotide composition, the separability verification based on Euclidean distance is used to sort them, and the four features with the highest discriminative power are selected from each feature set. Finally, the optimized 17-dimensional physicochemical features and 16-dimensional sequence features are concatenated to form a 33-dimensional comprehensive feature vector, which serves as the digital representation of each peptide.
7. The machine learning-based method for screening angiotensin-converting enzyme inhibitory peptides according to claim 6, characterized in that, Step 3 specifically includes: Model selection: Five algorithms that perform well on structured data were selected for comparison: XGBoost, LightGBM, CatBoost, GBDT, and MLP; Training and validation strategy: Five-fold cross-validation is used on the training set; for each model, grid search is used to optimize it within a predefined hyperparameter space. XGBoost / LightGBM / CatBoost / GBDT: Mainly adjust max_depth, learning_rate, n_estimators, and regularization parameters; MLP: Adjust the number of hidden layers, the number of neurons per layer, the activation function, the optimizer, and the dropout rate; Early stopping mechanism: Set up validation set performance monitoring during training. When the performance no longer improves within a certain number of consecutive rounds, terminate training early to prevent overfitting. The screening method also includes a comprehensive evaluation of model performance, specifically: evaluating the model comprehensively on an independent test set using accuracy, precision, recall, and F1 score.
8. A machine learning-based angiotensin-converting enzyme inhibitory peptide screening system, characterized in that, For implementing the machine learning-based angiotensin-converting enzyme inhibitory peptide screening method according to any one of claims 1-7, the system comprises: A peptide sequence dataset construction module is used to construct peptide sequence datasets containing known ACE-repressing and non-repressing peptides. A multidimensional feature extraction module is used to extract multidimensional features from the peptide sequence, the multidimensional features including physicochemical properties and sequence composition features; The model training and optimization module is used to train and optimize the machine learning model using the multidimensional features to obtain the ACE inhibitory peptide prediction model. The activity prediction and virtual screening module is used to perform activity prediction and virtual screening of the peptide sequences to be screened using the ACE inhibitory peptide prediction model, so as to obtain candidate ACE inhibitory peptides.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the machine learning-based angiotensin-converting enzyme inhibitory peptide screening method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the machine learning-based angiotensin-converting enzyme inhibitory peptide screening method as described in any one of claims 1-7.