Blood brain barrier penetrating peptide prediction method based on multi-feature fusion and data enhancement
By constructing a balanced dataset and integrating multi-dimensional features, FBDiffusion was used to generate synthetic B3PPs, and CatBoost was used to filter features. Finally, LightGBM was used to build a classification model, which solved the problems of sample imbalance and insufficient feature extraction in B3PP prediction. This achieved efficient and accurate B3PP prediction, providing strong support for the development of drugs for central nervous system diseases.
Patent Information
- Application Number
- CN202511578305.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-02-03
AI Technical Summary
Existing B3PPs prediction methods suffer from problems such as imbalanced samples, incomplete feature extraction, and insufficient prediction accuracy and generalization ability, making it difficult to meet the needs of drug development for central nervous system diseases.
By constructing a balanced dataset, integrating multi-dimensional features, generating synthetic B3PPs using FBDiffusion, combining stratified sampling and CatBoost feature selection, and using LightGBM to build a classification model, we can achieve efficient and accurate prediction of B3PPs.
It significantly improves the model's prediction accuracy and generalization ability, reduces computational complexity, provides an efficient B3PPs screening tool, and supports the rapid development of drugs for the treatment of central nervous system diseases.
Smart Images

Figure CN121459947A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics and drug design technology, specifically to a method for predicting blood-brain barrier-penetrating peptides based on multi-feature fusion and data augmentation. Background Technology
[0002] The blood-brain barrier (BBB) is a highly selective semi-permeable membrane barrier composed of brain microvascular endothelial cells, tight junctions, and astrocytes. Its core function is to maintain the homeostasis of the central nervous system (CNS) microenvironment and ensure normal brain physiological activity. However, this barrier also serves as a major obstacle to drug delivery for CNS diseases (such as Alzheimer's disease, stroke, and brain tumors)—approximately 98% of small molecule drugs and almost all large molecule drugs cannot penetrate the blood-brain barrier to enter brain tissue, severely limiting the research and application of neurotherapeutic drugs.
[0003] Blood-brain barrier-penetrating peptides (B3PPs), as a class of short peptide molecules, can cross the blood-brain barrier through endogenous transcytosis without damaging its structure, exhibiting characteristics such as high biosafety, ease of synthesis, and excellent penetration efficiency. These peptides can not only be used directly as neurotherapeutic drugs but also as carriers to deliver bioactive molecules such as proteins, small interfering RNA (siRNA), and plasmid DNA into the brain. Therefore, rapid and accurate identification of B3PPs is of great significance for the development of drugs for the treatment of central nervous system diseases.
[0004] Currently, the identification of B3PPs mainly relies on experimental methods such as phage display and radionuclide labeling. However, these methods are time-consuming, labor-intensive, and costly, failing to meet the needs of large-scale screening. To date, only a few hundred B3PPs have been experimentally validated, far from sufficient to support the needs of drug development.
[0005] To address the limitations of experimental methods, researchers have developed machine learning-based B3PP computational prediction methods. For example, Dai et al. extracted sequence features based on multiple descriptors and combined them with logistic regression to build a prediction model; Kumar et al. proposed the B3Pred tool based on random forests; Gu et al. used Borderline SMOTE for data augmentation and combined it with random forests for classification; Tang et al. used generative adversarial networks (GANs) to generate synthetic peptide sequences and combined them with Transformers and multilayer perceptrons (MLPs) to build a DeepB3P model.
[0006] However, existing calculation methods still have the following key problems: 1) Data set limitations: The small number of training samples and the imbalance between positive and negative samples (scarce B3PPs samples) result in poor model generalization ability and difficulty in adapting to peptide sequence prediction scenarios of different lengths and structures. 2) Single feature extraction: Most models rely on only single types of features such as amino acid physicochemical properties or sequence information, which cannot fully capture the key structure-function associations of B3PPs in penetrating the blood-brain barrier. This easily leads to high-dimensional redundant features, increases computational overhead, and causes overfitting. 3) Imbalance in model performance: Some models (such as B3Pred) have excessively high specificity (SP) but excessively low sensitivity (SN), resulting in a large number of potential B3PPs being misclassified as non-penetrating peptides; other models (such as SCMB3PP) have poor overall performance and cannot meet the actual screening requirements.
[0007] Therefore, developing a B3PPs prediction method that can solve the problem of sample imbalance, integrate multi-dimensional effective features, and balance prediction accuracy and computational efficiency has become the key to breaking through the bottleneck of central nervous system drug development. Summary of the Invention
[0008] To address the problems of sample scarcity and imbalance, incomplete feature extraction, and insufficient prediction accuracy and generalization ability in existing B3PPs prediction methods, this invention provides a blood-brain barrier penetration peptide prediction method based on multi-feature fusion and data augmentation. By constructing a balanced dataset, integrating multi-dimensional features, and optimizing feature selection and classification models, efficient and accurate prediction of B3PPs can be achieved.
[0009] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for predicting blood-brain barrier-penetrating peptides based on multi-feature fusion and data augmentation includes the following steps; S1. Dataset Construction and Balancing: Collect basic B3PPs and Non-B3PPs sequences, use the feedback diffusion model (FBDiffusion) to generate pseudo B3PPs (gen-B3PPs), and combine stratified sampling to supplement Non-B3PPs to construct multiple balanced training sets and independent test sets; S2. Multi-dimensional feature extraction: Features extracted from peptide sequences in the balanced training set and the independent test set are combined to generate the original feature matrix; S3. For the balanced training set constructed in S1, use all the features in the generated original feature matrix to select the optimal training set based on the model performance on the independent test set. S4. Based on CatBoost feature selection, calculate the feature importance of all features and set the optimal threshold to select a subset of effective features; S5. Build a classification model based on LightGBM, using the optimal training set determined in S3 and the effective feature subset selected in S4, and train and optimize the model using 5-fold cross-validation. S6. Extract and normalize the features of the polypeptide sequence to be predicted, input the optimized model from S5, and output the predicted category and the predicted probability value.
[0010] In one embodiment, step S1 requires the following steps: Step 1) Collect the basic dataset. Obtain 324 experimentally validated B3PPs sequences and 324 non-B3PPs sequences from the BBPpredict database as the basic training set. The B3PPs sequences and non-B3PPs sequences are kept equal. We obtained 98 experimentally validated B3PPs sequences and 99 Non-B3PPs sequences as the basic independent test set; we retrieved B3PPs-related literature published from 2018 to 2024 from the PubMed database to supplement thousands of experimentally validated Non-B3PPs sequences; and we manually verified and removed 3 contradictory sequences from the independent test set. The final basic dataset contains 324 B3PPs and 324 Non-B3PPs as the training set, and 97 B3PPs and 97 Non-B3PPs as the independent test set. Step 2) The FBDiffusion model is trained using the blood-brain barrier penetrating peptide B3PPs sequences in the basic dataset as training data. The model learns the core biological characteristics of B3PPs, including amino acid composition (such as the proportion of basic amino acids such as lysine and arginine), sequence length distribution (5-50 residues), and key physicochemical properties (such as isoelectric point and hydrophilicity index). Finally, four sets of synthetic B3PPs sequences gen-B3PPs with different numbers of sequences are generated, specifically 324, 648, 1296, and 2592, respectively, to provide data support for the positive samples in the subsequent training set.
[0011] Step 3) To construct a balanced training set without length bias, based on the sequence length distribution of non-blood-brain barrier penetrating peptides (Non-B3PPs) and B3PPs in the basic dataset, a stratified sampling method is used to select corresponding numbers of Non-B3PPs samples, as well as the gen-B3PPs samples generated in Step 2) (which together with the real B3PPs constitute positive samples). Then, by combining positive and negative samples in equal quantities, five balanced training sets (train1~train5) are finally constructed. The sample sizes of each training set are 648 (324 B3PPs + 324 Non-B3PPs), 1296, 2592, 5184, and 10368, respectively. This solves the problem of scarce B3PPs samples and class imbalance in the original data, providing high-quality data for model training.
[0012] In step 2), the FBDiffusion model is described as follows: The diffusion model represents B3PPs as latent variables. And construct a pair of forward / reverse Markov chains; the forward process is performed by progressively adding Gaussian noise to the real training samples. Perturbation is a latent variable The latent variables satisfy Specifically, the forward process in discrete time steps The upper continuous corrosion sample forms in For the first step of the forward process Step by step, the distribution of latent variables is made to approximate the standard normal distribution. Accordingly, a learnable inverse Markov chain is constructed, and noise is gradually reduced from... Reconstructed sequence Their joint distribution is in The parameters to be learned For the reverse process Step 1. During training, variational inference is employed to maximize sample likelihood, a process that is equivalent to maximizing the lower bound of evidence (ELBO). Due to the limited number of real B3PPs, underfitting is inevitable during model training. This means that the sample space reconstructed through the reverse process still differs from the real B3PPs space. The FBDiffusion model nests a sequence BLAST analyzer on top of the diffusion model to evaluate the similarity between the samples generated by the diffusion model and real B3PPs, thereby guiding the model to generate peptide sequences that are closer to real B3PPs, i.e., gen-B3PPs. In S2, the features include amino acid composition, sequence adjacency, physicochemical properties, and molecular structure features, namely, AAC, APAAC, ASDC, CKSAAP, CTD, DDE, DPC, TPC, SEP, SER, QSO, SE, SOCN, and molecular structure features. For each polypeptide sequence in the training and test sets, 14 features are extracted, specifically including: 1) Amino acid composition characteristics: AAC, APAAC, DPC, TPC, and DDE are used to describe the amino acid composition characteristics of the polypeptide sequence; 2) Sequence adjacency characteristics: The adjacency combination patterns of amino acid residues are described using QSO, SOCN, CKSAAP, and ASDC. 3) Physicochemical properties: CTD, SE, SEP (Sequence-Order Penalty), and SER descriptors are used to quantify the hydrophilicity, charge, polarity, and spatial structure of amino acids. 4) Molecular structural characteristics: Extract the electronic states, topological complexity, atomic composition and other molecular theoretical characteristics of the polypeptide to reflect the structural basis of the interaction between the polypeptide and the blood-brain barrier endothelial cells; A total of 11,200-dimensional original feature matrices were generated.
[0013] In step S3, to determine the most suitable dataset for subsequent model training and feature selection, model training and performance evaluation are carried out sequentially on the five balanced training sets of different sizes (sample sizes of 648, 1296, 2592, 5184, and 10368) constructed in S1. The steps are as follows: 1) For each training set, all 11200 original features extracted in S2 are used to train the initial classification model. Here, the LightGBM algorithm is selected as the initial classification model, and the default parameters of the algorithm are used for training to ensure that the training conditions are consistent across different training sets. 2) The model trained on each training set is tested on an independent test set. During the testing process, several key model performance metrics are monitored. include: ①AUC (Area Under the ROC Curve): The area under the ROC curve can comprehensively reflect the true positive rate and false positive rate of the model under different classification thresholds. The closer the AUC value is to 1, the better the classification performance of the model. ②Accuracy: The proportion of correctly classified samples out of the total number of samples, reflecting the overall accuracy of the model's classification; ③ Sensitivity (Sn): Also known as the true positive rate, this is the proportion of samples that are actually positive but are correctly identified as positive by the model, reflecting the model's ability to capture positive samples. The formula is: Where TP represents the number of true positives and FN represents the number of false negatives; ④ Specificity (Sp): The proportion of samples that are actually negative but are correctly identified as negative by the model, reflecting the model's ability to capture negative samples. The formula is... TN represents the number of true negatives, and FP represents the number of false positives; ⑤ Matthews correlation coefficient (MCC): The correlation coefficient that comprehensively considers true positives, true negatives, false positives, and false negatives. The formula is as follows: The range of values is The closer the value is to 1, the better the model performance. 3) By comparing and analyzing the performance of the models trained on these 5 training sets on the independent test set, the training set with the highest AUC (training set 3) is selected as the basic dataset for subsequent feature selection (S4) and model optimization (S5), laying the foundation for obtaining a high-precision blood-brain barrier penetration peptide prediction model.
[0014] The specific steps of step 1) are as follows: ①GOSS sampling: Gradient-based One-Sided Sampling (GOSS) is used to filter the training samples, retaining only key samples for subsequent calculations, as shown in the following formula: For the sample The original gradient (reflecting the degree of influence of the sample on the model loss); (Retain 20% of samples with large gradients); (Randomly sample 10% from 80% of the small gradient samples); This is the adjusted gradient.
[0015] For each training set, samples are selected according to this ratio, which not only retains large gradient samples that are key to model optimization, but also reduces the amount of computation through sampling, and the uniform values of a and b ensure that the sampling logic of different training sets is consistent. ② Histogram discretization (high-dimensional feature compression): For the original 11200-dimensional features, LightGBM discretizes continuous feature values into histograms (with 256 bins) by default, transforming high-dimensional features into low-dimensional histogram features, thus reducing memory consumption and computational complexity.
[0016] After discretization, the model selects the optimal feature split point by calculating the split gain, as shown in the following formula: These are the gradients of the left and right child nodes, respectively; These are the Hessian sums of the left and right child nodes, respectively; (L2 regularization parameter, to prevent overfitting) (Splitting cost, controlling the complexity of the tree); For each training set of 11200 features, the split gain of each feature is calculated using this formula. The feature with the largest gain and the split point are selected to construct a decision tree, ensuring training efficiency under high-dimensional features. ③ Leaf-Wise growth strategy for tree construction: The Leaf-Wise growth strategy is adopted, and the node with the largest splitting gain is selected from all leaf nodes for growth each time until the default number of iterations or the stopping condition is reached. In step S4, the CatBoost algorithm is used to filter the original high-dimensional features, remove redundant features, and reduce computational complexity. The steps are as follows: 1) Feature standardization: For the generated 11200-dimensional features, StandardScaler is used to normalize them into a standard normal distribution with a mean of 0 and a standard deviation of 1, eliminating scale interference caused by differences in the representation dimensions of different features; 2) Feature Importance Calculation: For the standardized 11200-dimensional features, the importance of each feature is calculated based on the change in the loss function during the decision tree splitting process by training a CatBoost classifier. The calculation formula is as follows: 3) Feature selection threshold optimization: Set the feature importance threshold range to 0~0.00001, and use the AUC, ACC, SN, SP, and MCC of the independent test set as evaluation indicators to select the optimal threshold; When the threshold is 0.0001, the number of features decreases from 11200 to 1211 (compression rate 89.2%), and the test set AUC decreases by only 0.04% (from 0.8857 to 0.8853), balancing prediction accuracy and computational efficiency. 4) Feature set construction after filtering: retain features with an importance of not less than 0.0001 to form an optimized feature subset for subsequent model training.
[0017] S5 is based on the construction and training of the B3PPs classification model using LightGBM; A B3PPs classification model is constructed using the LightGBM algorithm. Its gradient one-sided sampling (GOSS), histogram discretization, and leaf-wise growth strategies are utilized to improve model training efficiency and prediction accuracy. The steps are as follows: 1) Model parameter initialization: Set key parameters for LightGBM; gradient one-sided sampling GOSS, large gradient sample ratio. Small gradient sample ratio Histogram discretization and Leaf-Wise growth strategy, regularization parameters Cost of generating new leaf nodes Learning rate 0.05, maximum depth 30, maximum number of leaf nodes 64, number of iterations 1000; 2) Model Training: Using the feature subset selected in step S3 as input, the LightGBM model is trained using 5-fold cross-validation. During training, the GOSS technique is used to retain samples with large gradients and randomly sample samples with small gradients. At the same time, gradient weights are adjusted using (Formula 2) to reduce computation. Continuous features are converted into discrete intervals using the histogram discretization method. First, the number of discrete intervals is preset (256). Then, the value distribution of each continuous feature in the training set is statistically analyzed. The feature value range is evenly divided into the preset number of intervals according to the value size. Finally, the specific value of each continuous feature is mapped to the corresponding discrete interval to complete the conversion and reduce memory consumption. The Leaf-Wise growth strategy is used to select the leaf node with the largest splitting gain for growth, which accelerates model convergence. 3) Model optimization: Early stopping strategy is adopted to prevent overfitting - when the validation set AUC has no improvement for 20 consecutive rounds, training is stopped and the current optimal model parameters are saved.
[0018] The steps for predicting and outputting a B3PPs in S6 are as follows: 1) Input sequence preprocessing: For the polypeptide sequence to be predicted, generate the corresponding multi-dimensional feature vector according to the feature extraction method in step S2, and normalize the features using the standardized parameters in step S3. 2) Prediction and inference: Input the normalized feature vector into the trained LightGBM classification model and output the probability value of the sequence being B3PPs; 3) Result determination and output: Set the probability threshold to 0.5 (when the predicted probability is ≥0.5, it is determined as B3PPs; otherwise, it is determined as non-B3PPs), and output the predicted category and the corresponding probability value.
[0019] The beneficial effects of this invention are: This invention discloses a method for predicting blood-brain barrier penetration peptides based on multi-feature fusion and data augmentation. First, synthetic B3PPs are generated through FBDiffusion and combined with stratified sampling to construct a balanced training set. Then, 14 features of the peptide sequence are extracted to form a multi-dimensional feature space. Next, CatBoost is used to screen effective features, remove redundancy, and reduce computational complexity. Finally, a classification model is constructed based on LightGBM to accurately predict the blood-brain barrier crossing performance of B3PPs.
[0020] This invention generates bio-rational synthetic B3PPs through FBDiffusion and supplements them with non-B3PPs using stratified sampling, constructing a multi-scale balanced dataset that ensures stable model performance on the independent test set. The model corresponding to the optimal training set (sample size 2592) achieves an AUC of 0.8853 on the independent test set, significantly outperforming existing models such as B3Pred (AUC=0.7286) and DeepB3P (AUC=0.8360). This invention integrates 14 features across four categories: amino acid composition, sequence adjacency, physicochemical properties, and molecular structure, to comprehensively capture the structure-function relationship of B3PPs. By using CatBoost feature filtering to remove redundant features, it compresses 89.2% of the feature dimensions while only losing 0.4% of the AUC, balancing accuracy and efficiency. The GOSS and histogram discretization techniques of the LightGBM model in this invention significantly improve training efficiency compared to traditional random forests, reduce memory consumption, and can be deployed on ordinary computing devices or drug development platforms to meet the needs of large-scale peptide screening. This invention provides key feature contribution values along with the prediction results, helping researchers understand the core mechanism of peptides penetrating the blood-brain barrier and providing guidance for the artificial design and modification of B3PPs.
[0021] In summary, this invention provides a method for predicting blood-brain barrier-penetrating peptides based on multi-feature fusion and data enhancement, enabling rapid and accurate identification of B3PPs and providing an efficient tool for the development of drugs for the treatment of central nervous system diseases. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the overall framework of the method of the present invention; it sequentially shows the process of dataset construction and balancing, multi-feature extraction, CatBoost feature selection, and LightGBM model training and prediction.
[0023] Figure 2 This describes the probability distribution and transformation process of the forward and reverse Markov chains in the FBDiffusion model.
[0024] Figure 3 To optimize the process of generating B3PPs sequences by combining the FBDiffusion model with a BLAST analyzer.
[0025] Figure 4 The flowchart for expanding the training set shows the process of constructing five balanced training sets (train1~train5) of different sizes from the base dataset.
[0026] Figure 5 The graph shows the performance comparison of the models on different training sets, displaying the AUC values of each training set on the validation set and the independent test set.
[0027] Figure 6 shows a comparison of performance and feature quantity for different feature importance thresholds. (a) shows the AUC values of the validation set and test set under different thresholds, and (b) shows the number of features corresponding to different thresholds.
[0028] Figure 7 shows a comparison of the AUC, ACC, SN, SP, and MCC indices of the model of this invention and existing models on an independent test set. (a) is a bar chart of each model, and (b) is a radar chart of each model. Detailed Implementation
[0029] The present invention will now be described in further detail with reference to the accompanying drawings.
[0030] This invention discloses a method for predicting blood-brain barrier-penetrating peptides (B3PPs) based on multi-feature fusion and data augmentation. Treating the B3PP prediction problem as a typical class-imbalanced classification task, a highly efficient prediction model, B3Ppredict, is designed by integrating feature selection and advanced classification algorithms. The B3Ppredict model first uses FBDiffusion to generate synthetic peptides with physicochemical properties highly similar to real B3PPs, and then constructs multiple balanced training sets using stratified sampling. Subsequently, it extracts multi-dimensional features covering amino acid composition, physicochemical properties, sequence order, and molecular structure using 14 methods. Next, the CatBoost algorithm is used to calculate feature importance and select key features with a threshold of 0.0001, reducing the feature dimension from 11,200 to 1,211 with only a 0.4% loss in AUC. Finally, a classification model is constructed based on the LightGBM algorithm. This model employs GOSS sampling, histogram discretization, and Leaf-Wise growth strategies, achieving an AUC of 0.8853 on the independent test set, significantly outperforming existing prediction tools and providing strong support for efficient B3PP screening.
[0031] Please refer to the details. Figure 1 The present invention provides a model-driven method for DOA estimation of complex neural networks, comprising the following steps: S1. In specific implementation, to address the issues of scarce experimental verification samples and natural positive-negative sample imbalance in B3PPs, while ensuring data authenticity and reasonable distribution, and laying the foundation for subsequent model selection of optimal input data, please refer to [the relevant documentation / reference]. Figure 2 Experimentally validated data were collected from the BBPpredict database, published literature, and public databases. Questionable data was manually removed to filter out noise interference, ultimately yielding 324 B3PPs sequences (positive samples) and 324 Non-B3PPs sequences (negative samples) as the basic training set, and 97 B3PPs sequences and 97 Non-B3PPs sequences as the independent test set. The independent test set did not overlap with the training set, ensuring the objectivity of subsequent model evaluation and preventing data leakage that could lead to inflated performance.
[0032] Subsequently, the FBDiffusion model was trained to learn the amino acid composition, physicochemical properties (such as hydrophilicity and hydrophobicity, charge distribution) and sequence length distribution characteristics of real B3PPs, generating synthetic positive samples gen-B3PPs that highly match the structure and functional characteristics of real B3PPs, supplementing the scarce positive samples. At the same time, based on the sequence length distribution of Non-B3PPs in the basic dataset (stratified into intervals of 5-10, 10-20, 20-30, 30-40, and 40-50 amino acid residues), synthetic negative samples were selected by stratified sampling, and synthetic positive and negative samples were gradually added to the basic dataset in a 1:1 ratio, constructing 5 balanced training sets with sample sizes of 648, 1296, 2592, 5184, and 10368, respectively. This process not only yielded high-quality, non-overlapping basic data and independent test sets, solving the problems of sample scarcity and imbalance, but also provided data support for the gradient scale in subsequent selection of the optimal training set, avoiding prediction bias caused by sample bias in the model. It also provided data conditions for verifying that "too many synthetic samples lead to overfitting".
[0033] S2. To comprehensively capture the structural and functional association information of peptide sequences and avoid the problem that a single feature cannot fully characterize the blood-brain barrier penetration mechanism of B3PPs, leading to insufficient model discrimination ability, this step integrates traditional feature descriptors and molecular structural features to construct a high-dimensional feature matrix based on the structural and functional association characteristics of peptide sequences. The traditional feature descriptors cover AAC, APAAC, DPC, TPC, DDE, QSO, SOCN, CKSAAP, ASDC, CTD, SE, SEP, and SER. Combined with molecular structural features that reflect molecular topology, electronic states, and other attributes, a feature matrix of over 11,200 dimensions is finally constructed. This matrix covers key dimensions such as peptide composition, arrangement, terminal characteristics, and molecular structure, maximizing the preservation of the core structural and functional differences between B3PPs and non-B3PPs, avoiding the limitations of a single feature, and laying a solid feature foundation for subsequent model learning of key discrimination information and improvement of prediction accuracy.
[0034] S3. Small sample sizes can lead to insufficient model learning and poor generalization ability, while too many synthetic samples can cause the model to overfit synthetic data details and reduce its adaptability to real data. This step aims to select the optimal training set that balances sample size and generalization ability, ensuring that subsequent experiments have uniform and efficient input data, and improving experimental reproducibility and model generalization ability. Specifically, for the 5 balanced training sets (train1~train5) constructed in S1, as follows... Figure 4 As shown, all features generated by S2 were used, and the model performance of each training set was comprehensively evaluated based on AUC (reflecting the overall discriminative ability of the model), ACC (accuracy), SN (sensitivity, measuring the ability to recognize B3PPs), SP (specificity, measuring the ability to recognize non-B3PPs), and MCC (comprehensive evaluation index) on the independent test set. According to the analysis results in Figure 4, the AUC of training set 3 (sample size 2592) was significantly better than that of other training sets, while training sets 4 and 5 showed obvious overfitting and poor generalization ability due to the addition of too many gen-B3PPs. Therefore, training set 3 was determined as the unified input data for subsequent experiments. This result not only solved the problem of model performance fluctuation caused by inappropriate sample size, but also reduced experimental redundancy by unifying the training set, providing stable input for subsequent feature selection and model training, and improving the prediction reliability of the final model.
[0035] S4. Irrelevant features increase computational complexity and slow down training. Scale differences between features can cause the model to overly favor large-scale features, reducing the fairness of importance calculations. This step uses the CatBoost algorithm for feature selection, balancing prediction accuracy and computational efficiency, and solving the "curse of dimensionality" problem. The implementation process consists of three steps: First, feature standardization is performed. StandardScaler is used to normalize all features to a standard normal distribution with a mean of 0 and a standard deviation of 1, eliminating scale differences and ensuring the fairness of subsequent feature importance calculations. The CatBoost classifier was then trained with 500 iterations, a learning rate of 0.1, a tree depth of 6, and a fixed random seed of 42. The contribution of each feature to the model decision was quantified based on the change in the loss function during the decision tree splitting process, and the feature importance was calculated. Finally, the feature selection threshold was optimized, with a range of 0 to 0.00001. The AUC, ACC, SN, SP, and MCC of the independent test set were used as evaluation metrics to select the optimal threshold. According to Figure 5, when the threshold is 0.0001, the number of features decreased from 11200 to 1211 (compression rate of 89.2%), significantly reducing computational complexity and memory consumption. Furthermore, the test set AUC only decreased by 0.4% (from 0.8857 to 0.8853), with almost no loss of accuracy. This process not only eliminated feature scale differences and avoided misjudgment of importance but also selected key features, achieving a balance between dimensionality reduction and accuracy preservation. This provides lightweight and high-value feature inputs for efficient model training and accurate prediction.
[0036] S5. To leverage LightGBM's efficient optimization strategy to improve model training efficiency and prediction accuracy, and to prevent overfitting through parameter optimization and regularization measures, ensuring the model's robustness and generalization ability, this step constructs and trains a LightGBM classification model.
[0037] First, the model parameters are initialized, fully utilizing LightGBM's gradient one-sided sampling (GOSS), histogram discretization, and leaf-wise growth strategy: GOSS parameters are set (large gradient sample ratio a=0.2, small gradient sample ratio b=0.1) to reduce computational cost; regularization parameter λ=0.01 (controlling model complexity); and new leaf node generation cost γ=0.1 (avoiding redundant splitting). Simultaneously, the learning rate is configured to 0.05, maximum depth 30, maximum number of leaf nodes 64, and number of iterations 1000, balancing model performance and complexity through parameter combinations. Next, using the feature subset filtered by S3 as input, 5-fold cross-validation is employed for model training. 5-fold cross-validation allows the model to learn fully on different data subsets, reducing bias from a single training set. An early stopping strategy is also used: when the validation set AUC shows no improvement after 50 consecutive iterations, training is stopped and the current optimal model parameters are saved, preventing the model from overfitting the training data. Through this process, GOSS, histogram discretization, and Leaf-Wise strategies significantly improved training efficiency and reduced memory consumption, while 5-fold cross-validation and early stopping strategies effectively prevented overfitting. The final trained model has high prediction accuracy and strong generalization ability on independent test sets, providing a reliable model foundation for subsequent prediction applications.
[0038] S6. To facilitate the practical application of the model and establish a standardized prediction process to avoid prediction biases caused by inconsistent processing methods, while simultaneously reducing the time and economic costs of traditional experimental screening through model prediction and accelerating the development of B3PPs-related central nervous system drugs, this step designs a complete B3PPs sequence prediction process. First, the trained LightGBM model is saved. For the B3PPs sequence to be predicted, the corresponding multi-dimensional feature vector is generated strictly according to the feature extraction method in S2, ensuring that the feature dimensions are consistent with those used during model training. Then, the standardized parameters in S3 are used for feature normalization to avoid feature distribution shifts due to differences in normalization parameters. Subsequently, the normalized feature vector is input into the trained model, outputting the probability value that the sequence is a B3PPs sequence, with a probability threshold of 0.5. When the predicted probability is ≥0.5, it is determined to be a B3PPs sequence; otherwise, it is determined to be a non-B3PPs sequence. The predicted category and corresponding probability value are output simultaneously, providing researchers with an intuitive reference. This standardized process ensures the accuracy of the prediction results and avoids processing bias; the probabilistic results and clear classification criteria help researchers quickly determine whether further experimental verification is needed, which greatly reduces the cost of traditional experiments and effectively accelerates the development process of B3PPs.
[0039] like Figure 2 As shown above, the above illustrates a forward Markov chain, starting from the initial real training samples. Initially, through stepwise conditional probabilities ,go through The transformation up to step T ultimately makes the latent variables Approximating the standard normal distribution Their joint distribution is Below is a reverse Markov chain, starting from satisfying... of Starting with learnable conditional probabilities (in (The parameters to be learned) are used to gradually denoise and reconstruct the original sequence. The joint distribution is This enables the generation process from noise to real samples.
[0040] like Figure 3 As shown: First, the diffusion model generates peptide sequences including “DKWNYGL”, “QLHRLHKKGYWSKF”, and “KTCFKKLPE”. Next, these generated sequences are input into a BLAST analyzer, which scores their similarity to real B3PPs. Then, the scoring results and related sequences are used as training data to feed back into the diffusion model, guiding it to generate gen-B3PPs sequences that are even closer to real B3PPs, forming an optimized generation cycle.
[0041] like Figure 5 The chart shows a comparison of AUC for different training sets (training sets 1-5). The left side shows the validation set results, where the AUC increases with the amount of data. The right side shows the independent test set results, where the AUC first increases and then decreases with the amount of data. Excessive data can lead to overfitting. The AUC for training 3 on the independent test set is the highest, indicating the best generalization ability of the model.
[0042] Figure 6 shows a comparison of performance and feature quantity for different feature importance thresholds. (a) shows the AUC values of the validation and test sets under different thresholds, and (b) shows the number of features corresponding to different thresholds. When the threshold is 0.0001, the AUC of the independent test set is the highest, the model has the best generalization ability, and the number of features decreases from 11200 to 1211, which significantly reduces the computational cost.
[0043] Figure 7 shows a comparison of the AUC, ACC, SN, SP, and MCC metrics of the model of this invention and existing models on an independent test set. (a) is a bar chart of each model, and (b) is a radar chart of each model. It can be seen that the AUC and ACC of this invention are significantly higher than those of other models, and the SP accuracy is greatly improved while maintaining the accuracy of SN. In practical use, this effectively avoids misclassifying non-B3PPs as B3PPs, improving the success rate of experimental verification.
Claims
1. A method for predicting blood-brain barrier-penetrating peptides based on multi-feature fusion and data augmentation, characterized in that, Includes the following steps; S1. Collect basic B3PPs and Non-B3PPs sequences, generate pseudo-B3PPs using a feedback diffusion model, supplement Non-B3PPs with stratified sampling, and construct multiple balanced training sets and independent test sets. S2. Combine the features extracted from the peptide sequences in the balanced training set and the independent test set to generate the original feature matrix; S3. For the balanced training set constructed in S1, use all the features in the generated original feature matrix to select the optimal training set based on the model performance on the independent test set. S4. Based on CatBoost feature selection, calculate the feature importance of all features and set the optimal threshold to select a subset of effective features; S5. Build a classification model based on LightGBM, using the optimal training set determined in S3 and the effective feature subset selected in S4, and train and optimize the model using 5-fold cross-validation. S6. Extract and normalize the features of the polypeptide sequence to be predicted, input the optimized model from S5, and output the predicted category and the predicted probability value.
2. The method for predicting blood-brain barrier-penetrating peptides based on multi-feature fusion and data augmentation according to claim 1, characterized in that, The steps for S1 are as follows: Step 1) Collect the basic dataset. Obtain the experimentally verified B3PPs sequences and non-B3PPs sequences from the BBPpredict database as the basic training set. The B3PPs sequences and non-B3PPs sequences are kept equal. We obtained experimentally validated B3PPs sequences and Non-B3PPs sequences as the basic independent test set; we retrieved relevant literature on B3PPs from the PubMed database and supplemented it with thousands of experimentally validated Non-B3PPs sequences; after manual verification, we removed 3 contradictory sequences from the independent test set. The final basic dataset includes B3PPs and Non-B3PPs as the training set, and B3PPs and Non-B3PPs as the independent test set. Step 2) Using the blood-brain barrier penetrating peptide B3PPs sequence in the basic dataset as training data, the FBDiffusion model is trained: the model learns the core biological characteristics of B3PPs, including amino acid composition, sequence length distribution and key physicochemical properties, and finally generates four sets of synthetic B3PPs sequences gen-B3PPs with different numbers of sequences. Step 3) To construct a balanced training set without length bias, firstly, based on the sequence length distribution of non-blood-brain barrier penetrating peptides Non-B3PPs and B3PPs in the basic dataset, a stratified sampling method is used to select the corresponding number of Non-B3PPs samples and the gen-B3PPs samples generated in Step 2); then, by combining positive and negative samples in equal quantities, five balanced training sets are finally constructed.
3. The method for predicting blood-brain barrier-penetrating peptides based on multi-feature fusion and data augmentation according to claim 2, characterized in that, In step 2), the FBDiffusion model is described as follows: The diffusion model represents B3PPs as latent variables. And construct a pair of forward / reverse Markov chains; the forward process is performed by progressively adding Gaussian noise to the real training samples. Perturbation is a latent variable The latent variables satisfy Specifically, the forward process in discrete time steps The upper continuous corrosion sample forms in For the first step of the forward process Step by step, the distribution of latent variables is made to approximate the standard normal distribution. Accordingly, a learnable inverse Markov chain is constructed, and noise is gradually reduced from... Reconstructed sequence Their joint distribution is in The parameters to be learned For the reverse process Step; During training, variational inference is used to maximize the sample likelihood, which is equivalent to maximizing the lower bound of evidence: The FBDiffusion model nests a sequence BLAST analyzer on top of the diffusion model to evaluate the similarity between the samples generated by the diffusion model and real B3PPs, thereby guiding the model to generate peptide sequences that are closer to real B3PPs, i.e., gen-B3PPs.
4. The method for predicting blood-brain barrier penetration peptides based on multi-feature fusion and data augmentation according to claim 3, characterized in that, The features of S2 are amino acid composition, sequence proximity, physicochemical properties and molecular structure features, namely, AAC, APAAC, ASDC, CKSAAP, CTD, DDE, DPC, TPC, SEP, SER, QSO, SE, SOCN and molecular structure features. For each polypeptide sequence in the training and test sets, 14 features are extracted, specifically including: 1) Amino acid composition characteristics: AAC, APAAC, DPC, TPC, and DDE are used to describe the amino acid composition characteristics of the polypeptide sequence; 2) Sequence adjacency characteristics: The adjacency combination patterns of amino acid residues are described using QSO, SOCN, CKSAAP, and ASDC. 3) Physicochemical properties: CTD, SE, SEP, and SER descriptors are used to quantify the hydrophilicity, hydrophobicity, charge, polarity, and spatial structure of amino acids; 4) Molecular structural characteristics: Extract the electronic state, topological complexity, and atomic composition of the polypeptide to reflect the structural basis of the interaction between the polypeptide and the blood-brain barrier endothelial cells.
5. The method for predicting blood-brain barrier-penetrating peptides based on multi-feature fusion and data augmentation according to claim 4, characterized in that, In step S3, model training and performance evaluation are carried out sequentially on the five balanced training sets of different sizes constructed in S1, as follows: 1) For each training set, use the original features extracted in S2 to train the initial classification model; select the LightGBM algorithm as the initial classification model and use the default parameters of the algorithm for training to ensure that the training conditions are consistent on different training sets. 2) Test the model trained on each training set on an independent test set; during the testing process, pay attention to several key model performance metrics, including AUC, accuracy, sensitivity, specificity, and Matthews correlation coefficient. 3) By comparing and analyzing the performance of the model trained on the training set on the independent test set, the training set with the highest AUC index is selected and determined as the basic dataset for subsequent S4 feature selection and S5 model optimization, laying the foundation for obtaining a high-precision blood-brain barrier penetration peptide prediction model.
6. The method for predicting blood-brain barrier-penetrating peptides based on multi-feature fusion and data augmentation according to claim 5, characterized in that, The specific steps of step 1) are as follows: ① Gradient-based One-Sided Sampling is used to filter the training samples, retaining only key samples for subsequent calculations, as shown in the following formula: For the sample The original gradient; The adjusted gradient; For each training set, samples are selected according to this ratio, which not only retains large gradient samples that are key to model optimization, but also reduces the amount of computation through sampling, and the uniform values of a and b ensure that the sampling logic of different training sets is consistent. ② For the original features, LightGBM by default discretizes continuous feature values into histograms, transforming high-dimensional features into low-dimensional histogram features; After discretization, the model selects the optimal feature split point by calculating the split gain, as shown in the following formula: These are the gradients of the left and right child nodes, respectively; These are the Hessian sums of the left and right child nodes, respectively; ; For each set of training features, the split gain of each feature is calculated using this formula. The feature with the largest gain and the split point are selected to construct a decision tree, ensuring training efficiency under high-dimensional features. ③ The Leaf-Wise growth strategy is adopted, and the node with the largest splitting gain is selected from all leaf nodes for growth each time until the default number of iterations or the stopping condition is reached.
7. The method for predicting blood-brain barrier penetration peptides based on multi-feature fusion and data augmentation according to claim 6, characterized in that, In step S4, the CatBoost algorithm is used to filter the original high-dimensional features, remove redundant features, and reduce computational complexity. The steps are as follows: 1) For the generated features, StandardScaler is used to normalize them to a standard normal distribution with a mean of 0 and a standard deviation of 1, thereby eliminating scale interference caused by differences in the representation dimensions of different features; 2) For the standardized features, the importance of each feature is calculated based on the change in the loss function during the decision tree split by training a CatBoost classifier; The calculation formula is as follows: 3) Set the feature importance threshold range to 0~0.00001, and use the AUC, ACC, SN, SP, and MCC of the independent test set as evaluation indicators to select the optimal threshold; 4) Retain features with an importance of not less than 0.0001 to form an optimized feature subset for subsequent model training.
8. The method for predicting blood-brain barrier-penetrating peptides based on multi-feature fusion and data augmentation according to claim 7, characterized in that, S5 is based on the construction and training of the B3PPs classification model using LightGBM; A B3PPs classification model is constructed using the LightGBM algorithm. Its gradient one-sided sampling, histogram discretization, and leaf node-first growth strategy are utilized to improve model training efficiency and prediction accuracy. The steps are as follows: 1) Set the key parameters for LightGBM; 2) Using the feature subset filtered in step S3 as input, the LightGBM model is trained using 5-fold cross-validation. During training, large gradient samples are retained and small gradient samples are randomly sampled using GOSS technology. Gradient weights are adjusted using (Formula 2) to reduce computation. Continuous features are converted into discrete intervals using histogram discretization. First, the number of discrete intervals is preset. Then, the value distribution of each continuous feature in the training set is statistically analyzed. The feature value range is evenly divided into a preset number of intervals according to the value size. Finally, the specific value of each continuous feature is mapped to the corresponding discrete interval to complete the transformation and reduce memory consumption. The leaf node with the largest split gain is selected for growth through the Leaf-Wise growth strategy to accelerate model convergence. 3) An early stopping strategy is adopted to prevent overfitting. When the validation set AUC does not improve for 20 consecutive iterations, training is stopped and the current optimal model parameters are saved.
9. The method for predicting blood-brain barrier penetration peptides based on multi-feature fusion and data augmentation according to claim 8, characterized in that, The steps for predicting and outputting a B3PPs in S6 are as follows: 1) Input sequence preprocessing: For the polypeptide sequence to be predicted, generate the corresponding multi-dimensional feature vector according to the feature extraction method in step S2, and normalize the features using the standardized parameters in step S3. 2) Prediction and inference: Input the normalized feature vector into the trained LightGBM classification model and output the probability value of the sequence being B3PPs; 3) Result determination and output: Set a probability threshold. When the predicted probability is greater than or equal to the probability threshold, it is determined to be B3PPs. Otherwise, it is determined to be a non-B3PPs, and the predicted category and corresponding probability value are output.