A brain-targeting peptide screening method

By introducing various deep learning technologies and molecular docking tools, combined with a multi-level screening mechanism, the problems of high time consumption and low accuracy in brain-targeting peptide screening in traditional methods have been solved, achieving efficient and accurate brain-targeting peptide screening and improving the binding precision of peptides to targets.

CN119541618BActive Publication Date: 2025-12-16BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411597184.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-12-16
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

Existing methods for screening brain-targeting peptides are time-consuming, costly, and prone to experimental errors. Traditional molecular docking methods cannot accurately predict the binding sites and binding forces between proteins and ligands. Artificial intelligence-based methods still have room for improvement in specificity and accuracy, especially in the screening of brain-targeting peptides where complex structures are difficult to identify.

Method used

Using LSTM, trRosettaX, AlphaFold2, Point Transformer deep learning technology and Autodock Vina molecular docking tool, combined with molecular docking binding free energy score and Point Transformer spatial affinity score, a multi-level screening mechanism was used to screen for peptide sequences with high targeting and specificity, avoiding the influence of glycosylation sites.

Benefits of technology

It improves the efficiency and accuracy of brain-targeting peptide screening, shortens the screening cycle, reduces costs, and enhances the binding precision of peptides to targets, achieving rapid screening and efficient identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119541618B_ABST
    Figure CN119541618B_ABST
Patent Text Reader

Abstract

The application discloses a brain-targeting peptide screening method, introduces LSTM, trRosettaX, AlphaFold2, Point Transformer deep learning technology and Autodock Vina molecular docking tool, generates brain-targeting peptide sequences, predicts deep learning peptide structures, docks proteins and peptides and predicts spatial affinity, increases biochemical characteristic information of target points and peptides through Point Transformer deep learning screening, applies double structures of molecular docking combined free energy score and Point Transformer spatial affinity score for screening, and realizes brain-targeting peptide screening. The method can accelerate the screening process of brain-targeting peptides, shorten the experimental period, reduce the screening cost, realize rapid screening, improve the binding accuracy and screening accuracy of peptides and target points, and effectively screen out candidate peptides with low affinity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of targeted peptide screening technology, and more specifically to a method for screening brain-targeted peptides. Background Technology

[0002] Traditional brain-targeting peptide screening techniques primarily rely on animal and cell experiments to identify targeting peptides capable of crossing the blood-brain barrier. However, this method is time-consuming, costly, and prone to experimental errors and ethical issues. In practical applications, peptide screening that crosses the blood-brain barrier remains a significant technical challenge.

[0003] While existing computer-based screening methods can accelerate the molecular screening process to some extent through affinity prediction, they still face many limitations in the screening of brain-targeting peptides. Traditional molecular docking methods (such as Autodock Vina) are limited by their insufficient ability to simulate molecular flexibility, failing to effectively predict protein-ligand binding sites and binding forces, resulting in insufficient accuracy in target identification. While AI-based virtual screening offers advantages in efficiency and automation, there is still room for improvement in specificity and accuracy. Particularly in the screening of brain-targeting peptides, the limited diversity and coverage of training data restricts the specificity of AI models in identifying targets, making it difficult to reliably identify complex structures. Furthermore, existing deep learning models struggle to accurately capture complex intermolecular interactions in molecular structure prediction, failing to efficiently reconstruct protein-ligand complex binding patterns, thus limiting the identification and optimization of high-affinity peptides. These technologies have not yet achieved ideal results in the application of brain-targeting peptide screening. Therefore, further optimizing the method system integrating molecular modeling and deep learning to improve the prediction accuracy of target protein binding sites and enhance the accuracy of high-throughput screening is an important direction for advancing research on brain-targeting peptides. Summary of the Invention

[0004] In view of the above-mentioned shortcomings in the prior art, the purpose of this invention is to provide a method for screening brain-targeting peptides, so as to improve the efficiency and accuracy of brain-targeting peptide screening.

[0005] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:

[0006] A method for screening brain-targeting peptides is provided, comprising the following steps:

[0007] By introducing LSTM, trRosettaX, AlphaFold2, Point Transformer deep learning technologies and the Autodock Vina molecular docking tool, brain-targeting peptide sequences are generated, peptide structures are predicted through deep learning, proteins and peptides are docked, and spatial affinity is predicted. Point Transformer deep learning is used to screen and increase the biochemical characteristics of targets and peptides. A dual structure screening method combining molecular docking with free energy scoring and Point Transformer spatial affinity scoring is applied to achieve brain-targeting peptide screening.

[0008] Furthermore, the method for generating brain-targeting peptide sequences using p_LSTM specifically includes the following sub-steps:

[0009] (1) Train the pep_LSTM model on the active peptide database based on the LSTM network to enable it to generate a variety of peptide sequences;

[0010] (2) Fine-tuning the data with peptides related to the blood-brain barrier to make the model more accurate in its predictions;

[0011] (3) Input the target CD98hc sequence information into the model to generate a virtual peptide sequence with potential brain targeting;

[0012] (4) The target CD98hc and peptide were docked using docking technology to screen out relatively accurate candidate peptides.

[0013] Furthermore, the structural prediction method specifically includes the following sub-steps:

[0014] (1) The peptide sequences generated by the pep_LSTM model are used to predict the three-dimensional structure PDB or PDBQT using trRosettaX and Alpha Fold2 techniques.

[0015] (2) The folding conformation of peptides is predicted by deep learning models and co-evolutionary information, and the sequence is converted into PDB or PDBQT file format to provide a basis for subsequent peptide docking.

[0016] Furthermore, the Autodock Vina molecular docking method specifically includes the following sub-steps:

[0017] (1) Using Vina, the processed peptide structure was molecularly docked with the target CD98hc, and the binding free energy was calculated.

[0018] (2) Vina selected peptide sequences with lower binding energy through energy scoring as candidate peptides with potential targeting and blood-brain barrier crossing capabilities.

[0019] Furthermore, the Point Transformer deep learning filtering method specifically includes the following sub-steps:

[0020] (1) Obtain the protein-ligand interaction dataset from the RCSB PDB and BrainPep databases and preprocess the dataset;

[0021] (2) Divide the dataset according to the ratio of training set: test set = 8:2;

[0022] (3) By training and optimizing the Point Transformer model with the training set, the biochemical characteristics of the target and peptide, such as hydrophobicity, polarity, and charge distribution, are increased to make it more suitable for the needs of brain-targeted peptide screening.

[0023] (4) Input the target data and peptide dataset into the model to predict affinity;

[0024] (5) Point Transformer evaluates peptide-target interactions by analyzing three-dimensional spatial affinity, hydrophobicity, polarity and charge distribution, and screens out peptide sequences with higher targeting and specificity.

[0025] Furthermore, after screening using a dual structure of molecular docking binding free energy score and Point Transformer spatial affinity score, the following steps are also included: predicting peptide toxicity of candidate peptides using Ad met and Toxi n Pred, and deleting candidate peptides predicted to be toxic.

[0026] Furthermore, it also includes: labeling the glycosylation sites of the target CD98hc. N-glycosylation is found using the formula Asn-X-Ser / Thr (X≠Pro), and O-glycosylation is found using the formula Ser / Thr. In addition, there are glycosylation sites reported in the literature. After all glycosylation sites are labeled and recorded, regular expressions are used to filter whether the pocket docking results contain active sites. If they do, they are deleted.

[0027] The beneficial effects of this invention are as follows:

[0028] 1. Improved Point Transformer algorithm: In addition to the three-dimensional spatial location of the algorithm itself, biochemical characteristics of the target and peptides have been added, such as hydrophobicity, polarity, and charge distribution. This allows Point Transformer to not only rely on geometric information but also consider the biochemical properties between molecules, thereby improving the accuracy of screening.

[0029] 2. Multi-level affinity screening mechanism: This invention is the first to apply a dual structure screening mechanism of molecular docking combined with free energy score and PointTransformer spatial affinity score, which effectively improves the specificity and screening accuracy of brain-targeting peptides against the target CD98hc.

[0030] 3. Enhanced binding effect: The presence of glycosylation sites can lead to steric hindrance and increased interference from charge distribution and polarity, affecting the binding of the target and peptide. This invention uses algorithms to avoid the influence of glycosylation sites on the binding force of the target and peptide, making it easier to bind. Attached Figure Description

[0031] Figure 1 This is a flowchart of the prediction process for the generation of virtual peptides in the blood-brain barrier in step one.

[0032] Figure 2 The process and scoring of peptide structure transformation in step two;

[0033] Figure 3 This is a flowchart of the Point Transformer prediction and filtering process in step five.

[0034] Figure 4 The confusion matrix (left) and ROC curve (right) for the PointTransformer virtual filtering in step five.

[0035] Figure 5 This is an improvement on the Point Transformer algorithm in step five;

[0036] Figure 6 The flowchart for the Autodock Vina connection filtering in step four;

[0037] Figure 7 This is a flowchart illustrating the combined screening process using the two virtual screening methods in step six.

[0038] Figure 8 The active site (left) and glycosylation site (right) of the target in step six;

[0039] Figure 9 This is the structure of the peptide N-terminus linked to FITC in step seven;

[0040] Figure 10 is The results of cell transphagocytosis efficiency in step seven;

[0041] Figure 11 The data represents the fluorescence data of the mouse brain in step seven. Detailed Implementation

[0042] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0043] This invention proposes an AI-powered brain-targeting peptide screening method, incorporating deep learning technologies such as LSTM, trRosettaX, AlphaFold2, and Point Transformer, along with the Autodock Vina molecular docking tool. This enables functions such as related peptide generation models, deep learning peptide structure prediction, protein-peptide docking, and spatial affinity prediction, constructing a comprehensive intelligent screening system. The specific steps are as follows:

[0044] Step 1: Generate the target peptide sequence using p-LSTM

[0045] (1) The pep_LSTM model was trained on a large number of active peptide databases based on the LSTM network, so that it could generate a variety of peptide sequences to improve the diversity and coverage of the initial screening.

[0046] (2) Fine-tuning the data with peptides related to the blood-brain barrier to make the model more accurate in its predictions;

[0047] (3) Input the target CD98hc sequence information into the model to generate a virtual peptide sequence with potential brain targeting;

[0048] (4) The target CD98hc and peptide were docked using docking technology to screen out relatively accurate candidate peptides.

[0049] Step 2: Structural Prediction

[0050] (1) The peptide sequences generated by the pep_LSTM model were predicted by three-dimensional structure PDB or PDBQT using trRosettaX and Alpha Fold2 technology. Sequence lengths greater than 20 were predicted by AlphaFold2 technology, and sequence lengths between 0 and 20 were predicted by trRosettaX technology.

[0051] (2) The peptide's folding conformation is predicted using a deep learning model and co-evolutionary information. The sequence is then converted into a PDB or PDBQT file format, and a score image of the different residues is generated (the specific prediction results are scored in...). Figure 2 The bottom result provides a foundation for subsequent peptide docking.

[0052] Step 3: Open Babel processing

[0053] The target CD98hc and the generated peptide's PDB or PDBQT structure were further processed using Open Babel to add hydrogen atoms and electrons to ensure the integrity of the molecular structure. This processing step ensures that the peptide has reasonable atomic characteristics in the subsequent molecular docking process, improving the reliability of the docking results.

[0054] Step 4: Autodock Vina Molecular Docking

[0055] (1) Using Vina, the processed peptide structure was molecularly docked with the target CD98hc, and the binding free energy was calculated.

[0056] (2) Vina selected peptide sequences with lower binding energy through energy scoring as candidate peptides with potential targeting and blood-brain barrier crossing capabilities.

[0057] Step 5: Point Transformer Deep Learning Filtering

[0058] (1) Obtain the protein-ligand interaction dataset from the RCSB PDB and BrainPep databases and preprocess the dataset;

[0059] (2) Divide the dataset according to the result that the ratio of training set to test set is 8:2;

[0060] (3) By training and optimizing the Point Transformer model with the training set, the biochemical characteristics of the target and peptide, such as hydrophobicity, polarity, and charge distribution, are increased to make it more suitable for the needs of brain-targeted peptide screening.

[0061] (4) Input the target data and peptide dataset into the model to predict affinity;

[0062] (5) Point Transformer evaluates peptide-target interactions by analyzing three-dimensional spatial affinity, hydrophobicity, polarity and charge distribution, and screens out peptide sequences with higher targeting and specificity.

[0063] like Figure 4 As shown on the left: Confusion matrix analysis, which provides a deeper understanding of the model's performance across different categories. The confusion matrix data on the test set is shown below:

[0064] Top left: Low affinity predicted as low affinity (true negative class): 145;

[0065] Top right: Low affinity predicted as high affinity (false negative): 56v

[0066] Bottom left: High affinity predicted as low affinity (false positive class): 39;

[0067] Bottom right: High affinity predicted as high affinity (true class): 160.

[0068] Precision and recall can be calculated using the confusion matrix.

[0069] Accuracy: Of all the pairs predicted as high affinity, the model correctly identified 160 true class samples, therefore the accuracy is:

[0070]

[0071] Recall: Among all samples that are actually of high affinity, the model successfully identified 160 true class samples, therefore the recall rate is:

[0072]

[0073] On the right: AUC (Area Under Curve) is the core metric for measuring the classification performance of a model. Its value ranges from 0.5 to 1, and a higher value indicates better classification performance.

[0074] Step Six: Final Screening

[0075] (1) Combine the molecular docking screening results of Autodock Vina with the deep learning screening results of Point Transformer, and sort the binding energy and affinity scores in a comprehensive manner.

[0076] (2) Peptide toxicity prediction of candidate peptides was performed using Ad met and Toxi n Pred, and candidate peptides predicted to be toxic were deleted.

[0077] (3) The glycosylation sites of the target CD98hc were labeled. N-glycosylation was found by the formula Asn-X-Ser / Thr (X≠Pro), and O-glycosylation was found by the formula Ser / Thr. In addition, there are glycosylation sites reported in the literature. After all the glycosylation sites were labeled and recorded, regular expressions were used to filter whether the pocket docking results contained active sites. If they did, they were deleted.

[0078] (4) Finally, the optimal peptide sequence targeting the blood-brain barrier was selected.

[0079] Step Seven:

[0080] (1) The selected candidate peptides were labeled with FITC and synthesized;

[0081] (2) The targeting selection and efficiency of candidate peptides across the blood-brain barrier were verified by constructing an in vitro Transwell model;

[0082] (3) Verify whether the target peptide can cross the blood-brain barrier by injecting it into the tail vein of mice.

[0083] The verification results are as follows Figure 9-11 As shown. In summary, the method of this invention can accelerate the screening process of brain-targeting peptides, shorten the experimental cycle, reduce screening costs, achieve rapid screening, improve the binding precision of peptides to targets and the accuracy of screening, and effectively screen out candidate peptides with low affinity. The screening process of this invention has good versatility and can be extended to the screening needs of other targeted peptides, providing new ideas for the development of targeted peptides for various diseases.

[0084] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of equivalents of the claims be included within the present invention.

[0085] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A method for screening brain-targeting peptides, characterized in that, Includes the following steps: By introducing LSTM, trRosettaX, AlphaFold2, Point Transformer deep learning technologies and AutodockVina molecular docking tool, brain-targeting peptide sequences are generated, peptide structures are predicted through deep learning, proteins and peptides are docked, and spatial affinity is predicted. The biochemical characteristics of targets and peptides are increased through Point Transformer deep learning screening. The dual structure screening is performed by combining molecular docking with free energy scoring and Point Transformer spatial affinity scoring to achieve brain-targeting peptide screening. The method for generating brain-targeting peptide sequences using pep_LSTM specifically includes the following sub-steps: (1) Train the pep_LSTM model on the active peptide database based on the LSTM network to enable it to generate a variety of peptide sequences; (2) Fine-tuning the data with peptides related to the blood-brain barrier to make the model more accurate in its predictions; (3) Input the target CD98hc sequence information into the model to generate a virtual peptide sequence with potential brain targeting; (4) The target CD98hc and peptide were docked using docking technology to screen out relatively accurate candidate peptides. The structural prediction method specifically includes the following sub-steps: (1) The peptide sequences generated by the pep_LSTM model are used to perform three-dimensional structure PDB or PDBQT prediction using trRosettaX and AlphaFold2 techniques. (2) Predict the folding conformation of peptides using deep learning models and co-evolutionary information, and convert the sequences into PDB or PDBQT file formats to provide a basis for subsequent peptide docking. The Autodock Vina molecular docking method specifically includes the following sub-steps: (1) Using Vina, the processed peptide structure was molecularly docked with the target CD98hc, and the binding free energy was calculated. (2) Vina selected peptide sequences with lower binding energy through energy scoring as candidate peptides with potential targeting and blood-brain barrier crossing capabilities. The Point Transformer deep learning filtering method specifically includes the following sub-steps: (1) Obtain the protein-ligand interaction dataset from the RCSB PDB and BrainPep databases and preprocess the dataset; (2) Divide the dataset according to the result that the ratio of training set:test set = 8:2; (3) By training and optimizing the Point Transformer model with the training set, the biochemical characteristics of the target and peptide, such as hydrophobicity, polarity, and charge distribution, are increased to make it more suitable for the needs of brain-targeted peptide screening. (4) Input the target data and peptide dataset into the model to predict affinity; (5) Point Transformer evaluates peptide-target interactions by analyzing three-dimensional spatial affinity, hydrophobicity, polarity and charge distribution, and screens out peptide sequences with higher targeting and specificity.

2. The brain-targeting peptide screening method according to claim 1, characterized in that, After screening using a dual structure of molecular docking binding free energy score and Point Transformer spatial affinity score, the following steps are also included: predicting peptide toxicity of candidate peptides using Admet and ToxinPred, and deleting candidate peptides predicted to be toxic.

3. The brain-targeting peptide screening method according to claim 2, characterized in that, Also includes: The glycosylation sites of the target CD98hc were labeled. N-glycosylation was found using the formula Asn-X-Ser / Thr (X≠Pro), and O-glycosylation was found using the formula Ser / Thr. In addition, there were other glycosylation sites reported in the literature. After all the glycosylation sites were labeled and recorded, regular expressions were used to filter whether the pocket docking results contained active sites. If they did, they were deleted.

Citation Information

Patent Citations

  • Screening method of polypeptide compound and related device

    CN114724643A

  • Method for constructing target tumor polypeptide screening model by using structural biology

    CN115966248A