Screening method of LRP1 cross-blood lost barrier targeting peptide based on artificial intelligence
By using an artificial intelligence-based method to screen for high-affinity LRP1-targeting peptides, the problems of long screening cycles and poor stability in existing technologies have been solved, enabling precise drug delivery across the BLB and making it suitable for the treatment of inner ear diseases.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN SECOND PEOPLES HOSPITAL (SHENZHEN INST OF TRANSLATIONAL MEDICINE)
- Filing Date
- 2026-01-13
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies have difficulty effectively delivering drugs to the cochlear tissue via systemic administration, which limits the therapeutic effects of inner ear diseases such as sensorineural hearing loss. Furthermore, existing LRP1-targeting peptide screening methods suffer from problems such as long screening cycles, weak binding energy, poor structural stability, and insufficient reproducibility, making it difficult to achieve precise cross-BLB drug delivery.
Using an artificial intelligence-based approach, the peptide backbone was generated using RFdiffusion, sequence design was performed using ProteinMPNN, structure prediction was performed using AlphaFold3, and energy verification was performed using PRODIGY. High-affinity LRP1 targeting peptides were screened out, and their structural stability was verified by molecular dynamics simulations. The peptides were then applied to an exosome delivery system.
We have developed an LRP1 targeting peptide with high specificity, high affinity and good structural stability. It can serve as a molecular recognition element for cross-BLB drug delivery, and has versatility and scalability, making it suitable for other cross-barrier drug delivery systems.
Smart Images

Figure CN122050484A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and biomedicine, and in particular to an artificial intelligence-based method for screening LRP1-targeting peptides across the blood labyrinth barrier. Background Technology
[0002] The blood-labyrinth barrier (BLB) is an important structure for maintaining inner ear homeostasis. It is composed of capillary endothelial cells, tight junctions, basement membrane, and pericytes, and has high selective permeability to most macromolecular and small molecule drugs.
[0003] Due to the presence of BLB, systemic drug delivery makes it difficult for drugs to effectively enter the cochlear tissue, thus severely limiting the treatment efficacy of inner ear diseases such as sensorineural hearing loss (SNHL).
[0004] Low-density lipoprotein receptor-related protein 1 (LRP1) is a multifunctional transmembrane receptor that is widely distributed in brain blood vessels, inner ear endothelium, and neuronal membranes. It can mediate the cross-barrier transport of various ligands and is an important target for achieving precise cross-BLB drug delivery.
[0005] Existing LRP1-targeting peptides are mainly obtained through phage display, random peptide library screening, or natural protein fragment analysis, such as LDL receptor-related protein 1 (LRP1), a novel target for opening the blood-labyrinth barrier (BLB) (Signal Transduction and Targeted Therapy, 2022). However, these traditional methods generally suffer from problems such as long screening cycles, weak binding energies, poor structural stability, and insufficient reproducibility, making it difficult to meet the needs of precision drug delivery.
[0006] With the development of AI-driven protein structure prediction models (such as AlphaFold2, RFdiffusion, and ProteinMPNN), AI-driven structure-guided short peptide design has become possible. However, most existing models do not perform spatial constraints or binding energy assessments for specific receptor targets, making it difficult to achieve efficient screening and energy optimization for BLB-specific targets. Examples include: DeepB3P: A transformer-based model for identifying blood-brain barrier penetrating peptides with data augmentation using feedback GAN (2024) and ESM-BBB-Pred: a fine-tuned ESM 2.0 and deep neural networks for the identification of blood–brain barrier peptides (Briefings in Bioinformatics, 2025).
[0007] Therefore, there is an urgent need for a systematic strategy that integrates AI structure generation, sequence optimization and energy verification to achieve high-affinity de novo design and intelligent screening of LRP1 targeting peptides, providing a novel computational solution for cross-BLB precision drug delivery. Summary of the Invention
[0008] The purpose of this invention is to provide an artificial intelligence-based method for screening LRP1 target peptides across the BLB, thereby solving the aforementioned technical problems.
[0009] To address the aforementioned technical problems, this invention provides an artificial intelligence-based method for screening LRP1 transBLB target peptides, implemented as follows:
[0010] A method for screening LRP1 cross-BLB target peptides based on artificial intelligence includes the following steps: Target determination: The region of residues 5 to 117 in the crystal structure of LRP1 (low-density lipoprotein receptor-associated protein 1) is used as the target binding site; Backbone generation: Based on the spatial configuration of the target binding site, a peptide backbone of 10-15 amino acids is generated using the diffusion generation model RFdiffusion; Sequence design: The amino acid sequence of the generated peptide backbone is designed and optimized using the message passing neural network ProteinMPNN to obtain candidate peptide sequences; Structure prediction and preliminary screening: The complex structure of the candidate peptide sequence and LRP1 is predicted using AlphaFold3, and the candidate peptides are preliminarily screened based on the predicted template modeling score pTM, interface template modeling score ipTM, and minimum prediction alignment error min PAE. A comprehensive scoring function S is constructed to rank the preliminarily screened candidate peptides; Energy verification: The predicted binding free energy ΔG and dissociation constant Kd of the candidate peptides with LRP1 after preliminary screening are calculated using the PRODIGY model, and peptides with ΔG < -6 kcal / mol and Kd ≤ 1×10⁻ 7 M's high-affinity LRP1-targeting peptide.
[0011] Optionally, in the structure prediction and initial screening, a comprehensive scoring function S is constructed to rank the screened candidate peptides. The function is: S = a1 × ipTM + a2 × pTM - a3 × Disorder - a4 × Clash, where Disorder is the proportion of structural disorder, Clash is the atomic conflict index, and a1 to a4 are adjustable weight coefficients. The top 5% of candidate sequences are selected based on the S value to enter the energy verification.
[0012] Optionally, in the comprehensive scoring function, a1=0.8, a2=0.2, a3=0.5, a4=100.
[0013] Optionally, the preliminary screening criteria in the structure prediction and initial screening are: pTM > 0.8 and min PAE < 1.5 Å.
[0014] Optionally, the method further includes the step of performing molecular dynamics simulations on the screened high-affinity LRP1-targeting peptide to verify the structural stability of the complex formed with LRP1 under physiological conditions.
[0015] Optionally, the molecular dynamics simulation duration is 100 nanoseconds. When the root mean square deviation (RMSD) of the complex interface is stable within 2 Å and the hydrogen bond retention rate is greater than 80%, the structural stability of the targeting peptide is determined to meet the requirements.
[0016] A high-affinity LRP1-targeting peptide obtained by screening using any of the methods described above.
[0017] The present invention also provides a drug delivery system comprising the high-affinity LRP1 targeting peptide and a delivery carrier; wherein the high-affinity LRP1 targeting peptide is modified on the surface of the delivery carrier.
[0018] Optionally, the delivery vector is an exosome.
[0019] The present invention also provides the application of a drug delivery system in the preparation of a drug for treating inner ear diseases.
[0020] The artificial intelligence-based screening method for LRP1 cross-BLB target peptides provided by this invention differs from existing technologies DeepB3P and ESM-BBB-Pred in the following ways: (1) Core technical indicators: This invention is designed and functionally screened from scratch, targeting specific LRP1 cross-BLB target peptides; while the core technology of existing technologies is only classification and identification, judging whether a given peptide sequence can penetrate the BLB. (2) The essence of the problem in this invention is generation and optimization: the input is the target structure, and the output is a novel peptide sequence and structure that binds to the target with high affinity; while the input of existing technologies is the peptide sequence, and the output is the probability and score of penetrating the BLB. (3) AI model role, the core engine of this serial workflow: 1. RFdiffusion: structure generator, 2. ProteinMPNN: sequence designer, 3. AlphaFold3: structure validator, 4. PRODIGY / MD: energy estimator; existing technologies are: independent end-to-end predictors: DeepB3P: Transformer classifier; ESM-BBB-Pred: fine-tuning model based on protein language model (ESM). (4) Data dependence, this invention relies on the three-dimensional structure of the target and the physical force field, and has a low dependence on known peptide sequence data, belonging to "physical / structure-based generation"; existing technologies are highly dependent on labeled BLB penetrating peptide and non-penetrating peptide datasets, and the model performance is limited by data quality and scale, belonging to "data-based discrimination". (5) This invention outputs a novel peptide sequence and its complex three-dimensional structure with the target, and performs physical interpretation-based verification by combining free energy (ΔG), dissociation constant (Kd) and molecular dynamics simulation; existing technologies: output a penetration probability score. It is impossible to provide specific binding targets, affinity or three-dimensional mechanism of action, and the interpretability is relatively weak. (6) Application scenario: Precise targeted delivery of the present invention: Design "keys" for specific receptors (such as LRP1) to achieve active and efficient cross-barrier drug delivery; Existing technology: Preliminary screening: Rapidly screen molecules that may penetrate the barrier from a large number of candidate molecules, but the penetration mechanism and targeting are unknown.
[0021] The present invention has the following beneficial effects:
[0022] This invention achieves fully automated design from target modeling to affinity verification through an AI-driven structure-energy integrated design process. The resulting target peptide has high specificity, high affinity and good structural stability, and can be used as a molecular recognition element for cross-BLB drug delivery.
[0023] This invention is versatile and scalable, and can be applied to other cross-barrier drug delivery systems or membrane receptor target screening, thus having broad application value.
[0024] The main innovations of this invention include: a target-constrained generation mechanism: for the first time, LRP1 binding pocket (residues 5–117) spatial constraints are introduced into the RFdiffusion model to achieve target-guided peptide backbone generation; an intelligent sequence optimization algorithm: the ProteinMPNN message passing network is used to replace random sequence generation, significantly improving sequence foldability and energy matching; a two-layer energy verification system: combining AlphaFold3 structure prediction and PRODIGY energy assessment, a comprehensive judgment from static structure to thermodynamic stability is achieved; and a unified and interpretable scoring system: an S-value function is constructed to achieve unified quantitative ranking of results from multiple models, improving the transparency and interpretability of the AI design process. Attached Figure Description
[0025] Figure 1 This is the overall flowchart of the LRP1 target peptide screening based on artificial intelligence in this invention;
[0026] Figure 2 This is a flowchart of the structure-guided sequence optimization based on ProteinMPNN in this invention;
[0027] Figure 3 This is a three-dimensional structural model of the LRP1 bonding region of the present invention;
[0028] Figure 4 This is a schematic diagram of a representative short peptide backbone three-dimensional structure generated by artificial intelligence in this invention;
[0029] Figure 5 This is a diagram showing the interface structure of a representative short peptide and LRP1 generated by artificial intelligence in this invention.
[0030] Figure 6 This is a scatter plot of short peptide binding energy prediction generated by artificial intelligence in this invention;
[0031] Figure 7 This is a statistical diagram of short peptide binding energy and interfacial contact distribution generated by artificial intelligence in this invention;
[0032] Figure 8This is a graph showing the energy screening results and predicted dissociation constant (Kd) distribution of the composite structure generated by artificial intelligence in this invention;
[0033] Figure 9 This is a schematic diagram illustrating the application of LRP1 target peptide-modified exosomes for transBLB delivery provided by the present invention. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of this invention clearer, the following embodiments provide a more detailed description of the invention. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of the invention.
[0035] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0036] This invention is based on artificial intelligence and is performed on an Ubuntu 22.04 operating system with an NVIDIA A100 GPU environment. PyTorch, RFdiffusion, ProteinMPNN, and AlphaFold3 modules are installed for AI structure generation and energy assessment. All calculations are performed in the CUDA 12.0 environment.
[0037] This invention relates to an artificial intelligence-based method for screening LRP1 transBLB-targeting peptides, comprising the following steps, the flowchart of which is shown below. Figure 1 As shown.
[0038] S101. Binding Target Structure Preparation and Selection of Binding Proteins (Binding Target)
[0039] This invention targets low-density lipoprotein receptor-related protein 1 (LRP1), whose structure is derived from the crystal structure numbered 7D4I in the Protein Database (PDB). This structure comprises 76 protein chains and 3 nucleic acid chains, with a total molecular weight of approximately 5,300.1 kDa. The corresponding sequence of LRP1 in this structure is as follows:
[0040] >7D4I_R7MEDIEKIKPYVRSFSKALDELKPEIEKLTSKSLDEQLLLLSDERAKLELINRYAYVLSSLMFANMKVLGVKDMSPILGELKRVKSYMDKAKQYDNRITKSNEKSQAEQEKAKNIISNVLDGNKNQFEPSISRSNFQGKHTKFENDELAESTTTKIIDSTDHIRKASSKKSKRLDKVGKKKGGKK
[0041] Among them, amino acids 5 to 117 in the sequence are completely resolved and have a clear conformation in the crystal structure. Therefore, this study selects this region as the binding region and uses it as the target region for subsequent protein design.
[0042] Then, energy minimization and 100 ns equilibrium molecular dynamics simulations (CHARMM36m force field, 300 K, 1 bar) were performed using GROMACS software to obtain stable conformations as target inputs for later modeling.
[0043] In the S102 AI generation stage, the peptide backbone is generated by combining the peptidation framework generation (RF diffusion) and the RF diffusion (RFdiffusion) model.
[0044] Based on the identification of the active site of the bait protein in step S101, the protein structure generation algorithm RFdiffusion based on the diffusion model is used to design the structural backbone of the binding protein for this region (https: / / github.com / RosettaCommons / RFdiffusion).
[0045] Under target-constrained conditions, 1000 binding peptide backbones are generated. The preferred diffusion step size is 500–1500, the preferred noise intensity is 0.01–0.05, and the preferred binding pocket radius is 6–10 Å. In one specific embodiment, the diffusion step size is 1000, the noise intensity is 0.02, and the binding pocket radius is 8 Å. This ensures that the generated peptide matches the target surface. The binding agent length is set to 10–15 amino acids, and the Cα–Cβ distance is limited to ≤ 6 Å to ensure geometric rationality and conformational feasibility, match the spatial and chemical characteristics of interfacial interactions, and improve binding affinity and specificity.
[0046] S103, Binder Design Sequence Optimization Design
[0047] The ProteinMPNN model was used for sequence prediction and optimization of the generated backbone. Temperature parameters were set between 0.1 and 0.3, and five candidate sequences were generated for each backbone, resulting in approximately 5000 short peptides. The top 20% of candidate sequences were selected based on scores for hydrophobicity, charge distribution, and spatial compensation. This selection was based on a comprehensive evaluation of amino acid hydrophobicity distribution, charge complementarity, and spatial geometric matching, without limiting the specific scoring function or weighting form.
[0048] like Figure 2 The diagram shows a flowchart of structure-guided sequence optimization based on ProteinMPNN. The model takes the protein's 3D structure and evolutionary information as input and uses a deep learning algorithm to generate new sequences compatible with the target structure. Experimental results show that the designed protein exhibits significant improvements in physicochemical properties (such as better folding and stability), demonstrating the effectiveness of structure-driven sequence optimization in protein engineering.
[0049] Figure 2 The blue skeleton on the left represents the structural input, and the middle section shows a schematic of the neural network framework; the right side shows the gel filtration chromatography (SEC) curve: the horizontal axis represents the retention volume (mL), and the vertical axis represents the absorbance (Abs); the blue curve (Design) represents the AI-optimized sequence, and the gray curve (Parent) represents the original sequence.
[0050] The following is Figure 2 Explanation:
[0051] 1. INPUT — Model Input
[0052] The left side displays the input information for the design process, including two types of data:
[0053] Structure: Represents the three-dimensional spatial structure of the target protein (blue skeleton diagram, green represents active sites or ligand-binding residues). This structural information provides spatial constraints for the model, enabling it to understand the geometric relationships and interactions between amino acids.
[0054] Evolution (Evolutionary Information): Displays multiple sequence alignments (MSA). These sequences provide a reference for conservation and variability during natural evolution, enabling the model to balance structural stability and functional feasibility when generating new sequences.
[0055] 2. OUTPUT — Model Output
[0056] The middle section illustrates the role of the ProteinMPNN model. ProteinMPNN is a sequence design model based on a graph neural network (GNN). Given a protein's three-dimensional structure, it predicts amino acid sequences that are compatible with that structure and may possess superior properties. The output section lists several "Designed sequences" generated by the model, such as:
[0057] …LFNTDDGHSVS…
[0058] …LTNTSDGHSIS…
[0059] These are candidate sequences generated by the model based on structural constraints and statistical laws, used to replace or optimize the original (parent) protein.
[0060] 3. Improved Properties — Properties of Improvement
[0061] The right side shows the performance improvement of the redesigned sequence in experimental validation:
[0062] The blue curve (Design) in the figure represents the protein after model design;
[0063] The gray curve (Parent) represents the original natural sequence.
[0064] The vertical axis represents absorbance, and the horizontal axis represents retention volume (mL). This is typical gel filtration chromatography (SEC, Size Exclusion Chromatography) data, used to reflect the folding state and aggregation behavior of proteins.
[0065] The results show:
[0066] The designed protein peaks are more concentrated and the signal is stronger, indicating that it is more uniform in structure and more stable in folding.
[0067] The original protein peaks were weak and dispersed, suggesting partial misfolding or aggregation.
[0068] Therefore, it can be inferred that sequences designed using ProteinMPNN achieve better stability or expressive performance while maintaining their original functionality.
[0069] S104, LRP1-targeting peptide output, LRP1-Pep series candidate peptides, LRP1-peptide complex structure prediction.
[0070] The structure of LRP1 complexes with candidate peptides was predicted using AlphaFold3, an advanced model that accurately predicts the structures of complexes of almost all molecular types in a protein database. Batch complex structure predictions were performed on 1000 designed binding agent sequences with residues 5–117 of the LRP1 protein, and key metrics such as prediction template modeling score (pTM), interface prediction template modeling score (ipTM), and minimum pairing error (min PAE interaction) were systematically calculated.
[0071] like Figure 3 The diagram shows the three-dimensional structure of the LRP1 protein of this invention bound to the 444th of 1000 designed binding peptides. GLU-16, the 16th glutamic acid residue on the peptide chain, carries a negative charge and forms a salt bridge with a positively charged residue on LRP1 through electrostatic attraction, resulting in the strongest interaction. TYP-6, located at the 6th position of the binding peptide, contains an aromatic ring and participates in π-π stacking or hydrophobic interactions, playing a crucial role in stabilizing the binding interface. ARG-8, arginine, located at the 8th position of the peptide chain, is a positively charged basic amino acid that forms a salt bridge or hydrogen bond with negatively charged residues on LRP1 (such as GLU or ASP). LEU-12, the 12th leucine residue on the binding peptide chain, primarily contributes to hydrophobic interactions, helping the peptide chain anchor in the hydrophobic pocket of LRP1.
[0072] This invention also assesses the reliability of the predicted structure: the align command in PYMOL software is used to compare the batch composite structure predicted by AlphaFold3 with the crystal structure of LRP1 (PDB number 7D4I). The root mean square deviation (RMSD) value is calculated to be 1.104 Å, indicating that the two structures are highly consistent.
[0073] Figure 4 This is a schematic diagram of the representative short peptide backbone three-dimensional structure generated by AI in this invention. The yellow area represents the peptide backbone generated by RFdiffusion; the gray area represents the LRP1 surface; and the semi-transparent blue area represents the binding site. The coordinate grid represents the three-dimensional spatial reference frame (unit: Å).
[0074] AlphaFold3 prediction template modeling parameter description:
[0075] Extract the template modeling score (pTM), interface score (ipTM), and minimum prediction error (min PAEinteraction). The selection criteria are: pTM > 0.8, min PAE < 1.5 Å, fraction_disordered < 0.3, and no structural conflict (has_clash = False).
[0076] • pTM is used to evaluate the accuracy of the overall structure of the complex. It is a prediction TM score superimposed between the predicted structure and the assumed true structure. A pTM score higher than 0.5 means that the overall predicted folding of the complex is likely similar to the true structure. A pTM score lower than 0.5 means that the predicted structure may be incorrect.
[0077] • ipTM measures the accuracy of predicted relative positions of subunits forming protein-protein complexes. It is based on the concept of TM-score but focuses on the interfacial interactions between different proteins in the predicted protein complex. Values above 0.8 indicate high-quality predictions with confidence, while values below 0.6 indicate possible prediction failure. ipTM values between 0.6 and 0.8 represent a gray area where predictions may be correct or incorrect. It is important to note that disordered regions and regions with low pLDDT scores can negatively impact the ipTM score.
[0078] • PAE (Predicted Aligned Error) is a metric used by AlphaFold2 to measure the confidence level of the model in the relative positions of two residues in the predicted structure. PAE is defined as the error in the expected position of residue X when the predicted structure aligns with the actual structure at a certain residue Y, expressed in angstroms (Å).
[0079] In existing studies, predicted alignment error (PAE) is often established as a key indicator for evaluating the quality of protein complex models and as an important basis for screening and designing binding proteins. For example, in the design of binding proteins based on AlphaFold2 (AF2), if the PAE value of the interaction region is less than 10 Å, it is considered as the design standard for effective binders (Improving de novo Protein Binder Design with Deep Learning[4]). With the release of AlphaFold3, its accuracy has been significantly improved, and the corresponding screening threshold is also more stringent. The latest literature, such as AlphaProteo, uses a minimum interface PAE value of less than 1.5 Å and a PTM score of greater than 0.8 as the screening condition for high-quality complex structures. This study adopts the above AF3 standard, namely, min PAE interaction < 1.5 Å and PTM > 0.8, as the screening criterion for identifying high-confidence binding proteins.
[0080] • fraction_disordered: A scalar in the range of 0-1 that indicates which part of the predicted structure is disordered, measured by the accessible surface area.
[0081] • has_clash: A boolean value indicating whether the structure has a large number of conflicting atoms (more than 50% of the chain, or a chain with more than 100 conflicting atoms).
[0082] • anking_score: A scalar in the range of [-100, 1.5], which can be used for ranking prediction. It combines ptm, iptm, fraction_disordered, and has_clash into a single number, as shown in the formula: S = 0.8 × ipTM + 0.2 × pTM + 0.5 × disorder − 100 × has_clash. The top 5% of candidate sequences are selected based on the S value from high to low and proceed to step S105.
[0083] Figure 5 This is a structural diagram of the interface between the representative short peptide of this invention and LRP1. The solid black lines represent hydrogen bonds; the dashed black lines represent hydrophobic interactions; and the dashed yellow lines represent salt bridges. Key residues Arg61, Tyr88, and Glu92 are marked. The background is a potential distribution diagram of the LRP1 surface, with blue representing positive charges, red representing negative charges, and white representing neutral regions.
[0084] So far, it can be seen that AlphaFold3 has demonstrated high accuracy in predicting the structure of the system studied in this paper. Therefore, it can be used to efficiently screen candidate binders based on its scoring system. The following is a calculation of the comprehensive affinity of the predicted complex.
[0085] Calculation of the overall affinity of S105, LRP1 and peptide backbone complex
[0086] The PRODIGY energy model calculates the binding free energy (ΔG) and dissociation constant (Kd) of the complex. The screening criteria are set as ΔG < −6 kcal / mol and Kd ≤ 1×10⁻ 7 M, sequences that meet the criteria are identified as high-affinity binding peptides; or ΔG ≤ −5.5 kcal / mol, Kd ≤ 10⁻6 M.
[0087] PRODIGY (Protein Binding Energy Prediction) is a tool for predicting protein binding energies, utilizing the three-dimensional structure of protein-protein complexes to predict their binding affinity. PRODIGY is a highly efficient contact-based method that, while estimating binding free energies and dissociation constants, also reveals the structural determinants of protein-protein interactions. By combining interfacial contact properties with non-interacting surface features, PRODIGY can make reliable predictions, which is crucial for understanding intermolecular interactions, guiding the development of therapeutic approaches, and designing protein complexes.
[0088] like Figure 6 The image shown is a scatter plot of predicted binding energies for candidate peptides provided by this invention. It is commonly used to display the predicted binding energy distribution of each ligand in virtual screening or molecular docking results. A detailed analysis follows:
[0089] Overall meaning of the diagram
[0090] The horizontal axis (X-axis) is called binder_number, which represents the number of each short peptide (from 1 to approximately 1000).
[0091] The vertical axis (Y-axis) represents the predicted binding affinity (kcal / mol), indicating the predicted binding free energy. The lower the binding energy (the more negative), the more stable the predicted binding, and theoretically, the stronger the affinity.
[0092] Figure 6 Interpretation of elements: Scattered dots (colored dots) Each dot represents the predicted binding energy of a ligand.
[0093] The colors are represented by the color bar on the right, ranging from yellow (high energy, weak binding) to purple (low energy, strong binding).
[0094] It can be seen that the overall data distribution is between approximately -4.0 and -7.5 kcal / mol.
[0095] The red dashed line (Threshold = -6) represents the set filtering threshold, i.e.:
[0096] Molecules with scores below -6 kcal / mol are considered to be potentially good candidate ligands;
[0097] A score higher than -6 kcal / mol is considered to indicate a weaker binding.
[0098] The color bar is labeled with the unit kcal / mol, indicating that the color mapping is directly related to the strength of the binding energy.
[0099] Data Distribution and Interpretation
[0100] The predicted binding energies of most ligands are concentrated in the range of -5.0 to -5.5 kcal / mol.
[0101] A small number of ligands have a concentration below -6 kcal / mol; these may be strong binding agents (potential hits).
[0102] A small number of points reaching below -7 kcal / mol indicate high affinity.
[0103] Interpreting PRODIGY Output Results:
[0104] • No. of intermolecular contacts: 50: This indicates that there are a total of 50 intermolecular contacts in the protein-protein complex.
[0105] • No. of charged-charged contacts: 6: This refers to a number of contacts formed between charged groups of 6.
[0106] • No. of charged-polar contacts: 7: This indicates that there are 7 contacts between charged groups and polar groups.
[0107] • No. of charged-apolar contacts: 8: This means that 8 contacts are formed between the charged group and the nonpolar group.
[0108] • No. of polar - polar contacts: 7: This indicates that there are 7 contacts between polar groups.
[0109] • No. of apolar - polar contacts: 15: This indicates that there are 15 contacts between the nonpolar group and the polar group.
[0110] • No. of apolar - apolar contacts: 12: indicates that there are 12 contacts between nonpolar groups.
[0111] • Percentage of apolar NIS residues: 39.18: This refers to the percentage of nonpolar residues on the non-interacting surface of the protein-protein complex, which is 39.18%.
[0112] • Percentage of charged NIS residues: 29.48: This indicates that the percentage of charged residues on the non-interacting surface of the protein-protein complex is 29.48%.
[0113] • Predicted binding affinity (kcal.mol⁻¹): -6.4: The predicted binding free energy (ΔG) of the protein-protein complex is -6.4 kcal / mol. ΔG < 0 (negative value) indicates that the binding process is spontaneous. The more negative the value (i.e., the larger the absolute value), the more energy is released during the binding process, the more stable the complex, and the stronger the binding affinity.
[0114] • Predicted dissociation constant (M) at 25.0˚C: 1.1e - 07: At 25.0°C, the predicted dissociation constant of the protein-protein complex is 1.1 × 10⁻⁻⁷. 7 M. The dissociation constant is an indicator of the tendency of a complex to dissociate. The smaller the value, the more stable the complex is and the less likely it is to dissociate.
[0115] Figure 7 This is a combination of binding energy statistics and interfacial contact distribution diagrams. The left diagram shows the distribution of the composition ratio of interacting surface residues, displaying the differences in the distribution of polar and charged residues at the interface in the form of a violin diagram. The right diagram shows the statistical distribution of the number of interfacial contacts, used to analyze the contribution ratio of different types of interactions at the binding interface.
[0116] Figure 7 The left-middle panel is a violin plot that compares the percentage of polar and charged NIS residues on the interacting surface.
[0117] The X-axis displays two sets of data:
[0118] Percentage of polar NIS residues (Apolar NIS residues)
[0119] Percentage of charged NIS residues
[0120] The Y-axis shows the percentage of residues, that is, the percentage of each group of residues on the interacting surface.
[0121] Figure 7 The right figure shows the comparison results of the number of protein interface contacts, presented in the form of a violin plot.
[0122] The distribution of each violin shape represents a statistical result of a category of contact, including:
[0123] All intermolecular contacts.
[0124] And more specific contact types, such as charged-charged and charged-polar.
[0125] Overall trend
[0126] The total number of intermolecular contacts (blue) is significantly higher than other categories, typically concentrated in 30–45 contacts, indicating that the interface as a whole forms more stable interactions.
[0127] Contact of charged residues
[0128] The number of charged-to-charged (orange) and charged-to-polarity (red) contacts is relatively small, mostly between 0 and 5;
[0129] This indicates that although electrostatic effects exist, they account for a limited proportion of interfacial stability.
[0130] Charged-hydrophobic contact
[0131] (Brown) The number of residues is between 5 and 15, suggesting that a small number of charged residues are involved in the interaction under nonpolar conditions, possibly reflecting salt bridges or buried charges.
[0132] Contact between polar and hydrophobic residues
[0133] The relatively small number of polar-polar (purple) and hydrophobic-polar (yellow) contacts indicates that the interface is dominated by hydrophobic interactions.
[0134] Hydrophobic – Hydrophobic Contact
[0135] The number of (cyan) contacts is second only to the total number of contacts, concentrated between 20 and 30, indicating that a large number of hydrophobic interactions at the interface contribute the most to the stability of the complex.
[0136] Figure 8 This is a graph showing the distribution of energy screening results and predicted dissociation constant (Kd).
[0137] The horizontal axis represents the short peptide number, and the vertical axis represents the predicted Kd value (unit: M). Colors from purple to yellow indicate binding strength from strong to weak; the red dashed line indicates the screening threshold (10⁻⁻⁴). 6 M). The size of the scatter points in the graph corresponds to the absolute value of the binding energy (|ΔG|). The larger the point, the lower the binding energy (higher affinity).
[0138] Figure 8 This displays the distribution of the predicted dissociation constant (Kd) across different ligands, used to measure the binding strength between each candidate ligand and the receptor. A smaller predicted dissociation constant (Kd) indicates a tighter binding (stronger affinity).
[0139] Each point represents the predicted Kd value for a ligand.
[0140] The lower the Kd value, the stronger the binding of the ligand to the target;
[0141] A higher Kd value indicates a weaker binding.
[0142] The color represents the relative level of the Kd value, and the color bar on the right indicates the corresponding range:
[0143] Purple / Blue: Low Kd (High Affinity)
[0144] Yellow: High Kd (low affinity)
[0145] In the dynamics of combination:
[0146]
[0147] A lower Kd indicates a slower dissociation rate (k_off) and more stable binding. Therefore, this graph shows the distribution range of ligand affinity, helping to determine which short peptides are worth proceeding to the next stage of validation.
[0148] Kinetic simulations of the first 30 candidate peptide-LRP1 complexes were performed in GROMACS for 100 ns. The root mean square deviation (RMSD), hydrogen bond retention, and solvent accessible surface area (SASA) were calculated. The results showed that the RMSD of the LRP1-Pep5 and LRP1-Pep12 complexes remained stable within 1.8 Å in the 50–100 ns range, with hydrogen bond retention exceeding 85%, indicating stable complex structures.
[0149] Final result:
[0150] Thirty-one high-affinity short peptides (LRP1-Pep1 to LRP1-Pep31) were ultimately obtained, with binding energies ranging from −6.1 to −7.5 kcal / mol. These peptides can be further used for chemical synthesis and exosome membrane modification to verify their delivery efficiency across the blood-labyrinth barrier.
[0151] Example 2 verifies the predictive effect of the LRP1-Pep1~LRP1-Pep31 composite structure on the transvascular-labyrinthine barrier delivery of the present invention.
[0152] 2.1 AI Structural Prediction Validation
[0153] In the AlphaFold3 prediction phase, the LRP1 structure and candidate peptide sequences are input, and indicators such as pTM, ipTM, and PAE are extracted. If pTM > 0.8 and min PAE < 1.5 Å, the complex structure has high confidence. The interfacial energy of the first 50 complexes is calculated and compared with experimentally known binding modes. If RMSD < 2 Å, it indicates that the AI prediction results have high consistency and reliability.
[0154] 2.2 Molecular dynamics simulation verification
[0155] The dynamic stability of the complex was analyzed by performing a 100 ns simulation using GROMACS (CHARMM 36m force field, temperature 300 K, pressure 1 bar).
[0156] The results showed that the RMSD curves of LRP1-Pep5 and LRP1-Pep12 stabilized after 50 ns (fluctuation range 1.5–1.8 Å), with hydrogen bond retention exceeding 85% and energy fluctuation less than 5 kcal / mol, demonstrating that the screened complexes have good dynamic stability and binding persistence.
[0157] In summary, this invention, based on the RFdiffusion protein design model, uses the crystal structure of the LRP1 protein as a target to systematically generate and construct 1000 candidate binding agent sequences targeting this protein. To further screen high-affinity candidate molecules, AlphaFold3 was used to perform batch three-dimensional structure prediction of the aforementioned complexes. Rigorous screening was conducted based on multi-dimensional structural confidence indicators such as prediction template modeling score (pTM) and minimum pairing contact error (min PAE interaction), ultimately yielding 31 highly reliable candidate binding agent sequences.
[0158] In subsequent studies, users can rank these candidate sequences by their binding free energy to the LRP1 protein and select the best ones for experimental verification to evaluate their actual binding performance and functional activity.
[0159] This invention utilizes an artificial intelligence-designed LRP1 target peptide, which is modified onto the surface of an exosome membrane via lipid intercalation or chemical coupling, to deliver antioxidant enzymes, nucleic acids, or small molecule drugs across the blood-labyrinth barrier, enabling precise treatment of the cochlear region. Figure 9 This is a schematic diagram illustrating the application of LRP1 target peptide-modified exosomes for transBLB delivery in this invention.
[0160] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for screening LRP1 target peptides across the blood-labyrinth barrier based on artificial intelligence, characterized in that, Includes the following steps: Target identification: The region of residues 5 to 117 in the crystal structure of low-density lipoprotein receptor-associated protein 1 (LRP1) was selected as the target binding site; Backbone generation: Based on the spatial configuration of the target binding site, a binding peptide backbone of 10-15 amino acids in length is generated using the diffusion generation model RFdiffusion. Sequence design: ProteinMPNN message-passing neural network was used to design and optimize the amino acid sequence of the generated peptide backbone to obtain candidate peptide sequences; Structure prediction and initial screening: AlphaFold3 was used to predict the structure of the complex between the candidate peptide sequence and LRP1, and the candidate peptides were initially screened based on the predicted template modeling score pTM, interface template modeling score ipTM, and minimum prediction alignment error min PAE. A comprehensive scoring function S was constructed to rank the initially screened candidate peptides. Energy validation: The predicted binding free energy ΔG and dissociation constant Kd of the candidate peptides after preliminary screening with LRP1 were calculated using the PRODIGY model. ΔG ≤ −6.0 kcal / mol; Kd ≤ 10⁻ 7 M's high-affinity LRP1-targeting peptide.
2. The method according to claim 1, characterized in that, In the structure prediction and initial screening, a comprehensive scoring function S is constructed to rank the screened candidate peptides. The function is: S = a1 × ipTM + a2 × pTM - a3 × Disorder - a4 × Clash, where Disorder is the proportion of structural disorder, Clash is the atomic conflict index, and a1 to a4 are adjustable weight coefficients. The top 5% of candidate sequences are selected based on the S value to enter the energy verification.
3. The method according to claim 2, characterized in that, In the comprehensive scoring function, a1=0.8, a2=0.2, a3=0.5, a4=100.
4. The method according to claim 1, characterized in that, The preliminary screening criteria in the structure prediction and initial screening are: pTM > 0.8 and min PAE < 1.5 Å.
5. The method according to claim 1, characterized in that, The method further includes the step of performing molecular dynamics simulations on the screened high-affinity LRP1-targeting peptide to verify the structural stability of the complex formed with LRP1 under physiological conditions.
6. The method according to claim 5, characterized in that, The molecular dynamics simulation lasted for 100 nanoseconds. When the root mean square deviation (RMSD) of the complex interface was stable within 2 Å and the hydrogen bond retention rate was greater than 80%, the structural stability of the targeting peptide was determined to meet the requirements.
7. A high-affinity LRP1-targeting peptide obtained by screening using the method described in any one of claims 1-6.
8. A drug delivery system, characterized in that, The product comprises the high-affinity LRP1 targeting peptide of claim 7 and a delivery vector; wherein the high-affinity LRP1 targeting peptide is modified on the surface of the delivery vector.
9. The drug delivery system according to claim 8, characterized in that, The delivery vector is an exosome.
10. Use of the drug delivery system of claim 8 or 9 in the preparation of a medicament for treating inner ear diseases.