A method, apparatus, and program product for predicting blood brain barrier penetration efficacy of a nanoparticle
By extracting key features of nanoparticles using electron density field descriptors and AI models, the problem of data scarcity in nanomedicine crossing the blood-brain barrier was solved, enabling high-precision prediction of blood-brain barrier penetration and improving the efficiency of nanomedicine design.
Patent Information
- Application Number
- CN202610354358.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-19
- Estimated Expiration
- 2046-03-23
AI Technical Summary
In existing technologies, the scarcity of data on nanomedicines crossing the blood-brain barrier is a serious problem, leading to an unbalanced data distribution and limited prediction accuracy in nanomedicine design. In particular, there is a very limited standardized dataset for combinations of metal ions and organic compounds, making it difficult to effectively assess the blood-brain barrier penetration capability of nanoparticles.
Using the electron density field as a unified physical quantity descriptor across material categories, key features related to electron distribution on the surface of nanoparticles and BBB penetration are extracted through an AI model. By combining topological sequence, conformational sequence and electron density point cloud data, a universal penetration rate prediction model is established, and feature extraction and evaluation are performed using a three-dimensional network model and a discriminator model.
It enables high-precision prediction of blood-brain barrier penetration rates for various nanomaterials, reducing the risk of clinical application failure, improving the efficiency and accuracy of nanomedicine design, and shortening the R&D cycle from molecular design to stable nanoparticle construction.
Smart Images

Figure CN121885004B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent healthcare, specifically to a method, device, program product, and computer-readable storage medium for predicting the blood-brain barrier penetration efficiency of nanoparticles. Background Technology
[0002] The blood-brain barrier (BBB) is the primary obstacle in drug development for the central nervous system (CNS). This physiological barrier, formed by the tight junctions of brain microvascular endothelial cells and neurovascular units formed by pericytes and astrocyte terminale, strictly regulates the exchange of substances between blood and brain tissue, protecting the CNS from toxins and pathogens. However, this high selectivity also makes it difficult for over 98% of small molecule drugs and almost all large molecule biologics to penetrate the brain parenchyma. Consequently, although over 50 nanomedicines have been approved by regulatory agencies in the past thirty years, only NanoTherm®, approved by the European Medicines Agency, is truly used for the treatment of brain diseases, highlighting the enormous challenge of crossing the blood-brain barrier. Currently, researchers are exploring the introduction of artificial intelligence methods into the design process of brain-targeted nanomedicines. Recent research reports an AI model based on graph neural networks to predict the BBB penetration ability of small molecule ligands. This model screened functionalized lipid nanoparticles containing eight BBB-interacting small molecules (including acetylcholine, glucose, nicotine, and memantine). It found that acetylcholine-modified LNPs exhibited excellent brain targeting and gene expression efficiency in both in vitro BBB models and in vivo mouse models, and the AI predictions highly matched experimental results. While artificial intelligence methods have brought new possibilities to the design of brain-targeted nanomedicines, current technologies still face multiple challenges, with data scarcity being particularly prominent.
[0003] Standardized datasets specifically targeting the penetration of the BBB by combinations of "metal ions + organic compounds" are extremely limited. Unlike traditional small molecule organic drugs, nanomedicines have a higher dimensionality of composition (encompassing multiple variables such as metal centers, organic ligands, surface modifications, and particle size distribution), and experimental results are significantly affected by differences in synthesis processes, characterization methods, and model systems. In existing research, high-quality nano-biological effect data that can truly be used for model training often comes from scattered literature reports, and failed experimental cases are rarely published, resulting in a severely imbalanced data distribution. Summary of the Invention
[0004] To address the above problems, this invention provides a method for predicting the blood-brain barrier penetration efficiency of nanoparticles, specifically including:
[0005] Obtain the topological sequence, conformational sequence, and electronic density point cloud of nanoparticles;
[0006] The topological sequence and conformational sequence are vectorized and encoded to obtain the topological sequence representation and the conformational sequence representation, and the electron density point cloud is obtained by feature extraction through a three-dimensional network model to obtain the electron density representation.
[0007] The fused characterization is obtained by integrating topological sequence characterization, conformational sequence characterization and electronic density characterization.
[0008] The fused representation is fed into a neural network for feature extraction to obtain key features;
[0009] The blood-brain barrier penetration ability was assessed using the aforementioned key features, resulting in high, medium, and low assessment results.
[0010] Optionally, the nanoparticles include one or more of different types of nanoparticles such as small molecules, peptides, and RNA. The electron density point cloud is voxelized to obtain the electron cloud density, and the electron cloud density is characterized by cross-scale feature extraction through a three-dimensional network model. The three-dimensional network model includes any one or more of the following: 3DCNN, 3D ResNet, 3D DenseNet, 3D MobileNet, and Swin3D.
[0011] Optionally, the neural network is any one or more of the following discriminator models: SwinTransformer Discriminator, TransGAN Discriminator, DeiT Discriminator, CSWinTransformer Discriminator, BEiT Discriminator; the fused representation is input to the discriminator model for feature extraction to obtain key features.
[0012] Optionally, the discriminator model consists of N layers of discriminator modules, where N is a natural number greater than 1. Each discriminator module is a model constructed using a bidirectional self-attention mechanism, and the fusion representation extracts key features through the N-layer bidirectional self-attention mechanism.
[0013] Optionally, the discriminator model is trained by: acquiring electron cloud density data, topological data, and conformational data of small molecules, peptides, and small RNA nanoparticles; using the experimental efficiency of small molecules, peptides, and small RNA nanoparticles in penetrating the blood-brain barrier as a label; inputting the electron cloud density data, topological data, conformational data, and labels of small molecules, peptides, and small RNA nanoparticles into the discriminator model to be trained for training until the loss function remains unchanged, thus obtaining the discriminator model.
[0014] Optionally, the topological sequence of the nanoparticles includes the topological sequence of the metal ions added to the nanoparticles, the organic compound additives to the nanoparticles, and the target compound molecules; the conformational sequence of the nanoparticles includes the three-dimensional coordinates of atoms and non-atoms in the topological sequence.
[0015] Optionally, each element in the topological sequence represents a token. The atom includes a metal ion token added to the nanoparticle, an organic compound additive token for the nanoparticle, and a topological sequence token for the target compound molecule. The non-atoms include the start token, end token, metal ion symbol token, organic compound additive symbol token, and topological sequence symbol token for the target compound molecule. The three-dimensional coordinates of the atom are obtained according to the structure of the topological sequence, and the three-dimensional coordinates of the non-atoms are assigned a value of zero.
[0016] The purpose of this invention is to provide a computer program product comprising a computer program or instructions, wherein the computer program or instructions are executed by a processor to implement the above-described method for predicting the blood-brain barrier penetration efficiency of nanoparticles.
[0017] The purpose of this invention is to provide a computer device comprising a memory, a processor, and a computer program or instructions stored in the memory, wherein the computer program or instructions are executed by the processor to implement the above-described method for predicting the blood-brain barrier penetration efficiency of nanoparticles.
[0018] The purpose of this invention is to provide a computer-readable storage medium storing a computer program or instructions thereon, which is executed by a processor to implement the above-described method for predicting the blood-brain barrier penetration efficiency of nanoparticles.
[0019] Advantages of this invention:
[0020] 1. Overcoming the challenge of data scarcity in BBB penetration prediction: Existing BBB prediction models suffer from limited accuracy due to a lack of experimental data. This invention innovatively employs the electron density field as a unified physical quantity descriptor across material categories, and uses an AI model to extract key features related to the electron distribution on the nanoparticle surface and BBB penetration, thereby establishing a universal penetration rate prediction model.
[0021] 2. To assess and predict the BBB penetration capability of newly designed nanoparticles and reduce the chance of failure in later clinical applications, this invention proposes a method for predicting the BBB penetration capability of nanoparticles. By extracting key features from the topological sequence, conformational sequence, and electron density point cloud data of nanoparticles, the method predicts the BBB penetration capability, providing preliminary auxiliary assessment for the usability verification of designed nanoparticles.
[0022] 3. When designing nanomedicines capable of penetrating the blood-brain barrier, the selection of added metal ions and organic compound additives is often determined through trial and error, relying on the experience of pharmaceutical experts. The final selection of suitable metal ions and organic compound additives largely depends on human factors. Screening for metal ions and organic additives is inefficient and difficult to precisely control particle size, resulting in high trial-and-error costs. This invention utilizes artificial intelligence technology to construct a nanoparticle assembly prediction model, providing assembly recommendations for nanoparticles of specified sizes, including the metal ions, organic compound additives, and corresponding molecular conformations for those sizes. This significantly shortens the R&D cycle from molecular design to stable nanoparticle construction. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A schematic diagram of the process for predicting the blood-brain barrier penetration efficiency of nanoparticles provided in an embodiment of the present invention;
[0025] Figure 2 A schematic diagram of a blood-brain barrier penetration efficacy prediction system for nanoparticles provided in an embodiment of the present invention;
[0026] Figure 3 A schematic diagram of a computer device provided in an embodiment of the present invention;
[0027] Figure 4 A flowchart illustrating the penetration capability model of nanoparticles in BBB provided in an embodiment of the present invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0029] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as S101, S102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0030] Figure 1 A schematic diagram of the method for predicting the blood-brain barrier penetration efficiency of nanoparticles provided in this embodiment of the invention is shown, specifically including:
[0031] S1: Obtain the topological sequence, conformational sequence, and electronic density point cloud of nanoparticles;
[0032] In one embodiment, the topological sequence of the nanoparticles includes the topological sequence of the metal ions added to the nanoparticles, the organic compound additives to the nanoparticles, and the target compound molecules; the conformational sequence of the nanoparticles includes the three-dimensional coordinates of atoms and non-atoms in the topological sequence.
[0033] Each element in the topological sequence represents a token. The atoms include metal ion tokens added to the nanoparticles, organic compound additive tokens for the nanoparticles, and topological sequence tokens for the target compound molecule. The non-atoms include the start token, end token, metal ion symbol token, organic compound additive symbol token, and topological sequence symbol token for the target compound molecule. The three-dimensional coordinates of the atoms are obtained according to the structure of the topological sequence, and the three-dimensional coordinates of the non-atoms are assigned a value of zero.
[0034] In one embodiment, the nanoparticles include one or more of different types of nanoparticles such as small molecules, peptides, and RNA. The electron density point cloud of the nanoparticles is voxelized to obtain the electron cloud density. The electron cloud density is characterized by cross-scale feature extraction using a three-dimensional network model. The three-dimensional network model includes any one or more of the following: 3D CNN, 3D ResNet, 3D DenseNet, 3D MobileNet, and Swin3D.
[0035] In one specific embodiment, the topological sequence of nanoparticles is formatted as follows:
[0036] <sos> <metal> … <sol> … <mol> … <sos>
[0037] For ease of display, some content has been omitted here. <sos>This indicates the start of the entire topological sequence. <eos>This indicates the termination of the entire topological sequence. <metal>This indicates the metal ions added to form nanoparticles, followed by a discrete metal ion token. <sol>This indicates an organic compound additive that forms nanoparticles, followed by an additive topological sequence token that follows the SMILES serialization format. <mol>This represents the topological sequence of the target compound molecule.
[0038] In the conformational sequence of nanoparticles, <sos> 、 <metal> 、 <sol> 、 <mol> 、 <sos>Non-atomic tokens are assigned coordinates of (0,0,0). The rest are derived from the (x,y,z) three-dimensional coordinates of the atoms.
[0039] S2: The topological sequence and conformational sequence are vectorized and encoded to obtain the topological sequence representation and the conformational sequence representation. The electron density point cloud is obtained by feature extraction through a three-dimensional network model to obtain the electron density representation.
[0040] In one specific embodiment, the topological and conformational sequences of nanoparticles are characterized using an embedded vector module. The electron density of the nanoparticles is represented as a point cloud, which is then voxelized and extracted by a 3DCNN module to form an electron density characterization.
[0041] In one specific embodiment, the BBB penetration prediction model for nanoparticles is as follows: Figure 4 As shown, it includes a characterization module, a discriminator, and a nano-barrier penetration discrimination module. The function of the characterization module is to receive the input of nanoparticles and electron density, process them, and then characterize them into features that the model can understand.
[0042] The characterization module consists of an embedded vector module and a 3D CNN module. The embedded vector module is used to characterize the serialized topological and conformational sequences of the nanoparticles. The 3D CNN module in the characterization module is used to characterize the electron density of the nanoparticles. Finally, the embedded topological sequence characterization, embedded conformational sequence characterization, and electron density characterization of the nanoparticles are integrated to form features that can be understood by the AI model.
[0043] In one specific embodiment, existing BBB prediction models are typically built separately for a single material category (such as gold nanoparticles, polymer nanoparticles, etc.), and their accuracy is limited by the limited experimental data within that category. This invention innovatively uses the electron density field as a unified physical quantity descriptor across material categories, fundamentally changing this paradigm. Traditional methods face serious "data barriers" and "model silos." Experimental data from different material categories cannot be shared, requiring each category to accumulate data from scratch and train independent models. This not only results in high R&D costs and long development cycles but also prevents the model from generalizing to entirely new material systems with very little or no experimental data, severely limiting its practical application and predictive capabilities. This method, through the electron density field descriptor, abstracts and focuses on the common physical essence of the surface interactions of nanoparticles from different materials. The AI model can extract key electronic structure features related to BBB penetration across categories from data from disparate materials such as silicon, gold, and polymers. This breaks down the boundaries of material categories, enables knowledge transfer, and allows a single model to accurately predict the transmittance of various known and even unknown nanomaterials, greatly improving the model's universality, practicality, and efficiency in exploring new materials.
[0044] S3: The topological sequence characterization, conformational sequence characterization and electronic density characterization are fused to obtain the fused characterization;
[0045] In one embodiment, the fusion is achieved by fusing topological sequence characterization, conformational sequence characterization, and electronic density characterization through cross-attention to obtain a fused characterization.
[0046] Topological sequence characterization describes the connections between atoms (chemical bonds, molecular diagram structure), answering "how atoms are connected". Conformational sequence characterization describes the arrangement of atoms in three-dimensional space (conformational isomerism, spatial orientation), answering "how atoms are arranged". Electron density characterization describes charge distribution, frontier orbitals, and electrostatic potential, answering "how electrons are distributed", which is directly related to coordination ability and the strength of interaction with blood-brain barrier receptors.
[0047] The process of cross-attention fusion between topology and conformation:
[0048] Using topological representation as the query and conformational representation as the key and value, the dot product of the query and all keys is calculated and normalized using the Softmax function to obtain the attention score matrix. The attention score A is then used to perform a weighted summation of the value V to obtain the new representation after the topological modality has perceived the conformational information.
[0049] Symmetrically, the present invention performs topological and conformational interactions, topological and electronic interactions, and conformational and electronic interactions. After completing the above-mentioned multi-order cross-attention interactions, six interactive representations (one for each direction) are obtained. By splicing or weighted summation, all interactive representations are fused with the original modal representations to obtain a fused representation.
[0050] The interaction between topology and conformation yields new representations of topological modality-sensing conformational information and new representations of conformational modality-sensing topological information, which solves the problem of whether the connection mode of flexible molecules under a specific three-dimensional shape is stable.
[0051] The new characterization of topological mode sensing electron density information and the new characterization of electron density mode sensing topological information obtained by topological and electronic interaction solves the problem of whether a specific atomic connection mode (such as a coordinating group) has a suitable electron density to bind metal ions.
[0052] The interaction between conformation and electron yields new representations of conformation modality-sensing electron density information, new representations of electron density modality-sensing conformation information, new representations of topological modality-sensing electron density information, and new representations of electron density modality-sensing topological information.
[0053] S4: Input the fused representation into a neural network to extract key features;
[0054] In one embodiment, the neural network is any one or more of the following discriminator models: SwinTransformer Discriminator, TransGAN Discriminator, DeiT Discriminator, CSWinTransformer Discriminator, BEiT Discriminator; the fused representation is input to the discriminator model for feature extraction to obtain key features.
[0055] In one embodiment, the discriminator model consists of N layers of discriminator modules, where N is a natural number greater than 1. Each discriminator module is a model constructed using a bidirectional self-attention mechanism, and the fusion representation extracts key features through the N-layer bidirectional self-attention mechanism.
[0056] Each layer of the bidirectional self-attention mechanism includes a multi-head self-attention layer, a feedforward neural network layer (which performs a non-linear transformation on the representation at each location, typically consisting of two linear layers and a ReLU activation function), and a normalization layer. The input of the previous layer's bidirectional attention mechanism is fused with the output of the current layer via residual connections. The discriminator model uses a cross-entropy function as its loss function, and the optimizer during training is Adam / AdamW with an initial learning rate of 1e-4 ~ 1e-3. The final classification accuracy for the blood-brain barrier (BBB) is 85%.
[0057] The discriminator model is trained by acquiring electron cloud density data, topological data, and conformational data of small molecules, peptides, and small RNA nanoparticles, using the experimental efficiency of small molecules, peptides, and small RNA nanoparticles in penetrating the blood-brain barrier as a label. The electron cloud density data, topological data, conformational data, and labels of small molecules, peptides, and small RNA nanoparticles are input into the discriminator model to be trained for training until the loss function remains unchanged, thus obtaining the discriminator model.
[0058] In one specific embodiment, the discriminator's function is to receive the output of the characterization module and extract key features based on a bidirectional self-attention mechanism to facilitate the final determination of the nanoparticles' BBB penetration capability.
[0059] The discriminator adopts a Transformer-like architecture, consisting of N layers of discriminator modules. Each layer's discriminator module receives the output of the layer above as input. The first layer's discriminator module receives the output of the representation module as input. A bidirectional self-attention mechanism is used in all discriminator modules.
[0060] In one specific embodiment, data composed of different types (small molecules, peptides, small RNA) are considered to be of the same type from the perspective of electron density characterization, which alleviates the problem of limited data transmission for single-type BBB. Based on unified electron density characterization, different types of nanoparticles such as small molecules, peptides, and RNA are uniformly converted into multi-resolution electron density point clouds. These point clouds are then voxelized, and cross-scale features are extracted using three-dimensional convolutional kernels in a 3D CNN module, solving the prediction bias problem caused by data type differences in traditional methods. Data composed of different types (small molecules, peptides, small RNA) are considered to be of the same type from the perspective of electron density characterization, which alleviates the problem of limited data transmission for single-type molecule BBB.
[0061] When training the discrimination model, BBB transmission data of tens of thousands of small molecules, peptides, and small RNA nanoparticles that can be collected from literature / patents are used. Multi-resolution electron density is used as the characterization of nanoparticles, and the experimental efficiency of nanoparticles in BBB transmission is used as a label to train the discrimination model.
[0062] The characterization module and the discriminator module are integrated. By collecting BBB transmission data of tens of thousands of small molecules, peptides, and small RNA nanoparticles, the multi-resolution electron density is extracted through the three-dimensional network module as the characterization of the nanoparticles. The topological sequence and conformational sequence are then fused and input to the discriminator for discriminator training.
[0063] S5: The blood-brain barrier penetration ability is assessed using the aforementioned key features to obtain high, medium, and low assessment results.
[0064] In one embodiment, the assessment of the blood-brain barrier penetration capability of the nanoparticles is achieved by dividing the continuous variable of blood-brain barrier penetration rate into high, medium, and low categorical variables according to thresholds in a drug research scenario, and then dividing the predicted probability of the discriminator model into predicted classification results through threshold division.
[0065] In one specific embodiment, the function of the nanoparticle BBB penetration discrimination module is to receive the output of the discriminator module, output the BBB penetration capability label of the nanoparticles, and classify them into three categories: high, medium, and low.
[0066] Here, the present invention divides the BBB penetration rate, which is a continuous variable, into three categorical variables (high, medium, and low) according to the threshold in the actual drug research scenario. This transforms the regression problem into a classification problem, reduces the training difficulty of the discrimination model, and thus improves the accuracy of the discrimination model.
[0067] In one specific embodiment, nanoparticles are assembled and designed, and their BBB penetration capability is predicted and evaluated. Specifically, the topological sequence, conformational sequence, and electron density cloud of the assembled nanoparticles are obtained, vectorized, and then input to a discriminator for penetration capability evaluation (in the nanoparticle BBB barrier penetration discrimination module). Further, the topological and conformational sequences of the assembled nanomaterials are predicted using a nano-assembly prediction model, specifically:
[0068] Obtain the three-dimensional compound molecules and the size of the target nanoparticles;
[0069] The topological sequence is obtained by serializing the topological structure of a three-dimensional compound molecule.
[0070] The three-dimensional coordinates of each atom are obtained by the order of atoms in the topological structure, thus obtaining the molecular conformation sequence;
[0071] The size sequence of the target nanoparticles is obtained by discretizing the size.
[0072] The first topological sequence is obtained by integrating the topological sequence and the size sequence, and the first conformational sequence is obtained by integrating the molecular conformational sequence and the size sequence.
[0073] The first topological sequence and the first conformational sequence are input into the second neural network for feature extraction and prediction to obtain the predicted topological sequence and the predicted conformational sequence. The predicted topological sequence includes the metal ions and organic compound additives that need to be added to nanoparticles of a specified size. The predicted conformational sequence includes the three-dimensional coordinates of the atoms in the predicted topological sequence, and the corresponding three-dimensional conformation is obtained through the three-dimensional coordinates.
[0074] The second neural network includes any one or more of the following: Transformer, residual network, Mamba, Mamba-2, ERNIE, XLNet, and T5.
[0075] The second neural network includes an encoder and a decoder. The encoder includes an N-layer encoder module, and the decoder includes an N-layer decoder module, where N is a natural number greater than 1. The first topological sequence and the first conformation sequence are input into the first-layer encoder module and encoded by the N-layer encoder module to obtain encoded features. The encoded features are then decoded by the N-layer decoder module to obtain decoded features. The metal ions, organic compound additives, and corresponding three-dimensional conformations of nanoparticles of a specified size are predicted using the decoded features.
[0076] The encoder module includes a bidirectional self-attention mechanism, through which the first topological sequence and the first conformation sequence are extracted for features. The decoder module includes a unidirectional self-attention mechanism and a cross-attention mechanism, through which decoded features are obtained.
[0077] The neural network is trained using a two-stage training mode. The two-stage training mode first trains the neural network with crystal structure data to obtain an initial model, and then trains the model with nanoparticle data to obtain the neural network. The crystal structure data trains the neural network to learn the interaction of organic compounds, metal complexation, and crystal stacking patterns of complexes. The nanoparticle data trains the model to learn the regulation of assembly size by metal complexation and to learn the nanoparticle stacking pattern based on the crystal stacking pattern.
[0078] More specifically, prediction is made using a nanoparticle assembly model, which includes a first characterization module, an encoder, a decoder, and a prediction module.
[0079] For any compound molecule lacking conformational information, tools such as RDKIT and OpenBabel are used to generate the corresponding energy-minimizing conformation, thus transforming the 2D compound molecule into a 3D compound molecule. For 3D compound molecules, the SMILES syntax can be used to serialize the topological structure, forming a topological sequence in the form of c1ccccn1, where c, 1, c, c, c, c, n, and 1 are tokens in the sequence. Simultaneously, according to the order of atoms in the topology, their (x, y, z) three-dimensional coordinates are obtained, thus forming a molecular conformational sequence corresponding to the molecular topological sequence. For non-atom tokens such as 1, 2, and # that do not have three-dimensional coordinates in the topological sequence, their corresponding coordinates are assigned the value (0, 0, 0).
[0080] For the desired nanoparticle size, since size is a continuous variable, the continuous variable is first discretized and then transformed into a token that can be serialized.
[0081] Integrating the two inputs yields the integrated topological sequence and conformational sequence. The final topological sequence has the following format:
[0082] <sos> <size> … <mol> … <eos>
[0083] For ease of display, some content has been omitted here. <sos>This indicates the start of the entire topological sequence. <eos>This indicates the termination of the entire topological sequence. <size>This represents the desired nanoparticle size information, followed by a discretized size token. <mol>This represents the topological sequence of the target molecule.
[0084] In the final conformational sequence <sos> 、 <size> 、 <mol> 、 <eos>For tokens without three-dimensional coordinates, their corresponding coordinates are assigned the value (0,0,0). The rest are derived from the conformational sequence of the target molecule.
[0085] The function of the first characterization module is to receive inputs of compound molecules and nanoparticle sizes, process them, and then characterize them into features that the model can understand.
[0086] The encoder's role is to receive the output of the representation module and extract key features based on a bidirectional self-attention mechanism.
[0087] The encoder adopts a Transformer-like architecture, consisting of N layers of encoder modules. Each encoder module receives the output of the previous layer as input. The first-layer encoder module receives the output of the representation module as input. A bidirectional self-attention mechanism is used in all encoder modules.
[0088] The decoder's role is to receive the output of the encoder module and extract key features based on unidirectional self-attention and cross-attention mechanisms.
[0089] The decoder also adopts a Transformer-like architecture, consisting of N layers of decoder modules. Each layer's decoder module receives the output of the previous layer as input. The first layer's decoder module receives the encoder's output as input. Within the decoder module, a unidirectional self-attention mechanism and a cross-attention mechanism are employed; the unidirectional self-attention mechanism is performed first, followed by the cross-attention mechanism.
[0090] The prediction module receives the output of the decoder module, which can assemble the target compound into metal ions, organic compound additives, and the corresponding three-dimensional conformation of nanoparticles of a specified size.
[0091] The prediction module output consists of two sequences: a topological sequence and a conformational sequence. The topological sequence has the following format:
[0092] <sos> <metal> … <sol> … <mol> … <sos>
[0093] For ease of display, some content has been omitted here. <sos>This indicates the start of the entire topological sequence. <eos>This indicates the termination of the entire topological sequence. <metal>This indicates the amount of metal ions that need to be added to the nanoparticles, followed by a discretized metal ion token. <sol>This indicates the organic compound additives that need to be added to the nanoparticles, followed by an additive topology sequence token that follows the SMILES serialization format. <mol>This represents the topological sequence of the target compound molecule.
[0094] When outputting the conformational sequence, the following rule applies: for atoms in the topological sequence, output their three-dimensional coordinates; otherwise, output (0,0,0). <sos> 、 <metal> 、 <sol> 、 <mol> 、 <sos>All tokens are output as (0,0,0).
[0095] To train the nanoparticle assembly model, a two-stage training model was adopted: a pre-training truncation-fine-tuning phase. In the pre-training phase, 1.25 million crystal structure data sets from the Cambridge Structural Database (CSD) were used, enabling the nanoparticle assembly model to learn basic organic compound interactions, metal complexation, and complex crystal stacking patterns. In the fine-tuning phase, 100,000 nanoparticle assembly data sets from literature / patent reports were introduced, allowing the nanoparticle assembly model to learn the regulation of assembly size by metal complexation. Building upon the crystal stacking patterns learned in the pre-training model, the model then learned the nanoparticle stacking patterns.
[0096] The present invention also discloses a computer program product or system, including a computer program that, when executed by a processor, implements the above-described method steps.
[0097] Figure 2 A schematic diagram of the blood-brain barrier penetration efficacy prediction system for nanoparticles provided in this embodiment of the invention is shown, specifically including:
[0098] Acquisition module: Acquires the topological sequence, conformational sequence, and electron density point cloud of nanoparticles;
[0099] The representation module vectorizes the topological sequence and conformational sequence to obtain the topological sequence representation and conformational sequence representation. The electron density point cloud is used to extract features through a three-dimensional network model to obtain the electron density representation.
[0100] Fusion module: Fusion characterization is obtained by fusing topological sequence characterization, conformational sequence characterization and electronic density characterization;
[0101] Feature module: Inputs the fused representation into a neural network for feature extraction to obtain key features;
[0102] Assessment module: The blood-brain barrier penetration ability is assessed based on the key features to obtain high, medium, and low assessment results.
[0103] Figure 3 An embodiment of the present invention provides a schematic diagram of a computer device, specifically including:
[0104] A memory and a processor; the memory is used to store program instructions; the processor is used to invoke the program instructions, when the program instructions are executed any of the above-mentioned methods for predicting the blood-brain barrier penetration efficiency of nanoparticles.
[0105] The present invention also discloses a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, provides any of the above-described methods for predicting the blood-brain barrier penetration efficiency of nanoparticles.
[0106] The verification results of this verification embodiment show that assigning inherent weights to indications can improve the performance of this method compared to the default settings. Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, indirect coupling or communication connection of devices or units, and may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated; the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of this embodiment. Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0107] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0108] The computer device provided by the present invention has been described in detail above. For those skilled in the art, there will be changes in the specific implementation and application scope based on the ideas of the embodiments of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.< / sos> < / mol> < / sol> < / metal> < / sos> < / mol> < / sol> < / metal> < / eos> < / sos> < / sos> < / mol> < / sol> < / metal> < / sos> < / eos> < / mol> < / size> < / sos> < / mol> < / size> < / eos> < / sos> < / eos> < / mol> < / size> < / sos> < / sos> < / mol> < / sol> < / metal> < / sos> < / mol> < / sol> < / metal> < / eos> < / sos> < / sos> < / mol> < / sol> < / metal> < / sos>
Claims
1. A method for predicting blood brain barrier penetration efficiency of nanoparticles, characterized by, include: The topological sequence, conformational sequence, and electron density point cloud of nanoparticles are obtained; the topological sequence of the nanoparticles includes the topological sequence of metal ions added to the nanoparticles, organic compound additives to the nanoparticles, and target compound molecules; the conformational sequence of the nanoparticles includes the three-dimensional coordinates of atoms and non-atoms in the topological sequence. The topological sequence and conformational sequence are vectorized and encoded to obtain the topological sequence representation and the conformational sequence representation, and the electron density point cloud is obtained by feature extraction through a three-dimensional network model to obtain the electron density representation. The fused characterization is obtained by integrating topological sequence characterization, conformational sequence characterization and electronic density characterization. The fused representation is fed into a neural network for feature extraction to obtain key features; The blood-brain barrier penetration ability was assessed using the aforementioned key features, resulting in high, medium, and low assessment results.
2. The method of predicting blood brain barrier penetration efficacy of nanoparticles according to claim 1, wherein, The nanoparticles include one or more of different types of nanoparticles such as small molecules, peptides, and RNA. The electron density point cloud is voxelized to obtain the electron cloud density. The electron cloud density is characterized by cross-scale feature extraction through a three-dimensional network model. The three-dimensional network model includes any one or more of the following: 3D CNN, 3D ResNet, 3DDenseNet, 3D MobileNet, and Swin3D.
3. The method of predicting blood brain barrier penetration efficacy of nanoparticles according to claim 1, wherein, The neural network is any one or more of the following discriminator models: Swing Transformer Discriminator, TransGAN Discriminator, DeiT Discriminator, CSWin Transformer Discriminator, BEiT Discriminator; the fused representation is input to the discriminator model for feature extraction to obtain key features.
4. The method of predicting blood brain barrier penetration efficacy of nanoparticles according to claim 3, wherein, The discriminator model consists of N discriminator modules, where N is a natural number greater than 1. Each discriminator module is a model constructed using a bidirectional self-attention mechanism. The fusion representation extracts key features through the N-layer bidirectional self-attention mechanism.
5. The method for predicting the blood-brain barrier penetration efficiency of nanoparticles according to claim 4, characterized in that, The discriminator model is trained by acquiring electron cloud density data, topological data, and conformational data of small molecules, peptides, and small RNA nanoparticles, using the experimental efficiency of small molecules, peptides, and small RNA nanoparticles in penetrating the blood-brain barrier as a label. The electron cloud density data, topological data, conformational data, and labels of small molecules, peptides, and small RNA nanoparticles are input into the discriminator model to be trained for training until the loss function remains unchanged, thus obtaining the discriminator model.
6. The method of predicting blood brain barrier penetration efficacy of nanoparticles according to claim 1, wherein, Each element in the topological sequence represents a token. The atoms include metal ion tokens added to the nanoparticles, organic compound additive tokens for the nanoparticles, and topological sequence tokens for the target compound molecule. The non-atoms include the start token, end token, metal ion symbol token, organic compound additive symbol token, and topological sequence symbol token for the target compound molecule. The three-dimensional coordinates of the atoms are obtained according to the structure of the topological sequence, and the three-dimensional coordinates of the non-atoms are assigned a value of zero.
7. A computer program product comprising a computer program or instructions embodied thereon, characterized in that, The computer program or instructions are executed by a processor to implement the method for predicting the blood-brain barrier penetration efficiency of nanoparticles according to any one of claims 1-6.
8. A computer device comprising a memory, a processor, and a computer program or instructions stored on the memory, wherein, The computer program or instructions are executed by a processor to implement the method for predicting the blood-brain barrier penetration efficiency of nanoparticles according to any one of claims 1-6.
9. A computer readable storage medium having stored thereon a computer program or instructions, characterized in that, The computer program or instructions are executed by a processor to implement the method for predicting the blood-brain barrier penetration efficiency of nanoparticles according to any one of claims 1-6.
Citation Information
Patent Citations
Enhancement of transport of therapeutic molecules across the blood brain barrier
CN104159922A
Blood-brain barrier crossing antibodies
CN121358771A