A method, apparatus, and program product for predicting nanoparticle assembly
By simulating nanoparticle assembly schemes using generative artificial intelligence models, the problems of complex factor coupling and low efficiency of manual screening in nanomedicine development have been solved. This has enabled efficient and reliable nanomedicine assembly and blood-brain barrier penetration prediction, thereby improving the success rate of clinical translation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING NEUROSURGICAL INST
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-12
AI Technical Summary
Existing technologies for developing nanomedicines that can penetrate the blood-brain barrier suffer from problems such as complex factor coupling, low efficiency due to reliance on human screening, high costs and long cycles, and low success rate in clinical translation.
Generative artificial intelligence is used to learn about the stacking of nanoparticles. A neural network model is used to virtually simulate nanoparticle assembly schemes, generate metal ions and organic compound additives of specified sizes, and combine the prediction of blood-brain barrier penetration ability to achieve intelligent recommendation of compound combinations and assembly schemes.
It reduces trial-and-error costs, improves the reliability and efficiency of nanomedicine assembly, shortens the development cycle, and increases the success rate of clinical translation.
Smart Images

Figure CN121905345B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of drug delivery, specifically to a method, apparatus, program product, and computer-readable storage medium for predicting nanoparticle assembly. Background Technology
[0002] In the development of nanomedicines that penetrate the blood-brain barrier, traditional methods primarily rely on the experience-driven model of pharmaceutical research experts. Researchers, based on their understanding of the physiological characteristics of the blood-brain barrier, use a trial-and-error approach to screen combinations of metal ions and organic compound additives to optimize the physicochemical properties of nanoparticles. Existing research has clearly established that the particle size (10-100 nm), shape (aspect ratio approximately 2-5), surface charge (near-neutral), and surface functionalization modifications of nanomedicines are key parameters affecting their cross-BBB transport efficiency. Inorganic nanoparticles (such as metal-based, iron oxide, and carbon-based nanoparticles) have been extensively studied due to their unique magnetic and optical properties, but their surface modification strategies still require repeated optimization to balance BBB penetration ability with the risk of immune clearance. The limitations of this traditional trial-and-error model are mainly reflected in: 1) Multi-factor coupling leading to experimental space explosion. The factors affecting the brain delivery efficiency of nanomedicines are extremely complex, including numerous variables such as framework materials, particle size distribution, surface charge, target ligand structure, dosage, and disease models, and there are complex interactions between these factors. 1) The formulation development model relying on manual screening cannot systematically examine multiple factors simultaneously, resulting in low efficiency and extremely limited scope of exploration. 2) Subjectivity and non-replicability of expert experience. The selection of metal ions and organic compounds is highly dependent on the personal experience of researchers. The coordination behavior of specific metal ions (such as copper, iron, and manganese), the lipophilic modification of organic ligands, and the coupling strategies of target molecules (such as transferrin and angiopep-2 peptide) often originate from scattered literature inspiration or accidental experimental discoveries, lacking systematic design guidelines. 3) High cost and long cycle of trial and error. Each formulation adjustment requires a series of steps, including nanoparticle synthesis, characterization, in vitro BBB model screening, and in vivo validation in small animals, with a single experimental cycle lasting several weeks to months. Even if candidate combinations are successfully screened, they still face the risk of a significant drop in success rate during the clinical translation stage—the success rate of nanomedicines from Phase II to Phase III clinical trials is only 14%. Summary of the Invention
[0003] To address the above problems, this invention provides a method for predicting nanoparticle assembly, specifically including:
[0004] Obtain the three-dimensional compound molecules and the size of the target nanoparticles;
[0005] The topological sequence is obtained by serializing the topological structure of a three-dimensional compound molecule.
[0006] The three-dimensional coordinates of each atom are obtained by the order of atoms in the topological structure, thus obtaining the molecular conformation sequence;
[0007] The size sequence of the target nanoparticles is obtained by discretizing the size.
[0008] The first topological sequence is obtained by integrating the topological sequence and the size sequence, and the first conformational sequence is obtained by integrating the molecular conformational sequence and the size sequence.
[0009] The first topological sequence and the first conformation sequence are input into a neural network for feature extraction and prediction to obtain the metal ions, organic compound additives and corresponding three-dimensional conformations of nanoparticles of a specified size.
[0010] Optionally, the prediction yields a predicted topological sequence and a predicted conformational sequence. The predicted topological sequence includes the metal ions and organic compound additives required for nanoparticles of a specified size. The predicted conformational sequence includes the three-dimensional coordinates of atoms in the predicted topological sequence, and the corresponding three-dimensional conformation is obtained through the three-dimensional coordinates.
[0011] Optionally, each element in the topological sequence is a token in the sequence. The three-dimensional coordinates of each atom are obtained by the order of the atomic tokens in the topological structure. The three-dimensional coordinates of non-atomic tokens in the topological sequence are assigned a value of zero. The molecular conformation sequence is obtained by the three-dimensional coordinates of atomic tokens and non-atomic tokens in the topological sequence.
[0012] Optionally, the neural network includes any one or more of the following: Transformer, residual network, Mamba, Mamba-2, ERNIE, XLNet, and T5.
[0013] Optionally, the neural network includes an encoder and a decoder. The encoder includes an N-layer encoder module, and the decoder includes an N-layer decoder module, where N is a natural number greater than 1. The first topological sequence and the first conformation sequence are input into the first-layer encoder module and encoded by the N-layer encoder module to obtain encoded features. The encoded features are then decoded by the N-layer decoder module to obtain decoded features. The metal ions, organic compound additives, and corresponding three-dimensional conformations of nanoparticles of a specified size are predicted using the decoded features.
[0014] Optionally, the encoder module includes a bidirectional self-attention mechanism, through which the first topological sequence and the first conformation sequence are used for feature extraction. The decoder module includes a unidirectional self-attention mechanism and a cross-attention mechanism, through which decoded features are obtained.
[0015] Optionally, the neural network is trained using a two-stage training mode. The two-stage training mode first trains the neural network using crystal structure data to obtain an initial model, and then trains the model using nanoparticle data to obtain the neural network. The crystal structure data trains the neural network to learn the interaction of organic compounds, metal complexation, and crystal stacking patterns of complexes. The nanoparticle data trains the model to learn the regulation of assembly size by metal complexation and to learn the nanoparticle stacking pattern based on the crystal stacking pattern.
[0016] The purpose of this invention is to provide a computer program product that includes a computer program or instructions, which are executed by a processor to implement the above-described method for predicting nanoparticle assembly.
[0017] The purpose of this invention is to provide a computer device comprising a memory, a processor, and a computer program or instructions stored in the memory, wherein the computer program or instructions are executed by the processor to implement the above-described method for predicting nanoparticle assembly.
[0018] The purpose of this invention is to provide a computer-readable storage medium having a computer program or instructions stored thereon, which are executed by a processor to implement the above-described method for predicting nanoparticle assembly.
[0019] Advantages of this invention:
[0020] 1. The factors influencing the efficiency of nanomedicine delivery to the brain are extremely complex. Traditional formulation development models relying on manual screening cannot systematically examine multiple factors simultaneously, resulting in low efficiency and a very limited scope of exploration. This invention proposes using generative artificial intelligence to learn about the stacking knowledge of nanoparticles. Through modeling, it virtually simulates the assembly schemes between nanoparticle drugs, enabling intelligent recommendations for additive combinations and assembly schemes. Specifically, the topological and molecular conformational sequences of compound molecules are designed. The generative model simulates and extrapolates the specified nanoparticle size, topological sequence, and molecular conformational sequence, generating metal ions of specified sizes and the required organic compound additives, along with their corresponding three-dimensional conformations.
[0021] 2. Regarding the predicted penetration capability of nanoparticles of a specified size assembled on the blood-brain barrier, this invention constructs a blood-brain barrier penetration capability judgment method. The topological and conformational sequences of the predicted compound are used as inputs, and the electron cloud density of the compound is calculated. By judging and predicting the blood-brain barrier penetration capability of the compound through the topological sequence, conformational sequence, and electron cloud density, the blood-brain barrier penetration capability of the compound can be obtained. This can effectively reduce the trial and error cost and provide reliability and feasibility for the assembled nanoparticles. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A schematic diagram of the nanoparticle assembly prediction method provided in an embodiment of the present invention;
[0024] Figure 2 A schematic diagram of a nanoparticle assembly prediction system provided in an embodiment of the present invention;
[0025] Figure 3 A schematic diagram of a computer device provided in an embodiment of the present invention;
[0026] Figure 4 This is a flowchart of a nanoparticle assembly model provided in an embodiment of the present invention;
[0027] Figure 5 The training process for assembling nanoparticles is provided for embodiments of the present invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0029] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as S101, S102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0030] Figure 1 A schematic diagram of the nanoparticle assembly prediction method provided in this embodiment of the invention is shown, specifically including:
[0031] S1: Obtain the three-dimensional compound molecule and target nanoparticle size of the target compound;
[0032] In one embodiment, the target compound includes any one or more of small molecules, peptides, and small RNAs.
[0033] In one specific embodiment, for any compound molecule that does not have conformational information, tools such as RDKIT and OpenBabel are used to generate the corresponding energy-minimizing conformation, thereby transforming the 2D compound molecule into a 3D compound molecule.
[0034] S2: Sequencing the topological structure of a three-dimensional compound molecule to obtain a topological sequence;
[0035] In one specific embodiment, for 3D compound molecules, the SMILES syntax can be used to serialize the topological structure to form a topological sequence, such as c1ccccn1, where c, 1, c, c, c, c, n, and 1 are all tokens in the sequence.
[0036] S3: Obtain the three-dimensional coordinates of each atom by the order of atoms in the topological structure to obtain the molecular conformation sequence;
[0037] In one embodiment, each element in the topological sequence is a token in the sequence. The three-dimensional coordinates of each atom are obtained by the order of the atomic tokens in the topological structure. The three-dimensional coordinates of non-atomic tokens in the topological sequence are assigned a value of zero. The molecular conformation sequence is obtained by the three-dimensional coordinates of atomic tokens and non-atomic tokens in the topological sequence.
[0038] In one specific embodiment, the (x,y,z) three-dimensional coordinates of the atoms in the topology are obtained according to their order, thereby forming a molecular conformation sequence corresponding to the molecular topological sequence. For non-atom tokens such as 1, 2, and # that do not have three-dimensional coordinates in the topological sequence, their corresponding coordinates are assigned the value (0,0,0).
[0039] S4: Discretize the size of the target nanoparticles to obtain a size sequence;
[0040] In one specific embodiment, for the desired nanoparticle size, since size is a continuous variable, the continuous variable is first discretized and then transformed into a token that can be serialized.
[0041] S5: Integrate the topological sequence and the size sequence to obtain the first topological sequence, and integrate the molecular conformation sequence and the size sequence to obtain the first conformation sequence;
[0042] In one specific embodiment, the two inputs described above are integrated to obtain an integrated topological sequence and a conformational sequence. The final topological sequence has the following format:
[0043] <sos> <size> … <mol> … <eos>
[0044] For ease of display, some content has been omitted here. <sos>This indicates the start of the entire topological sequence. <eos>This indicates the termination of the entire topological sequence. <size>This represents the desired nanoparticle size information, followed by a discretized size token. <mol>This represents the topological sequence of the target molecule.
[0045] In the final conformational sequence <sos> 、 <size> 、 <mol> 、 <eos>For tokens without three-dimensional coordinates, their corresponding coordinates are assigned the value (0,0,0). The rest are derived from the conformational sequence of the target molecule.
[0046] In one embodiment, the three-dimensional coordinates of non-atomic elements in the molecular conformation sequence are assigned a value of zero. Non-atomic elements include non-atomic tokens such as 1, 2, and # in the topological sequence, as well as symbols representing the start of the sequence. <sos>), sequence termination symbol ( <eos>), size sequence symbols ( <size>), topological sequence symbols of the target molecule ( <mol>).
[0047] For the final topological and conformational sequences, an embedded approach is used to represent them as features that can be understood by the AI model.
[0048] S6: Input the first topological sequence and the first conformation sequence into the neural network for feature extraction and prediction to obtain the metal ions, organic compound additives and corresponding three-dimensional conformations of nanoparticles of a specified size.
[0049] In one embodiment, the prediction yields a predicted topological sequence and a predicted conformational sequence. The predicted topological sequence includes the metal ions and organic compound additives required to be added for nanoparticles of a specified size. The predicted conformational sequence includes the three-dimensional coordinates of atoms in the predicted topological sequence, and the corresponding three-dimensional conformation is obtained through the three-dimensional coordinates.
[0050] In one embodiment, the neural network includes any one or more of the following: Transformer, Residual Network, Mamba, Mamba-2, ERNIE, XLNet, and T5.
[0051] In one embodiment, the neural network includes an encoder and a decoder. The encoder includes an N-layer encoder module, and the decoder includes an N-layer decoder module, where N is a natural number greater than 1. The first topological sequence and the first conformation sequence are input into the first-layer encoder module and encoded by the N-layer encoder module to obtain encoded features. The encoded features are then decoded by the N-layer decoder module to obtain decoded features. The decoded features are used to predict the metal ions, organic compound additives, and corresponding three-dimensional conformations of nanoparticles of a specified size.
[0052] The encoder module includes a bidirectional self-attention mechanism, through which the first topological sequence and the first conformation sequence are extracted for features. The decoder module includes a unidirectional self-attention mechanism and a cross-attention mechanism, through which decoded features are obtained.
[0053] The neural network is trained using a two-stage training mode. The two-stage training mode first trains the neural network with crystal structure data to obtain an initial model, and then trains the model with nanoparticle data to obtain the neural network. The crystal structure data trains the neural network to learn the interaction of organic compounds, metal complexation, and crystal stacking patterns of complexes. The nanoparticle data trains the model to learn the regulation of assembly size by metal complexation and to learn the nanoparticle stacking pattern based on the crystal stacking pattern.
[0054] In one specific embodiment, the nano-assembly model is as follows: Figure 4 As shown, it includes a characterization module, an encoder, a decoder, and a nanoparticle prediction module. The first topological sequence and the first conformational sequence of the target compound are input to the characterization module for vectorization encoding to obtain the first topological sequence characterization and the first conformational sequence characterization, and then input to the encoder for encoding, the decoder for decoding, and the nanoparticle prediction module for prediction.
[0055] The encoder's role is to receive the output of the representation module and extract key features based on a bidirectional self-attention mechanism.
[0056] The encoder adopts a Transformer-like architecture, consisting of N layers of encoder modules. Each encoder module receives the output of the previous layer as input. The first-layer encoder module receives the output of the representation module as input. A bidirectional self-attention mechanism is used in all encoder modules.
[0057] The decoder's role is to receive the output of the encoder module and extract key features based on unidirectional self-attention and cross-attention mechanisms.
[0058] The decoder also adopts a Transformer-like architecture, consisting of N layers of decoder modules. Each layer's decoder module receives the output of the previous layer as input. The first layer's decoder module receives the encoder's output as input. Within the decoder module, a unidirectional self-attention mechanism and a cross-attention mechanism are employed; the unidirectional self-attention mechanism is performed first, followed by the cross-attention mechanism.
[0059] The bidirectional self-attention mechanism includes a bidirectional attention layer, a residual connection and a layer normalization layer, a feedforward neural network, a residual connection and a layer normalization layer, and its output is passed to the next layer of the bidirectional self-attention mechanism (encoder module).
[0060] Bidirectional self-attention layer: Used to capture the contextual relationships between all tokens in the input sequence. First "addition and normalization" layer: This is a residual connection followed by layer normalization. Specifically: layer normalization (output of the self-attention layer + input of this sub-layer), which helps stabilize training and gradient flow. Feedforward neural network layer: This is a fully connected feedforward network, typically containing two linear transformations and an intermediate non-linear activation function (such as ReLU or GELU) to independently transform the representation at each position. Second "addition and normalization" layer: Also a residual connection followed by layer normalization: layer normalization (output of the feedforward network layer + output of the first "addition and normalization" layer). Different numbers of residual connections and layer normalization layers can be constructed in the encoder and decoder modules as needed.
[0061] The decoder module consists of three core sub-layers, connected in sequence:
[0062] Masked One-Way Self-Attention Layer: This is a one-way self-attention mechanism. It uses a mask to ensure that when generating the current word, only the positional information before it is seen is visible, preventing information leakage.
[0063] "Addition and Normalization" layer: Residual connections and layer normalization are used to process the output of masked self-attention.
[0064] Cross-attention layer: Its key and value come from the final output of the encoder stack, while the query comes from the output of the previous "addition and normalization" layer. This allows the decoder to selectively focus on all the information from the encoder input when generating each lexical unit.
[0065] "Addition and Normalization" layer: handles the output of cross-attention.
[0066] Feedforward neural network layer: It functions the same as the feedforward network in the encoder, performing a non-linear transformation at each position.
[0067] "Addition and Normalization" layer: processes the output of the feedforward network.
[0068] The output is passed to the next decoder module or the final output layer.
[0069] The model's loss function is the cross-entropy function, the optimizer during training is Adam / AdamW, and the initial learning rate is 1e-4 ~ 1e-3.
[0070] The role of the nanoparticle prediction module is to receive the output of the decoder module, which can assemble the target compound into nanoparticles of a specified size, including metal ions, organic compound additives, and the corresponding three-dimensional conformation.
[0071] The prediction module output consists of two sequences: a topological sequence and a conformational sequence. The topological sequence has the following format:
[0072] <sos> <metal> … <sol> … <mol> … <sos>
[0073] For ease of display, some content has been omitted here. <sos>This indicates the start of the entire topological sequence. <eos>This indicates the termination of the entire topological sequence. <metal>This indicates the amount of metal ions that need to be added to the nanoparticles, followed by a discretized metal ion token. <sol>This indicates the organic compound additives that need to be added to the nanoparticles, followed by an additive topology sequence token that follows the SMILES serialization format. <mol>This represents the topological sequence of the target compound molecule.
[0074] When outputting the conformational sequence, the following rule applies: for atoms in the topological sequence, output their three-dimensional coordinates; otherwise, output (0,0,0). <sos> 、 <metal> 、 <sol> 、 <mol> 、 <sos>All tokens are output as (0,0,0).
[0075] In one specific embodiment, to train the nanoparticle assembly model, a two-stage training mode is adopted, namely a pre-training cutoff-fine-tuning stage, such as... Figure 5 As shown, after the first stage of pre-training using CSD data, the model is then trained in the second stage using patent / literature data related to nanoparticles. The model parameters are adjusted to obtain the final nanoparticle assembly model.
[0076] The purpose of the pre-training phase is to more fully initialize the model parameters based on crystal data in order to obtain better model performance.
[0077] During the pre-training phase, 1.25 million crystal structure data sets from the Cambridge Structural Database (CSD) were used to enable the nanoparticle assembly model to learn basic patterns of organic compound interactions, metal complexation, and complex crystal stacking.
[0078] The purpose of the fine-tuning phase is to generalize knowledge from crystals to nanoparticles and further adjust the model parameters specifically for nanoparticle data.
[0079] During the fine-tuning phase, nanoparticle assembly data from 100,000 literature / patent reports were introduced, enabling the nanoparticle assembly model to learn the regulation of assembly size by metal complexes. Based on the crystal stacking mode learned by the pre-trained model, the nanoparticle stacking mode was learned.
[0080] In one specific embodiment, the method further includes predicting the blood-brain barrier penetration capability of nanoparticles. This involves using the metal ions, organic compound additives (topological sequence), and corresponding three-dimensional conformational sequence (conformation sequence) of nanoparticles of a specified size, predicted by a nanoassembly model, as input to calculate the penetration capability of the nanoparticles of that specified size in the blood-brain barrier; specifically:
[0081] Obtain the topological sequence and conformational sequence of nanoparticles of a specified size, and calculate the electron density point cloud of the nanoparticles;
[0082] The topological sequence and conformational sequence are vectorized and encoded to obtain the topological sequence representation and the conformational sequence representation, and the electron density point cloud is obtained by feature extraction through a three-dimensional network model to obtain the electron density representation.
[0083] The fused characterization is obtained by integrating topological sequence characterization, conformational sequence characterization and electronic density characterization.
[0084] The fused representation is fed into the first neural network for feature extraction to obtain key features;
[0085] The blood-brain barrier penetration ability was assessed using the aforementioned key features, resulting in high, medium, and low assessment results.
[0086] The electron density point cloud is obtained by voxelizing the point cloud data of the nanoparticles to calculate the electron cloud density of the nanoparticles.
[0087] Optionally, the fusion is achieved by fusing topological sequence representation, conformational sequence representation, and electronic density representation through cross-attention to obtain the fused representation. The cross-attention fusion process of topology and conformation is as follows:
[0088] Using topological representation as the query and conformational representation as the key and value, the dot product of the query and all keys is calculated and normalized using the Softmax function to obtain the attention score matrix. The attention score A is then used to perform a weighted summation of the value V to obtain the new representation after the topological modality has perceived the conformational information.
[0089] Symmetrically, the present invention performs topological and conformational interactions, topological and electronic interactions, and conformational and electronic interactions. After completing the above-mentioned multi-order cross-attention interactions, six interactive representations (one for each direction) are obtained. By splicing or weighted summation, all interactive representations are fused with the original modal representations to obtain a fused representation.
[0090] The interaction between topology and conformation yields new representations of topological modality-sensing conformational information and new representations of conformational modality-sensing topological information, which solves the problem of whether the connection mode of flexible molecules under a specific three-dimensional shape is stable.
[0091] The new characterization of topological mode sensing electron density information and the new characterization of electron density mode sensing topological information obtained by topological and electronic interaction solves the problem of whether a specific atomic connection mode (such as a coordinating group) has a suitable electron density to bind metal ions.
[0092] The interaction between conformation and electron yields new representations of conformation modality-sensing electron density information, new representations of electron density modality-sensing conformation information, new representations of topological modality-sensing electron density information, and new representations of electron density modality-sensing topological information.
[0093] Optionally, during the training of the 3D model and the first neural network, the nanoparticles include multi-resolution electron density point clouds transformed from different types of nanoparticles such as small molecules, peptides, and RNA. The electron density point clouds are voxelized to obtain electron cloud density. The cross-scale electron cloud density is input into the 3D network model. After extracting key features by combining topological sequences and conformational sequences, the BBB penetration assessment is predicted. The real labels and predicted labels are compared, and iterative training is performed to obtain the model.
[0094] In application, the electron cloud density of nanoparticles is input into a three-dimensional network model for cross-scale feature extraction to obtain electron density characterization; the three-dimensional network model includes any one or more of the following: 3D CNN, 3D ResNet, 3DDenseNet, 3D MobileNet, and Swin3D.
[0095] The first neural network is any one or more of the following discriminator models: Swing TransformerDiscriminator, TransGAN Discriminator, DeiT Discriminator, CSWin TransformerDiscriminator, BEiT Discriminator; the fused representation is input to the discriminator model for feature extraction to obtain key features.
[0096] Furthermore, the discriminator model consists of N layers of discriminator modules, where N is a natural number greater than 1. Each discriminator module is a model constructed using a bidirectional self-attention mechanism. The fusion representation extracts key features through the N-layer bidirectional self-attention mechanism.
[0097] The training of the discriminator model involves: acquiring electron cloud density data, topological data, and conformational data of small molecules, peptides, and small RNA nanoparticles; using the experimental efficiency of small molecules, peptides, and small RNA nanoparticles in penetrating the blood-brain barrier as a label; and inputting the electron cloud density data, topological data, conformational data, and labels of small molecules, peptides, and small RNA nanoparticles into the discriminator model to be trained until the loss function remains unchanged, thus obtaining the discriminator model.
[0098] The topological sequence of the nanoparticles includes the topological sequence of the metal ions added to the nanoparticles, the organic compound additives to the nanoparticles, and the target compound molecules; the conformational sequence of the nanoparticles includes the three-dimensional coordinates of atoms and non-atoms in the topological sequence.
[0099] Specifically:
[0100] The prediction of blood-brain barrier penetration by nanoparticles is divided into a representation module, a discriminator, and a blood-brain barrier penetration discrimination module. The representation module receives inputs of nanoparticles and their electron density, processes them, and then represents them as features that the model can understand. The representation module consists of an embedded vector module and a 3D CNN module. The embedded vector module is used to represent the serialized topological and conformational sequences of the nanoparticles, while the 3D CNN module is used to represent the electron density of the nanoparticles. The electron density of the nanoparticles is represented in the form of a point cloud. After voxelization, it is extracted by the 3D CNN module to form the electron density representation.
[0101] Finally, the topological sequence embedded characterization, conformational sequence embedded characterization, and electronic density characterization of nanoparticles are integrated (such as feature splicing operations) to form features that can be understood by the AI model.
[0102] The discriminator adopts a Transformer-like architecture, consisting of N layers of discriminator modules. Each layer's discriminator module receives the output of the layer above as input. The first layer's discriminator module receives the output of the representation module as input. A bidirectional self-attention mechanism is used in all discriminator modules.
[0103] The function of the nanoparticle BBB penetration discrimination module is to receive the output of the discriminator module and output the BBB penetration capability label of the nanoparticles, classifying them into three categories: high, medium, and low.
[0104] Here, the present invention divides the BBB penetration rate, which is a continuous variable, into three categorical variables (high, medium, and low) according to the threshold in the actual drug research scenario. This transforms the regression problem into a classification problem, reduces the training difficulty of the discrimination model, and thus improves the accuracy of the discrimination model.
[0105] This invention treats data composed of different types of nanoparticles (small molecules, peptides, small RNA) as belonging to the same type from the perspective of electron density characterization, thus alleviating the problem of limited data transmission for single-type BBB. Based on the unified characterization of electron density, different types of nanoparticles such as small molecules, peptides, and RNA are uniformly transformed into multi-resolution electron density point clouds. These point clouds are then voxelized, and cross-scale features are extracted using three-dimensional convolutional kernels in a 3D CNN module, solving the prediction bias problem caused by data type differences in traditional methods.
[0106] The two-stage cascaded model of nanoparticle assembly and nanoparticle BBB penetration prediction proposed in this invention provides a technical solution that differs from existing technologies in terms of improving additive screening efficiency, prediction accuracy, and data utilization efficiency. It is a novel technical concept with great application value.
[0107] The present invention also discloses a computer program product or system, including a computer program that, when executed by a processor, implements the above-described method steps.
[0108] Figure 2 A schematic diagram of the nanoparticle assembly prediction system provided in this embodiment of the invention specifically includes:
[0109] Acquisition Unit: Acquires the three-dimensional compound molecules and the size of the target nanoparticles;
[0110] Topological unit: A topological sequence is obtained by serializing the topological structure of a three-dimensional compound molecule;
[0111] Conformational unit: The three-dimensional coordinates of each atom are obtained by the order of atoms in the topological structure, thus obtaining the molecular conformation sequence;
[0112] Size unit: The size sequence is obtained by discretizing the size of the target nanoparticles;
[0113] Integration unit: Integrating the topological sequence and the size sequence to obtain the first topological sequence, and integrating the molecular conformation sequence and the size sequence to obtain the first conformation sequence;
[0114] Prediction unit: Input the first topological sequence and the first conformation sequence into the neural network for feature extraction and then perform prediction to obtain the metal ions, organic compound additives and corresponding three-dimensional conformations of nanoparticles of a specified size.
[0115] Figure 3 An embodiment of the present invention provides a schematic diagram of a computer device, specifically including:
[0116] A memory and a processor; the memory is used to store program instructions; the processor is used to invoke the program instructions when any of the above-described methods for predicting nanoparticle assembly are executed.
[0117] The present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, provides any of the above-described methods for predicting nanoparticle assembly.
[0118] The verification results of this verification embodiment show that assigning inherent weights to indications can improve the performance of this method compared to the default settings. Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, indirect coupling or communication connection of devices or units, and may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated; the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of this embodiment. Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0119] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0120] The computer device provided by the present invention has been described in detail above. For those skilled in the art, there will be changes in the specific implementation and application scope based on the ideas of the embodiments of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.< / sos> < / mol> < / sol> < / metal> < / sos> < / mol> < / sol> < / metal> < / eos> < / sos> < / sos> < / mol> < / sol> < / metal> < / sos> < / mol> < / size> < / eos> < / sos> < / eos> < / mol> < / size> < / sos> < / mol> < / size> < / eos> < / sos> < / eos> < / mol> < / size> < / sos>
Claims
1. A method for predicting nanoparticle assembly, characterized in that, include: Obtain the three-dimensional compound molecules and the size of the target nanoparticles; The topological sequence is obtained by serializing the molecular topological structure of a three-dimensional compound. The three-dimensional coordinates of each atom are obtained by the order of atoms in the topological structure, thus obtaining the molecular conformation sequence; The size sequence of the target nanoparticles is obtained by discretizing the size. The first topological sequence is obtained by integrating the topological sequence and the size sequence, and the first conformational sequence is obtained by integrating the molecular conformational sequence and the size sequence. The first topological sequence and the first conformation sequence are input into a neural network for feature extraction and prediction to obtain the metal ions, organic compound additives and corresponding three-dimensional conformations of nanoparticles of a specified size. The neural network is trained using a two-stage training mode. The two-stage training mode first trains the neural network with crystal structure data to obtain an initial model, and then trains the model with nanoparticle data to obtain the neural network. The crystal structure data trains the neural network to learn the interaction of organic compounds, metal complexation, and crystal stacking patterns of complexes. The nanoparticle data trains the model to learn the regulation of assembly size by metal complexation and to learn the nanoparticle stacking pattern based on the crystal stacking pattern.
2. The method for predicting nanoparticle assembly according to claim 1, characterized in that, The prediction yields a predicted topological sequence and a predicted conformational sequence. The predicted topological sequence includes the metal ions and organic compound additives required for nanoparticles of a specified size. The predicted conformational sequence includes the three-dimensional coordinates of atoms in the predicted topological sequence, and the corresponding three-dimensional conformation is obtained through the three-dimensional coordinates.
3. The method for predicting nanoparticle assembly according to claim 1, characterized in that, Each element in the topological sequence is a token in the sequence. The three-dimensional coordinates of each atom are obtained by the order of the atomic tokens in the topological structure. The three-dimensional coordinates of non-atomic tokens in the topological sequence are assigned to zero. The molecular conformation sequence is obtained by the three-dimensional coordinates of atomic tokens and non-atomic tokens in the topological sequence.
4. The method for predicting nanoparticle assembly according to claim 1, characterized in that, The neural network includes any one or more of the following: Transformer, Residual Network, Mamba, Mamba-2, ERNIE, XLNet, and T5.
5. The method for predicting nanoparticle assembly according to claim 1, characterized in that, The neural network includes an encoder and a decoder. The encoder includes an N-layer encoder module, and the decoder includes an N-layer decoder module, where N is a natural number greater than 1. The first topological sequence and the first conformation sequence are input into the first-layer encoder module and encoded by the N-layer encoder module to obtain encoded features. The encoded features are decoded by the N-layer decoder module to obtain decoded features. The decoded features are used to predict the metal ions, organic compound additives, and corresponding three-dimensional conformations of nanoparticles of a specified size.
6. The method for predicting nanoparticle assembly according to claim 5, characterized in that, The encoder module includes a bidirectional self-attention mechanism, through which the first topological sequence and the first conformation sequence are extracted for features. The decoder module includes a unidirectional self-attention mechanism and a cross-attention mechanism, through which decoded features are obtained.
7. A computer program product comprising a computer program or instructions, characterized in that, The computer program or instructions are executed by a processor to implement the method for predicting nanoparticle assembly as described in any one of claims 1-6.
8. A computer device comprising a memory, a processor, and a computer program or instructions stored in the memory, characterized in that, The computer program or instructions are executed by a processor to implement the method for predicting nanoparticle assembly as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instructions are executed by a processor to implement the method for predicting nanoparticle assembly as described in any one of claims 1-6.
Citation Information
Patent Citations
Drug target affinity prediction method and system based on multi-scale protein attention mechanism
CN120783870A
Method and apparatus for determining drug molecule property, and storage medium
US20220415452A1