A quasi-smiles and recurrent neural network based nanoparticle inverse design method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU UNIVERSITY
- Filing Date
- 2022-10-10
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies struggle to quickly and cost-effectively predict the hydrophilicity/hydrophobicity and cellular uptake properties of nanoparticles, especially given the vast variety and number of nanomaterials. Experimental methods are time-consuming, labor-intensive, and dependent on equipment and technical expertise.
A nanoparticle reverse design method based on Quasi-SMILES and recurrent neural networks is adopted. By constructing generative and predictive models, and utilizing nanoparticle structural information in Quasi-SMILES format, combined with recurrent neural networks and random forest algorithms, the hydrophilicity/hydrophobicity and cellular uptake of nanoparticles can be predicted.
A predictive model covering a variety of nanoparticles was constructed. It is simple, fast, low-cost, and easy to operate, conforming to the development principles of OECD QSAR models. It provides data support and theoretical guidance for the rational design of nanomaterials, avoiding blind experimentation.
Smart Images

Figure CN115547434B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the research field of artificial intelligence-assisted prediction and design of nanoparticle properties, specifically a nanoparticle reverse design method based on Quasi-SMILES and recurrent neural networks. Background Technology
[0002] The hydrophilicity or hydrophobicity of nanoparticles reflects their hydrophilic or lipophilic properties, typically expressed as the oil-water partition coefficient (logP). Its value is highly correlated with numerous biological effects, including cellular uptake, in vivo distribution, protein adsorption, cytotoxicity, immune response, and pharmacokinetics. Hydrophilicity or hydrophobicity can be considered a crucial prerequisite and indicator for evaluating the biological effects of nanomaterials. Due to their small size and large specific surface area, nanoparticles are prime candidates for drug delivery systems. They can effectively deliver drugs or imaging agents to lesion sites for disease diagnosis and treatment. Furthermore, the uptake and retention of nanoparticles by various organs directly impacts their safety in the environment. Nanoparticles primarily enter cells through endocytosis and infiltration. Numerous experimental and theoretical analyses have demonstrated that particle size, shape, and surface modification directly influence their interaction with cells. Understanding cellular uptake of nanoparticles helps us better design nanomaterials. Therefore, the precise design of nanoparticles with specific logP and cellular uptake is of great significance for applications in nanotechnology medicine, toxicology, and other fields.
[0003] The hydrophilicity / hydrophobicity and cellular uptake of nanomaterials can be measured experimentally. However, experimental methods are expensive, time-consuming, and dependent on equipment and the technical skill level of the testers. Given the vast variety and number of nanomaterials, it is difficult to determine their hydrophilicity / hydrophobicity and cellular uptake individually using experimental methods. Therefore, there is an urgent need for non-experimental methods to rapidly predict and evaluate the hydrophilicity / hydrophobicity and cellular uptake of nanomaterials to meet the needs of screening and designing novel nanomaterials.
[0004] Artificial intelligence methods, represented by machine learning and deep learning, can accurately predict the properties of unknown objects by building reliable models from existing data. In the past decade or so, with the emergence of big data and the rapid development of computer hardware, artificial intelligence has been widely applied in many fields such as facial recognition, autonomous driving, and smart healthcare. In the field of cheminformatics, machine learning and deep learning have also been successfully used to predict the properties of compounds. However, compared with small molecule compounds, nanomaterials have much more complex structures. Therefore, the application of machine learning and deep learning to predict the properties of nanomaterials is greatly limited. To address this, we propose a nanoparticle reverse design method based on Quasi-SMILES and recurrent neural networks. Summary of the Invention
[0005] (a) Technical problems to be solved
[0006] To address the shortcomings of existing technologies, this invention provides a nanoparticle reverse design method based on Quasi-SMILES and recurrent neural networks, which solves the aforementioned problems.
[0007] (II) Technical Solution
[0008] To achieve the above-mentioned objectives, the present invention provides the following technical solution: a nanoparticle reverse design method based on Quasi-SMILES and recurrent neural networks, comprising the following steps:
[0009] Step 1: Construction of nanoparticle generation model dataset and conversion of nanostructure information;
[0010] Step 2: Building a deep learning generative model;
[0011] Step 3: Extract nanoparticle structure information through a deep learning model and output new nanostructure information;
[0012] Step 4: Evaluation of the rationality and synthetic feasibility of the nanoparticle structure;
[0013] Step 5: Construction of a dataset on the hydrophilicity / hydrophobicity of nanoparticles and their uptake by cells;
[0014] Step 6: Generation of electronic files for the three-dimensional structure of nanoparticles and calculation of tetrahedral descriptors;
[0015] Step 7: Machine learning prediction model construction and five-fold cross-validation;
[0016] Step 8: Prediction of the hydrophilicity / hydrophobicity and cellular uptake of nanoparticles.
[0017] Preferably, the nanoparticles comprise a core material and small molecule ligands, wherein the core material includes gold, silver, platinum, and palladium, and the small molecule ligands comprise organic compounds composed of carbon, hydrogen, oxygen, nitrogen, sulfur, phosphorus, and fluorine.
[0018] Preferably, the generative model employs a recurrent neural network, and the input features are in the Quasi-SMILES format, which encompasses nanoparticle structural information.
[0019] Preferably, the evaluation indicators for the rationality of the generated nanostructure include particle size, ligand loading, and syntheticability.
[0020] Preferably, the prediction model is constructed using the random forest method, and the predicted properties include the hydrophilicity / hydrophobicity of nanoparticles and cellular uptake.
[0021] (III) Beneficial Effects
[0022] Compared with existing technologies, this invention provides a nanoparticle reverse design method based on Quasi-SMILES and recurrent neural networks, which has the following advantages:
[0023] 1. The model constructed in this invention can be used to predict the hydrophilicity / hydrophobicity and cellular uptake of various nanoparticles. This method is simple, rapid, low-cost, and easy to operate, making it readily usable even for researchers without a computational chemistry background. The method for predicting nanoparticle hydrophilicity / hydrophobicity and cellular uptake conforms to the OECD's principles for the development and use of QSAR models. Therefore, using the prediction results of nanoparticle hydrophilicity / hydrophobicity and cellular uptake using this patented invention can provide data support and theoretical guidance for the rational design of nanomaterials, and is of great significance for evaluating nano-biological effects.
[0024] 2. The model dataset generated by this invention contains 455 nanoparticles, covering three core materials and 244 different surface ligand small molecules; the hydrophobicity prediction model dataset contains 147 nanoparticles, covering three core materials and 91 different surface ligand small molecules, which is currently the largest and most diverse nanoparticle hydrophobicity prediction model containing nanomaterials; the cellular uptake prediction model dataset contains 71 nanoparticles, including one core material and 45 different surface ligand small molecules.
[0025] 3. The model of this invention is constructed using recurrent neural network method and traditional machine learning method. It has high code integration and strong usability, and can maximize the realization of automated operation.
[0026] 4. This invention constructs and evaluates models in accordance with the OECD's principles for the construction and use of QSAR models. The constructed model has good fitting ability, stability and predictive ability, and can be used for predicting the hydrophilicity and hydrophobicity of nanomaterials and cellular uptake, as well as for the rational design of nanomaterials.
[0027] 5. This invention avoids researchers blindly conducting experiments in a vast chemical space.
[0028] 6. Researchers can use this model to design nanoparticles with desired properties without prior knowledge. Attached Figure Description
[0029] Figure 1 This is a flowchart of a method for predicting the hydrophilicity / hydrophobicity and cellular uptake of nanoparticles based on recurrent neural networks and machine learning.
[0030] Figure 2 This is a diagram of the recurrent neural network model framework in Embodiment 1 of the present invention;
[0031] Figure 3 This is a PCA diagram of the input and output data of the generated model in Embodiment 1 of the present invention;
[0032] Figure 4 This is the predicted range distribution of the hydrophilicity / hydrophobicity and cellular uptake of the generated nanoparticles in Example 1 of the present invention. Detailed Implementation
[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0034] Please see Figure 1-4 Example 1:
[0035] A method for predicting the hydrophilicity / hydrophobicity and cellular uptake of nanoparticles based on recurrent neural networks, such as... Figure 1 As shown, it includes the following steps:
[0036] Specifically as follows:
[0037] First, basic information (nuclear atom, shape, size, smiles of ligands, number of ligands) of 455 nanoparticles was collected. A generative model constructed using a recurrent neural network was then used to concatenate the collected information (e.g., Au(Sphere)(5)176{147}.NCCCNC(=O)CCCCC1CCSS1.O=C(OC1=CC=C(NC(CCCC)=O)C=C1)CCCCC2SSCC2), resulting in 455 lines of input data for the generative model. The training epochs of the generative model were adjusted to 2,000,000, ensuring the generated results matched the structure of the concatenated input information. The generated results were then filtered, removing non-standard and unreasonable data points. The generated data was used to construct electronic files (PDB format) containing three-dimensional structural information of nanoparticles (such as particle size, atom type, atomic coordinates, atomic connections, etc.) using VINAS (Virtual Nanostructure Simulation) software developed by Hao Zhu's laboratory at Rutgers University.
[0038] Simultaneously, hydrophilicity / hydrophobicity (Nano-logP) experimental values of 147 nanoparticles were collected. These nanoparticles, synthesized using nanocombinatorial chemistry, encompassed three core materials (gold, platinum, and palladium) and 91 small molecule ligands. Additionally, nano-cell uptake experimental values of 71 nanoparticles were collected, all with gold cores and encompassing 45 small molecule ligands. Their physicochemical properties (particle size, ligand number, Nano-logP, Nano-cell uptake, etc.) underwent rigorous characterization and testing to ensure the reliability of the experimental data. Based on the experimentally measured nanoparticle size, ligand number, and other information, electronic files (PDB format) containing three-dimensional structural information of the nanoparticles (such as particle size, atom type, atomic coordinates, and atomic connections) were constructed using VINAS software. Furthermore, the tetrahedral descriptors of the nanoparticles were calculated using the obtained PDB files and imported as input variables into a prediction model, ultimately outputting the predicted values of nanoparticle hydrophilicity / hydrophobicity and nano-cell uptake.
[0039] The tokens in the generative model contain all the symbols in the input data. The tokens are composed as follows: tokens=['<','>','#',')','(','[',']','{','}','@',':','*','+','-',' / ',','.','1','0','3','2','5','4','7','T','6','9','8','=','A','B','C','D','E','F','G','H','I','J','K', 'L','M','N','[',']','O','P','Q','R','S','T','U','V','W','X','Y','Z','r','s','t','u','v','w','x','y','z','\\','a','b','c','d','e','f','g','h','i','j','k','l','m','n','o','p','-']. Where '<' and '>' serve as the start and end symbols for each data point, respectively.
[0040] The generative model used employs a stack-augmented recurrent neural network (Stack-RNN) framework. Stack-RNN defines a new neuron or cell structure on top of a standard GRU cell, with two additional multiplication gates called a memory stack. This allows the stack-RNN to learn meaningful long-range interdependencies. The stack memory is a differentiable structure where continuous vectors can be inserted and removed. The data generated by the generative model is compared with the input data, removing unreasonable data and output data identical to the input data from the output nanoparticle information. Next, ligand small molecules are extracted from the generated nanoparticle information, and the ease of synthesis of the generated ligand small molecules is scored using a comprehensive reachability score (SAS). This method relies on knowledge extracted from known synthetic reactions and adds a penalty for macromolecular complexity. For ease of interpretation, the SAS is scaled between 1 and 10. Molecules with high SAS values (typically above 6) are considered difficult to synthesize, while molecules with low SAS values are easy to synthesize.
[0041] The dataset containing the hydrophilicity / hydrophobicity of nanomaterials is the largest dataset on this topic, containing 147 nanomaterials, covering three core materials (gold, platinum, and palladium) and 91 small molecule ligands, exhibiting diversity in nanostructures; the hydrophilicity / hydrophobicity values of the nanomaterials have a wide range of distribution, from -2.68 to 2.72. The dataset containing cellular uptake of nanomaterials contains 71 nanomaterials, all with gold cores, covering 45 small molecule ligands, exhibiting diversity in nanostructures; the cellular uptake values of the nanomaterials have a wide range of distribution, from 0.74 to 22.81. The diversity of nanostructures and the wide range of predicted values are beneficial for building robust predictive models.
[0042] The prediction model used is the classic machine learning algorithm—random forest. Tetrahedral descriptors calculated from experimental test values are divided into training and test sets (the test set accounts for 20% of the total data). Five-fold cross-validation is performed on the training set, and the training and test sets with the best initial scores are saved. The model parameters (random_state, n_estimators, max_depth, max_features) are adjusted using the training set with the best initial scores, and the parameters with the best model performance are saved (the optimal parameters for the hydrophobicity prediction model are random_state = 36, n_estimators = 131, max_depth = 12, max_features = 109; the optimal parameters for the cell uptake prediction model are random_state = 21, n_estimators = 51, max_depth = 8, max_features = 37). The model is then rebuilt using the optimally partitioned training and test sets and the optimal parameters. The tetrahedral descriptor calculated from the nanoparticle information output by the generative model is substituted into the trained model to obtain the predicted values of hydrophilicity / hydrophobicity and cellular uptake of the nanoparticles generated by the generative model.
[0043] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A nanoparticle reverse design method based on Quasi-SMILES and recurrent neural networks, characterized in that, Includes the following steps: Step 1: Construction of nanoparticle generation model dataset and conversion of nanostructure information; Step 2: Building a deep learning generative model; The generative model uses a recurrent neural network, and the input features are in the Quasi-SMILES format, which includes information on the structure of nanoparticles. Step 3: Extract nanoparticle structure information through a deep learning model and output new nanostructure information; Step 4: Evaluation of the rationality and synthetic feasibility of the nanoparticle structure; Among them, the evaluation indicators for the rationality of generating nanostructures include particle size, ligand loading, and syntheticability. The nanoparticles contain a core material and small molecule ligands, wherein the core material includes gold, silver, platinum, and palladium; Step 5: Construction of a dataset on the hydrophilicity / hydrophobicity of nanoparticles and their uptake by cells; Step 6: Generation of electronic files for the three-dimensional structure of nanoparticles and calculation of tetrahedral descriptors; Step 7: Machine learning prediction model construction and five-fold cross-validation; Step 8: Prediction of the hydrophilicity / hydrophobicity and cellular uptake of nanoparticles.
2. The nanoparticle reverse design method based on Quasi-SMILES and recurrent neural networks according to claim 1, characterized in that: Small molecule ligands include organic compounds composed of carbon, hydrogen, oxygen, nitrogen, sulfur, phosphorus, and fluorine.
3. The nanoparticle reverse design method based on Quasi-SMILES and recurrent neural networks according to claim 1, characterized in that: The prediction model was constructed using the random forest method, and the predicted properties included the hydrophilicity / hydrophobicity of nanoparticles and cellular uptake.