Nanoparticle preparation method based on machine learning and high-throughput screening

Through nanoparticle preparation methods based on machine learning and high-throughput screening, a pharmacological model of metal polyphenol network is constructed, which solves the problems of particle size regulation and structural consistency in the preparation of metal polyphenol network nanoparticles, and achieves efficient and stable nanoparticle preparation and large-scale production.

CN120220883APending Publication Date: 2025-06-27SOUTHERN MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510273780.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the process of preparing metal polyphenol network nanoparticles, it is difficult to achieve precise control of particle size and structural consistency, resulting in large fluctuations in performance between product batches. Some preparation conditions require strict pH, temperature or metal ion concentration, which limits the stability of large-scale production.

Method used

Using nanoparticle preparation method based on machine learning and high-throughput screening, a calculation pharmacological model is constructed in metal polyphenol networks, the complex interaction laws between multiple metal ions and polyphenol compounds are analyzed, and the formula with the best physical and chemical characteristics meets the preset requirements is generated, and the formula with the best antioxidant activity is screened through high-throughput methods.

Benefits of technology

It significantly improves the efficiency of formula R&D, reduces the cost of testing, realizes the uniformity and stability control of nanoparticles, reduces performance fluctuations between batches, and supports the stability of large-scale production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220883A_ABST
    Figure CN120220883A_ABST
Patent Text Reader

Abstract

The invention provides a nanoparticle preparation method based on machine learning and high-throughput screening, which comprises the following steps: carrying out nanoparticle preparation experiments under different experiment conditions by using polyphenol and metal ions, and recording experiment results; extracting feature tags based on experimental conditions and experimental results, and constructing a preliminary screening data set; performing model training according to the relationship between the experimental feature data in the preliminary screening data set and the experimental result and a least square lifting algorithm to obtain a basic prediction model; expanding the preliminary screening data set through different prediction verification means, and performing iterative training on the basic prediction model based on the expanded data set to obtain a standard prediction model; generating a formula with physicochemical properties meeting preset requirements through a standard prediction model, and screening out a formula with optimal antioxidant activity based on a high-throughput method; and completing a preparation experiment by adopting an optimal formula to obtain the nanoparticles, and characterizing the nanoparticles. According to the method, the formula research and development efficiency is remarkably improved, and the test cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical technologies, and particularly to a method for preparing nanoparticles based on machine learning and high-throughput screening. Background Art

[0002] In recent years, the booming development of nanotechnology has brought revolutionary progress to many fields. Among them, metal polyphenol network nanoparticles, as an emerging nanomaterial, have attracted a great deal of attention. These materials combine the coordination ability of metal ions and the unique chemical properties of polyphenol compounds, and through a simple self-assembly process, nanoparticles with stable structures and diverse functions can be prepared. Metal polyphenol networks not only show great potential in the field of drug delivery, such as achieving efficient encapsulation and controlled release of drug molecules, but are also widely used in fields such as tumor diagnosis and treatment, antibacterial therapy, tissue engineering, and photothermal therapy. In addition, thanks to their diverse metal-phenolic hydroxyl complexation modes, the components of these nanoparticles can be flexibly adjusted, enabling tunable properties in optical sensing, magnetic resonance imaging, and catalytic reactions. For example, introducing rare earth metals into the network can enhance the magnetic resonance contrast, and introducing metals with redox activity can enhance the catalytic efficiency. Therefore, the research and application of metal polyphenol network materials are gradually moving from the laboratory to practical application scenarios. Their unique physical and chemical properties and design flexibility make them have broad application prospects in the fields of future materials science, biomedical engineering, and environmental protection.

[0003] In drug R & D and formulation design, traditional experimental methods are usually time-consuming and costly. In recent years, computational pharmaceutics has gradually become an effective research tool. By applying mathematical models and computer algorithms, it can quickly screen and optimize formulations, improving R & D efficiency. In this process, the introduction of machine learning technology has enabled computational pharmaceutics to gradually evolve from early quantitative structure-activity relationship analysis (QSAR) into a complex, multi-dimensional modeling tool. Machine learning can predict key parameters such as the solubility, stability, metabolic properties of compounds, and their interactions with carriers by learning and generalizing a large amount of experimental data, thus providing guidance for the optimized design of nanocarriers and drug molecules. A variety of algorithms represented by supervised learning and unsupervised learning can analyze the coordination modes of different metal ions and polyphenol molecules, and find formulation conditions conducive to network stability and biocompatibility; while through deep learning, potential laws can be mined in a more complex multi-parameter space, providing more accurate predictions for the design of new materials and personalized drug delivery systems. In the field of pharmaceutics, this technology has been used to predict drug release curves, improve the physicochemical properties of nanoparticles, and accelerate the high-throughput screening of new materials, thus significantly shortening the R & D cycle, reducing R & D costs, and improving the quality and efficacy of the final products. Therefore, with the continuous progress of machine learning technology, the application potential of computational pharmaceutics models in materials science, drug development, and personalized medicine is becoming increasingly significant, bringing new impetus to technological innovation in the biomedical field.

[0004] Currently, in the preparation process, controlling the homogeneity and stability of metal polyphenol networks is a major challenge. Conventional methods often struggle to achieve precise control of particle size and structural consistency, resulting in significant performance fluctuations between product batches; in addition, certain preparation conditions require strict pH values, temperatures, or metal ion concentrations, which pose limitations to the stability of large-scale production. Due to the wide variety of polyphenolic compounds, their reaction conditions and product characteristics may vary significantly, so it remains quite difficult to establish a highly generalizable and efficient preparation strategy. Summary of the Invention

[0005] In view of this, the present invention provides a method for preparing nanoparticles based on machine learning and high-throughput screening to solve the above problems.

[0006] The present invention provides a method for preparing nanoparticles based on machine learning and high-throughput screening, including: conducting experiments on the preparation of metal polyphenol network nanoparticles by using polyphenols and metal ions under different experimental conditions, and recording the experimental results; extracting characteristic tags based on the experimental conditions and results; constructing a primary screening data set based on the experimental characteristic data corresponding to the characteristic tags; performing model training according to the relationship between the experimental characteristic data in the primary screening data set and the experimental results and the least squares boosting algorithm to obtain a basic prediction model; expanding the primary screening data set through different prediction verification means, and performing iterative training on the basic prediction model based on the expanded data set to obtain a standard prediction model; generating a formulation with physicochemical properties meeting the preset requirements through the standard prediction model, and screening out the formulation with the optimal antioxidant activity therefrom based on the high-throughput method; completing the preparation experiment by using the optimal formulation to obtain nanoparticles, and characterizing the nanoparticles.

[0007] In another implementation manner of the present invention, the polyphenols used in the experiment on the preparation of metal polyphenol network nanoparticles are danshensu, salvianolic acid B, protocatechuic acid, and hydroxysafflor yellow A; the metal ion used is cerium ion.

[0008] In another implementation manner of the present invention, the characteristic tags include ratio parameter characteristic tags and process parameter characteristic tags; the ratio parameter characteristic tags include the concentrations of each component, metal concentration, the ratio relationship between each component and the metal, total component concentration, total metal concentration, and the ratio relationship between phenolic hydroxyl and the metal; the process parameter characteristic tags include reaction volume, reaction time, rotor speed, and solution pH value.

[0009] In another implementation manner of the present invention, it further includes: preprocessing the experimental characteristic data corresponding to the characteristic tags, and the preprocessing includes data cleaning, standardization processing, normalization processing, and data dimensionality reduction processing.

[0010] In another implementation manner of the present invention, the data dimensionality reduction processing uses the t-SNE method to reduce the experimental characteristic data corresponding to the characteristic tags to two-dimensional data; performing visual analysis on the two-dimensional data to obtain the distribution law.

[0011] On the other hand, the present invention provides a nanoparticle preparation system based on machine learning and high-throughput screening, including: a data acquisition module: performing experiments on the preparation of metal polyphenol network nanoparticles by using polyphenols and metal ions under different experimental conditions, and recording the experimental results; a data processing module: extracting characteristic tags based on the experimental conditions and experimental results; constructing a preliminary screening data set based on the experimental characteristic data corresponding to the characteristic tags; a model construction module: performing model training according to the relationship between the experimental characteristic data in the preliminary screening data set and the experimental results and the least squares boosting algorithm to obtain a basic prediction model; a parameter optimization module: expanding the preliminary screening data set through different prediction verification means, and performing iterative training on the basic prediction model based on the expanded data set to obtain a standard prediction model; a formulation prediction module: generating a formulation with physical and chemical properties meeting the preset requirements through the standard prediction model, and screening out the formulation with the optimal antioxidant activity therefrom based on the high-throughput method; a nanoparticle preparation module: completing the preparation experiment by using the optimal formulation to obtain nanoparticles, and characterizing the nanoparticles.

[0012] On the other hand, the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the steps of a nanoparticle preparation method based on machine learning and high-throughput screening as described in any one of the above are implemented.

[0013] On the other hand, the present invention provides a computer storage medium, characterized in that a computer program is stored on the computer storage medium, and when the computer program is executed by a processor, the steps in a nanoparticle preparation method based on machine learning and high-throughput screening as described in any one of the above are implemented.

[0014] The nanoparticle preparation method based on machine learning and high-throughput screening of the present invention realizes intelligent prediction and optimization of formulations by constructing a metal polyphenol network computational pharmaceutics model based on machine learning; by using this model, the complex interaction rules between various metal ions and polyphenol compounds can be effectively analyzed, so that without relying on a large number of experimental screenings, the best formulation with specific physical and chemical properties, stability and biological functions can be quickly predicted; significantly improving the formulation R & D efficiency, reducing the test cost, and providing a scientific basis for subsequent preparation and functional modification, further promoting the application of metal polyphenol network materials in the fields of drug delivery and biomedicine. Description of the Drawings

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. By reading the detailed description of the following embodiments, the advantages and benefits in the solutions will become clear to those skilled in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. In the drawings:

[0016] Figure 1 Schematic flow chart of the preparation method of nanoparticles based on machine learning and high-throughput screening according to an embodiment of the present invention.

[0017] Figure 2 Schematic diagram of the polyphenol structure according to an embodiment of the present invention.

[0018] Figure 3 Schematic diagram of the particle size-pdi distribution of the primary screening data according to an embodiment of the present invention.

[0019] Figure 4 Schematic diagram of the t-sne distribution after dimensionality reduction of the formulation parameters of each group in the primary screening according to an embodiment of the present invention.

[0020] Figure 5 Schematic diagram of the fitting situation of the primary screening data predicted by six algorithms according to an embodiment of the present invention.

[0021] Figure 6 Schematic diagram of the comparison of the performance and training weights of the LsBoost model before and after data iteration according to an embodiment of the present invention.

[0022] Figure 7 Schematic diagram of the formulation prediction of the optimal interval results before and after model iteration according to an embodiment of the present invention.

[0023] Figure 8 Schematic diagram of the active high-throughput detection results of each batch according to an embodiment of the present invention.

[0024] Figure 9 Schematic diagram of the dynamic light scattering detection results of the nanoparticles with the optimal formulation according to an embodiment of the present invention.

[0025] Figure 10 Picture taken by a transmission electron microscope according to an embodiment of the present invention. Detailed implementation manners

[0026] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following will clearly and detailedly describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present invention.

[0027] Figure 1 Schematic diagram of a nanoparticle preparation method based on machine learning and high-throughput screening provided by an embodiment of the present invention. As Figure 1 shown, this embodiment mainly includes:

[0028] S101. Conduct experiments on the preparation of metal polyphenol network nanoparticles by using polyphenols and metal ions under different experimental conditions, and record the experimental results.

[0029] Exemplarily, in the experimental design, two variation methods are set: one is to use the same formulation and change the reaction conditions; the other is to adopt different formulations and keep the reaction conditions unchanged.

[0030] S102. Extract feature tags based on the experimental conditions and experimental results.

[0031] Exemplarily, by analyzing the internal relationships between the formulations and combining the experimental conditions and results, a set of parameters for describing the characteristics of the experimental data are extracted, totaling 22 feature tags. These tags include the concentrations of SAA, SAB, PCA, HYSA, metal concentration, drug-to-metal ratio, Tris volume fraction, pH, rotation speed, time, total reaction volume, and the specific quantity values and proportional relationships of components such as SAA, SAB, PCA, and HYSA.

[0032] S103. Construct a primary screening dataset based on the experimental feature data corresponding to the feature tags.

[0033] Exemplarily, by quantifying the experimental preparation method and randomly designing experiments, the experimental feature data corresponding to the feature tags are used as input variables of the experimental results to form a primary screening dataset for training and predicting the experimental results. As Figure 3 shown, the analysis results of the primary screening dataset show that the particle size distribution is concentrated in the range of 100 - 1000 nm, and the distribution ratio is concentrated between 0.2 - 0.6.

[0034] S104. Perform model training according to the relationship between the experimental feature data and the experimental results in the primary screening dataset and the least squares boosting algorithm to obtain a basic prediction model.

[0035] S105. Expand the preliminary screening dataset through different prediction verification means, and iteratively train the basic prediction model based on the expanded dataset to obtain a standard prediction model.

[0036] Exemplarily, according to the characteristics of the Least Squares Boosting (LSBoost) algorithm architecture, call the feature importance score to conduct a scientific quantitative analysis of the model influence factors, such as Figure 6 shown. The model analysis results show that the main parameters affecting particle size include reaction time, Tris dosage, and the pH value of the reaction environment, while the main parameters affecting particle size distribution include reaction time, Tris dosage, and the molar ratio of drug to metal ion. Based on the basic prediction, by performing floating calculations on the superior formulation parameters, several formulations are randomly generated, and the generated formulations are experimentally verified according to the model prediction results.

[0037] As Figure 7 shown, the area selected by the red box is the range of particle size 100 - 120 nm and pdi value 0.1 - 0.3. Among the formulations randomly generated by the basic prediction model, the formulation data that meet the target particle size (100 - 120 nm) and distribution (0.1 - 0.3) range are relatively scarce, indicating that the basic prediction model is not yet sufficient to accurately predict the formulation parameters that meet the conditions. Therefore, through different prediction verification means, the dataset is gradually expanded, and the model is iteratively trained multiple times to improve the overall prediction ability.

[0038] After five iterations, the prediction accuracy of the particle size of the model is improved to 0.84, and the prediction accuracy of the distribution is improved to 0.89, significantly enhancing the prediction reliability of the model. From the analysis results of the latest training weights, the pH value is still the most important parameter affecting particle size, while the secondary parameters include the overall drug concentration, and this conclusion is consistent with the experience in the previous screening experiments.

[0039] Through the above optimization, the model has achieved a significant improvement in the formulation prediction accuracy, and at the same time further verified the key influencing factors of particle size and distribution, providing a scientific basis and technical support for subsequent formulation design and optimization of drug delivery systems.

[0040] Compared with the traditional optimization method based on experience and experiments, the computational pharmaceutics model constructed using the Least Squares Boosting (LsBoost) algorithm gradually optimizes the dataset and model performance through floating calculations and experimental verification, greatly improving the prediction accuracy of formulation parameters and particle size distribution, enabling rapid screening and optimization of formulations, thus promoting the development of innovative drug formulations and new drug delivery technologies, providing important technical support for the biomedical field. This data-driven optimization method significantly reduces the number of experiments and R & D costs, and improves the R & D efficiency.

[0041] S106. Generate a formulation with physicochemical properties meeting the preset requirements through the standard prediction model, and screen out the formulation with the optimal antioxidant activity from it based on the high-throughput method.

[0042] Exemplarily, through the standard prediction model, a relatively optimal formulation meeting the preset requirements can be generated, and the formulation with the optimal antioxidant activity can be screened out from multiple relatively optimal formulations in combination with the ROS high-throughput method.

[0043] S107. Complete the preparation experiment using the optimal formulation to obtain nanoparticles, and characterize the nanoparticles.

[0044] The method for preparing nanoparticles based on machine learning and high-throughput screening of the present invention realizes the intelligent prediction and optimization of the formulation by constructing a computational pharmaceutics model of metal polyphenol network based on machine learning; using this model, the complex interaction rules between various metal ions and polyphenol compounds can be effectively analyzed, so that without relying on a large number of experimental screenings, the best formulation with specific physicochemical properties, stability and biological functions can be quickly predicted; significantly improves the formulation R & D efficiency, reduces the test cost, and provides a scientific basis for subsequent preparation and functional modification, further promoting the application of metal polyphenol network materials in the fields of drug delivery and biomedicine.

[0045] In another implementation manner of the present invention, the polyphenols used in the metal polyphenol network nanoparticle preparation experiment are danshensu, salvianolic acid B, protocatechuic acid, hydroxysafflor yellow A; the metal ion used is cerium ion.

[0046] Exemplarily, using the basic results of previous experimental studies, the ratio and reaction conditions of polyphenols and metal ions are designed and optimized, and metal polyphenol network units and quaternary structures with different structural characteristics are successfully prepared. These metal polyphenol networks not only achieve the responsive release of polyphenols, but also improve the in vivo stability and bioavailability of polyphenols by optimizing the complexation structure. The characteristic of responsive release enables polyphenols to be rapidly released under specific physiological conditions while maintaining higher stability in a stable environment, thus improving their overall performance in vivo.

[0047] Such as Figure 2As shown in the figure, the study screened the important active ingredients in Danshen-Safflower, identified four polyphenolic ingredients as the core ingredients for treating atherosclerosis, and selected these four polyphenols as carrier drugs. In addition, the study used iron-o-diphenol hydroxyl to construct metal polyphenol network coordination polymer nanoparticles. There are many disadvantages of iron ions in the body, such as iron ions will cause certain interference with the redox balance in the biological environment, which may induce oxidative stress response, thereby bringing potential toxicity, and excessive iron ions may form uncontrollable free radicals in the body, leading to cell damage and inflammatory response, thereby affecting the normal function of tissues. Based on this, in the new preparation process, other metal ions (cerium ions) were selected as metals in the formation of polyphenol networks. Cerium ions have the following advantages, which will help promote the treatment of atherosclerosis.

[0048] First, cerium ions have unique redox capabilities, which can 3+ and Ce 4+ This property enables dynamic transformation between the two, and this property gives it antioxidant and free radical scavenging functions, thereby significantly reducing the potential toxicity caused by oxidative stress reactions. Secondly, under certain conditions, cerium ions are better than traditional antioxidants in scavenging reactive oxygen species (ROS). This property can directly act on the inflammatory environment in atherosclerotic lesions, helping to reduce inflammation and oxidative damage to the blood vessel walls. Finally, cerium ions show good stability and low toxicity in the biological environment, and the network formed after complexing with polyphenols has efficient drug carrier capabilities, which can not only achieve the sustained release and targeted release of polyphenol components, but also further enhance the anti-inflammatory and antioxidant effects of these polyphenols.

[0049] The present invention aims to utilize the unique characteristics of the metal polyphenol network, and to construct a drug carrier with responsive release and high bioavailability by optimizing the ratio of polyphenols to metal ions and reaction conditions. Combining the above characteristics, the use of cerium ions to construct metal polyphenol network nanoparticles can not only solve the biosafety problem of iron ion introduction, but also give full play to the unique biological advantages of cerium ions, improve the biosafety and stability of drug carriers, thereby improving the overall performance of the drug delivery system, and providing a more efficient and safer solution for the treatment of atherosclerosis.

[0050] In another implementation of the present invention, the feature tags include ratio parameter feature tags and process parameter feature tags; the ratio parameter feature tags include the concentration of each component, metal concentration, the ratio of each component to the metal, the total component concentration, the total metal concentration, and the ratio of phenolic hydroxyl to metal; the process parameter feature tags include reaction volume, reaction time, rotor speed, and solution pH.

[0051] Exemplarily, as shown in Table 1, a primary screening data set containing 22 feature tags was constructed, covering formulation parameters and process parameters, laying a solid data foundation for the construction and optimization of the machine learning model.

[0052] Table 1 Feature Tags

[0053]

[0054] Combined with the analysis of feature importance in machine learning, the present invention clarifies the key influencing factors of particle size and distribution, such as pH value, reaction time, and the ratio of metal to polyphenol. This quantitative analysis provides a scientific basis for subsequent formulation adjustment. Compared with traditional technologies, the optimization process is more transparent and more instructive.

[0055] In another implementation manner of the present invention, it further includes: preprocessing the experimental feature data corresponding to the feature tags, and the preprocessing includes data cleaning, standardization processing, normalization processing, and data dimensionality reduction processing.

[0056] Exemplarily, through further processing of the data, by using the methods of data cleaning and standardization, the influence of dimension on data in different dimensions is removed, ensuring that all variables are in the same order of magnitude, and data normalization processing is completed. This step of processing makes the subsequent model fitting process more accurate.

[0057] In another implementation manner of the present invention, the data dimensionality reduction processing uses the t-SNE method to reduce the experimental feature data corresponding to the feature tags to two-dimensional data; visual analysis is performed on the two-dimensional data to obtain the distribution law.

[0058] Exemplarily, in the data dimensionality reduction analysis, the 22-dimensional parameters are reduced to two dimensions by the t-SNE method, and the distribution law is visualized. As Figure 4 shown, the data after dimensionality reduction shows obvious distribution characteristics in the t-SNE space, which directly verifies the effectiveness of the primary screening data construction method and lays a foundation for subsequent model establishment and performance optimization.

[0059] In another implementation manner of the present invention, as Figure 8 shown, pharmacodynamic evaluation is carried out. Through in vitro cell experiments, the activities of different formulated nanoparticles in removing endogenous ROS are detected. For the generated 100 model formulations, high-throughput activity screening is carried out, and finally the E61 (batch number 6, number 1) formulation is determined as the preparation plan with the best activity.

[0060] The specific formulation composition and preparation method are as follows:

[0061] First, prepare the mother liquors of danshensu, salvianolic acid B, protocatechuic aldehyde, and hydroxysafflor yellow A, and their respective concentrations are all 2 mg / mL.

[0062] Prepare a stock solution of Ce(NO3)3·6H2O with a concentration of 15 mg / mL.

[0063] Prepare a stock solution of tris(hydroxymethyl)aminomethane (Tris) with a concentration of 100 nM.

[0064] Operate in the following order:

[0065] Sequentially add 0.361 mL of danshensu solution, 0.197 mL of salvianolic acid B solution, 0.212 mL of protocatechuic aldehyde solution, 0.165 mL of hydroxysafflor yellow A solution, and 0.307 mL of Ce(NO3)3·6H2O solution, and mix well on a magnetic stirrer.

[0066] Subsequently, add 3.535 mL of pure water, stir for 1 minute, and then add 0.134 mL of Tris solution.

[0067] Stir the mixed solution at 1000 rpm for 24 hours at room temperature.

[0068] After stirring is completed, centrifuge at 8000 rpm for 5 minutes, take the supernatant to obtain a suspension containing nanoparticles.

[0069] As Figure 9 and Figure 10 shown, the nanoparticle size and its distribution were measured by dynamic light scattering method, and characterized by transmission electron microscopy to observe the nanoparticle morphology. The results show that its diameter size is consistent with the detection data.

[0070] In another implementation of the present invention, as shown in Table 2, based on the primary screening dataset, regression fitting analysis was performed on the relationship between experimental features and results using six mainstream machine learning algorithms, specifically including decision tree (DT), random forest (RF), Gaussian process regression (GPR), support vector machine (SVM), least squares boosting (LsBoost), and partial least squares (PLS).

[0071] As Figure 5 shown, the fitting results of the LsBoost model are within the red frame. After systematic comparison, it is found that the LsBoost algorithm performs best in predicting particle size and distribution, and the prediction accuracies reach 0.79 and 0.76 respectively. In contrast, the other algorithms either have poor fitting effects or obvious overfitting problems. The excellent performance of LsBoost benefits from its strong anti-overfitting ability, good adaptability to high-dimensional features, and high stability in different data batches.

[0072] Table 2 Parameters for running six mainstream machine learning algorithms

[0073]

[0074]

[0075] In the modeling process, considering the advantages and disadvantages of various machine learning algorithms comprehensively, the Least Squares Boosting (LSBoost) algorithm with stable performance and good interpretability was selected as the modeling tool. The advantage of the least squares boosting algorithm is that it can effectively process high-dimensional features and has good anti-overfitting ability, which is used to accurately predict the relationship between formulation parameters and particle size distribution. By inputting the cleaned data into this algorithm for training, a computational pharmaceutics model of formulation parameters - particle size, that is, a standard prediction model, was finally successfully constructed.

[0076] The solution of the present invention not only has broad application prospects in the fields of pharmaceutical industry, precision medicine and drug research and development, but also can enhance the therapeutic effect, reduce side effects and reduce the research and development cost through a more efficient drug delivery system.

[0077] On the other hand, the present invention provides a nanoparticle preparation system based on machine learning and high-throughput screening, including:

[0078] Data acquisition module: Using polyphenols and metal ions to conduct experiments on the preparation of metal polyphenol network nanoparticles under different experimental conditions and recording the experimental results.

[0079] Data processing module: Extracting feature tags based on the experimental conditions and experimental results; constructing a primary screening data set based on the experimental feature data corresponding to the feature tags.

[0080] Model construction module: Training a model according to the relationship between the experimental feature data in the primary screening data set and the experimental results and the least squares boosting algorithm to obtain a basic prediction model.

[0081] Parameter optimization module: Expanding the primary screening data set through different prediction verification means and iteratively training the basic prediction model based on the expanded data set to obtain a standard prediction model.

[0082] Formulation prediction module: Generating a formulation with physical and chemical properties meeting preset requirements through the standard prediction model and screening out the formulation with the optimal antioxidant activity from it based on a high-throughput method.

[0083] Nanoparticle preparation module: Completing the preparation experiment with the optimal formulation to obtain nanoparticles and characterizing the nanoparticles.

[0084] The nanoparticle preparation system based on machine learning and high-throughput screening of the present invention realizes intelligent prediction and optimization of formulations by constructing a metal polyphenol network computational pharmaceutics model based on machine learning. Using this model, the complex interaction rules between various metal ions and polyphenol compounds can be effectively analyzed, so that the best formulation with specific physicochemical properties, stability and biological functions can be quickly predicted without relying on a large number of experimental screenings. The formulation R & D efficiency is significantly improved, the test cost is reduced, and a scientific basis is provided for subsequent preparation and functional modification, further promoting the application of metal polyphenol network materials in the fields of drug delivery and biomedicine.

[0085] On the other hand, the electronic device of the present invention includes: a processor, a memory, and a communication bus and a communication interface.

[0086] Wherein:

[0087] The processor, the memory, and the communication interface complete communication with each other through the communication bus.

[0088] The communication interface is used to communicate with other electronic devices or servers.

[0089] The processor is used to execute a program, and specifically can execute the steps of any one of the nanoparticle preparation methods based on machine learning and high-throughput screening in the above embodiments.

[0090] Specifically, the program may include program code, and the program code includes computer operation instructions.

[0091] The processor may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0092] The memory is used to store the program. The memory may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0093] The program can specifically be used to cause a processor to execute steps for implementing any one of the nanoparticle preparation methods based on machine learning and high-throughput screening described in the embodiments. For the specific implementation of each step in the program, reference can be made to the corresponding descriptions in the steps and units of any one of the nanoparticle preparation methods based on machine learning and high-throughput screening described above, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and modules described above can refer to the corresponding process descriptions in the foregoing method embodiments.

[0094] The method according to the embodiment of the present invention can be implemented in a server equipped with a central processing unit (CPU) and an image processing unit (GPU).

[0095] So far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result.

[0096] It should be noted that all directional indications (such as up, down, left, right, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship between components in a specific order (as shown in the drawings). If this specific order changes, the directional indications will change accordingly.

[0097] In the description of the present invention, the terms "first" and "second" are only used to conveniently describe different components or names, and cannot be construed as indicating or implying an order relationship, relative importance, or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features.

[0098] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0099] It should be noted that although the specific embodiments of the present invention have been described in detail in conjunction with the accompanying drawings, it should not be construed as a limitation on the protection scope of the present invention. Within the scope described in the claims, various modifications and deformations that can be made by those skilled in the art without creative efforts still fall within the protection scope of the present invention.

[0100] The examples of the embodiments of the present invention are intended to concisely illustrate the technical features of the embodiments of the present invention, so that those skilled in the art can intuitively understand the technical features of the embodiments of the present invention, and are not an improper limitation of the embodiments of the present invention.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for preparing nanoparticles based on machine learning and high-throughput screening, characterized in that: include: Use polyphenols and metal ions to prepare metal polyphenol network nanoparticles under different experimental conditions, and record the experimental results; Based on the experimental conditions and experimental results, extract feature labels; Constructing a preliminary screening data set based on the experimental feature data corresponding to the feature labels; Performing model training based on the relationship between the experimental feature data in the primary screening data set and the experimental results and a least squares lifting algorithm to obtain a basic prediction model; Expanding the initial screening data set by different prediction verification methods, and iteratively training the basic prediction model based on the expanded data set to obtain a standard prediction model; Generate a formula whose physical and chemical properties meet the preset requirements through the standard prediction model, and screen out the optimal formula for antioxidant activity based on a high-throughput method; The optimal formula is used to complete the preparation experiment, obtain nanoparticles, and characterize the nanoparticles.

2. The method according to claim 1, characterized in that The polyphenols used in the metal polyphenol network nanoparticle preparation experiment are danshensu, salvianolic acid B, protocatechuol, and hydroxysafflor yellow A; The metal ion used is cerium ion.

3. The method according to claim 2, characterized in that The feature tags include a ratio parameter feature tag and a process parameter feature tag; The ratio parameter characteristic labels include the concentration of each component, the metal concentration, the ratio of each component to the metal, the total component concentration, the total metal concentration, and the ratio of phenolic hydroxyl group to the metal; The process parameter characteristic tags include reaction volume, reaction time, rotor speed, and solution pH value.

4. The method according to claim 1, characterized in that: Also includes: The experimental feature data corresponding to the feature labels are preprocessed, and the preprocessing includes data cleaning, standardization, normalization and data dimension reduction.

5. The method according to claim 4, characterized in that The data dimension reduction process uses the t-SNE method to reduce the experimental feature data corresponding to the feature labels into two-dimensional data; Visual analysis is performed on the two-dimensional data to obtain distribution rules.

6. A nanoparticle preparation system based on machine learning and high-throughput screening, characterized in that: include: Data acquisition module: Use polyphenols and metal ions to conduct metal polyphenol network nanoparticle preparation experiments under different experimental conditions and record the experimental results; Data processing module: extracting feature labels based on the experimental conditions and experimental results; constructing a preliminary screening data set based on the experimental feature data corresponding to the feature labels; Model building module: performing model training according to the relationship between the experimental feature data in the primary screening data set and the experimental results and the least squares lifting algorithm to obtain a basic prediction model; Parameter optimization module: expanding the initial screening data set through different prediction verification methods, and iteratively training the basic prediction model based on the expanded data set to obtain a standard prediction model; Formula prediction module: Generates a formula whose physical and chemical properties meet the preset requirements through the standard prediction model, and selects the optimal formula for antioxidant activity based on a high-throughput method; Nanoparticle preparation module: using the optimal formula to complete the preparation experiment, obtain nanoparticles, and characterize the nanoparticles.

7. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of a method for preparing nanoparticles based on machine learning and high-throughput screening as described in any one of claims 1 to 5 are implemented.

8. A computer storage medium, characterized in that The computer storage medium stores a computer program, and when the computer program is executed by the processor, the steps in the method for preparing nanoparticles based on machine learning and high-throughput screening as described in any one of claims 1 to 5 are implemented.