Method for calculating phase boundary line of self-assembly structure of reversely designed nanoparticles
Through the reverse design method combined with computer simulation and machine learning, the problem of high efficiency and low efficiency in drawing the phase boundary of the target structure in the nanoparticle self-assembly system is solved, and efficient and accurate drawing of the phase boundary and discovery of new structures is achieved.
Patent Information
- Application Number
- CN202510017610.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is cost-effective when drawing the phase boundary of the target structure in a nanoparticle self-assembly system, making it difficult to efficiently discover new self-assembly structures.
The reverse design method is adopted, combined with computer simulation, machine learning and global optimization algorithms, and the formation conditions of the target structure of complex self-assembly systems are efficiently obtained and phase boundaries are drawn by constructing candidate target structure data sets, nanoparticle self-assembly simulation, structural similarity prediction and target structure optimization.
The "experimental" cost and time of new structural discovery has been greatly reduced, and the efficiency and accuracy of self-assembled structural phase boundary drawing is improved, especially in three-dimensional space, which significantly shortens the calculation time.
Smart Images

Figure CN119964692A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a calculation method for reversely designing a phase boundary line of a nanoparticle self-assembly structure. Background Art
[0002] The self-assembly of nanoparticles (NPs) has great potential for fabricating complex functional metamaterials via a bottom-up approach. In nanoparticle self-assembly systems, the ordered structures of these assemblies are determined by the size, shape, composition, interactions of NPs, and environmental conditions.
[0003] In the simulation of nanoparticle self-assembly system, by introducing multiple free parameters to characterize the self-assembly system of nanoparticles, rich self-assembly structures can be discovered. However, the increase of independent parameters expands the parameter space, making manual search for the stable region (phase boundary) of the target structure time-consuming and tedious, thus complicating the discovery of new self-assembly structures.
[0004] Zhao et al. (ACS Macro Lett. 2021, 10, 598-602) autonomously constructed a block polymer phase diagram based on active machine learning. Although the active learning model can efficiently draw phase diagrams, it still requires a small sample data set to initialize the model. At the same time, the model is mainly used to draw phase diagrams. When our focus is on the phase region of a certain structure, this method is not effective enough. In addition, the automatic annotation of the structure in the model does not use more general standards.
[0005] How to develop a powerful inverse design solution to efficiently draw the phase boundaries of the desired target structure in the nanoparticle self-assembly system is a technical problem that needs to be solved urgently. Summary of the invention
[0006] (1) Technical issues to be solved
[0007] The present invention aims to solve the defects of the current self-assembly structure phase region mapping technology, which is high cost and low efficiency, and provides a method for calculating the phase boundary of the self-assembly structure of nanoparticles by reverse design. The present invention combines computer simulation, machine learning and global optimization algorithms to efficiently obtain the formation conditions of the target structure of the complex self-assembly system and draw the phase boundary of the target structure. The present invention uses experimental empirical structures and mathematical models to design candidate structures as data sets, which accelerates the discovery of new structures and greatly reduces the required "experimental" cost and "experimental" time.
[0008] (2) Technical solution
[0009] The present invention mainly solves the above technical problems through the following technical solutions.
[0010] The present invention provides a calculation method for reverse designing nanoparticle interface self-assembly structure, which comprises:
[0011] A candidate target structure construction module, used to construct a candidate structure data set through an empirical structure data set output by experimental simulation and a structure data set satisfying a mathematical geometric equation;
[0012] A nanoparticle self-assembly simulation module, which uses Monte Carlo to simulate the binary nanoparticle interface self-assembly system;
[0013] A structural similarity prediction module, which uses a convolutional neural network (CNN) as a proxy model to measure the similarity between the target structure and the search structure;
[0014] A target structure optimization result output module uses the free parameters of the self-assembly system as the parameter space, specifies a candidate structure as the initial structure of the simulation, and uses the BO algorithm to optimize the similarity function constructed by the CNN proxy model in the parameter space, and iterates until the fitness function is close to 1.0; the optimized similarity function is:
[0015]
[0016] Where Ω represents the parameter search space, P target (X) is the predicted probability component of the target structure of the CNN model at the free parameter vector X.
[0017] The present invention is suitable for self-assembly experiments or simulations of binary nanoparticles with complex free parameters, and is particularly easy to automate in simulation computing systems. In addition, it can also be used to solve the problem of drawing two-dimensional or three-dimensional phase boundaries of target structures. For simulation systems that can obtain diffraction patterns, CNN is a universal prediction and classification proxy model.
[0018] In the present invention, the candidate structure data set comes from the existing empirical structures in the field. In addition, the design of the candidate structure is based on the two-dimensional two-hard disk triangle stacking equation. Solving the equation can obtain a large number of binary nanoparticle self-assembly structure arrangements. According to the summary of the mathematical field, the particle coordinates of these ordered arrangement structures can be directly constructed as the candidate structure data set.
[0019] In the present invention, the construction of the diffraction pattern is a conventional calculation method in the art, which generally specifies a fixed box size and is obtained based on the following formula:
[0020]
[0021] in Fourier transform of density, r i is the position of particle i, v is the wave vector, and v is determined by the following formula:
[0022]
[0023] Among them, n x and n y are two integers in the interval [-64, 64], L is the length and width of the box, and the diffraction patterns considered in this work are built on a 128 × 128 grid.
[0024] The BO algorithm used in the present invention is a commonly used algorithm in optimization algorithms. The algorithm can converge to the stable region of the target structure within a shorter number of iteration steps, and the acquisition function uses the commonly used LCB acquisition function.
[0025] In the present invention, the target structure is searched in two-dimensional and three-dimensional parameter spaces respectively. The sample sampling of the BO algorithm is generally to establish a parameter vector in the specified parameter space according to the dimension of the parameter space; the free parameter can be ε and χ or introduce more free parameters.
[0026] In the present invention, according to the reverse design convention in the field, it is necessary to design an optimization function that describes the degree of similarity between the search sample structure and the target structure, and use the diffraction patterns calculated separately from the two nanoparticles to train the CNN proxy model as a function for evaluating the similarity. The higher the degree of similarity between the search sample structure and the target structure, the closer the CNN will give a prediction to 1.0.
[0027] In the present invention, the iteration termination condition of the optimization algorithm is that the similarity prediction is less than 10 -4 And it terminates after 5 iterations or the number of iterations reaches 50.
[0028] The active learning model in the present invention uses the label propagation algorithm as a conventional graph classification algorithm, and the minimum confidence method is the preferred uncertain sampling method. The specific algorithm flow is as follows:
[0029] First, the parameter space is divided into a finite grid, where all grid points constitute the sampling set G. LPA classifies all sample points in the parameter space according to the labeled samples. Specifically, we use the converged samples obtained from the inverse search as the initial sampling points. We use the vector p = [α, β, ...] to represent the probability distribution of a sample point belonging to different structures. LPA will derive the probability distribution P(G, p) of the labeled phase p of all unchecked points at position X in the parameter space. P(G, p) is used to calculate the uncertainty score S(G) through the marginal sampling estimator. The sample with the largest uncertainty score will be used as the next sampling sample. Next, the sample structure is obtained through the simulation module and labeled using the CNN model. Assuming that the sample structure is consistent with the target structure, it is marked as 1. Any other structure will be marked as 2 (non-target structure), which improves the algorithm to be suitable for drawing the target structure phase region. Repeat the above steps until the phase boundary prediction no longer changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 The logic block diagram of the method for calculating the phase boundary of the nanoparticle self-assembly structure in Example 1;
[0031] Figure 2 The representative structures in the experimental data set of the binary interface nanoparticle self-assembly structure of Example 1 and the data set constructed by the mathematical equation;
[0032] Figure 3 It is the reverse search process for the HST target structure in the two-dimensional parameter space;
[0033] Figure 4 It is the process of drawing the phase boundary of HST target structure in two-dimensional parameter space;
[0034] Figure 5 The results of drawing the phase boundary of the HST target structure in the three-dimensional parameter space;
[0035] Figure 6 It is the training framework for CNN models. DETAILED DESCRIPTION
[0036] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0037] The present invention provides a method for calculating the phase boundary of a nanoparticle self-assembly structure by reverse design. The method of the present invention first obtains stable sample parameters by reverse searching the target structure, and then quickly draws a high-precision phase boundary by an active learning method. Compared with directly calculating all structures and phase diagrams in the entire parameter space, the present invention designs a more purposeful scheme for a single structure, especially in three-dimensional space, and the amount of calculation is greatly reduced and the calculation time is greatly shortened.
[0038] Example 1
[0039] like Figure 1 As shown in the figure, it is the logical framework corresponding to the method of the present invention in Example 1. As can be seen from the figure, in Example 1 of the present invention, the target structure is first specified, and then random sampling is performed in the parameter space. Then, the stable structure of the binary nanoparticle self-assembly system at the sample point is calculated using a Monte Carlo-based computer simulation, and the three-dimensional coordinates of the structure are projected onto a plane parallel to the interface. The diffraction patterns of the two nanoparticles are calculated respectively and input into the CNN model to obtain a similarity prediction between the sample structure and the target structure. The Bayesian optimization algorithm is further combined, specifically including Gaussian process regression and acquisition function sampling to predict the sample point with the highest similarity, and the sample point is simulated and calculated again to obtain a stable structure and evaluate the similarity. This iteration is repeated until the similarity prediction result is obtained for five consecutive iterations>0.9999. The iteration stops, that is, it is considered that the sample parameters that can be formed in the search space for the target structure are obtained.
[0040] This embodiment is directed to a method for calculating the phase boundary of a nanoparticle self-assembly structure by reverse design, which includes a candidate target structure data set construction module; a nanoparticle interface self-assembly simulation module; a structure similarity prediction module; a target structure free parameter optimization module; and a phase boundary drawing module. The specific implementation process is as follows:
[0041] (1) In the candidate target structure construction module, the candidate structure dataset is constructed by obtaining the empirical structure dataset and the structure dataset satisfying the mathematical equation through experimental simulation.
[0042] Figure 2 This is a representative top view of each candidate structure in the candidate target structure data set of Example 1. By using different parameters to regulate the formation of the structure in the simulation system, the characteristics of the orderly arrangement of the structure can be quantified, which is the basis for constructing a candidate data set for reverse design. In addition to the structures obtained empirically in the past, the structures included in the figure also provide a set of structures that satisfy the mathematical equation of binary particle arrangement, including DTRI, HT and SF.
[0043] (2) In the nanoparticle interface self-assembly simulation module, we use a binary nanoparticle interface self-assembly system described by three dimensionless free parameters. The system mainly defines the interaction and energy calculation based on the following assumptions:
[0044] Assuming that the difference in surface energy Δ between the nanoparticles (NPs) and the two fluids is equal, the assembly thermodynamics of the NPs is completely described by three dimensionless free parameters: ε and χ.
[0045]
[0046] in is the radius ratio of the two nanoparticles A and B, ε is the relative strength of the interaction between the NPs relative to the surface tension γ, and χ is the ratio of the surface energy difference of NP A to the surface tension γ of the two fluids. Here, we fix R A = 3σ and use r to represent R B σ represents the intrinsic length scale of the system.
[0047] The model considers these two types of NPs as having a radius of R A and R B The interaction between particles is described by a shifted Lennard-Jones (LJ) pairwise potential function:
[0048]
[0049] ij = AA, AB and BB represent the interactions between the same and different NPs, d is the distance between NPs, σ represents the intrinsic length scale of the system, △ ij =R i +R j -σ is a distance displacement to prevent NPs from hard contact with each other. The strength of the interaction between other particles follows the Lorentz-Berthelot rule.
[0050] Assume that the difference in surface energy △γ between the two NPs and the two fluids is equal. The interfacial free energy of the two particles is described by two piecewise functions:
[0051]
[0052] where γ is the surface tension between the two fluids.
[0053] The total energy of the system is described by the interactions between particles and the interactions at the interfaces:
[0054]
[0055] where r Nrepresents the three-dimensional coordinates of NPs, and I is the function that determines the type of particle (NP A or B).
[0056] (3) In the structural similarity prediction module, the data set constructed in the candidate target structure construction module will be converted into diffraction patterns as a CNN model training set for training. The purpose of the CNN model is to predict the structural similarity P between the sample structure and the target structure in the reverse search process based on the pre-trained model. target (X). Figure 6 As shown in the figure, the CNN model uses dual-channel input data, passes through multiple convolutional layers and pooling layers, and finally inputs a Softmax layer for prediction and classification.
[0057] (4) In the target structure free parameter optimization module, we first need to specify the parameter space to be searched, that is, the free parameters describing the self-assembly system. In this simulation system, we use r, ε, and χ as optional parameters to construct the parameter space.
[0058] like Figure 3 (a) shows the search optimization process in the parameter space composed of r and χ as free parameters and HST as the target structure. The direction in which the color of the sample point becomes lighter is the direction of iteration, and the red sample point is the position of the sample point when convergence. First, a sample X is randomly initialized, and then a stable structure is obtained through simulation and the diffraction pattern is calculated. The diffraction pattern data is input into the CNN model to predict the similarity P between the sample and the target structure. target (X), and then predict the next sampling point X through Gaussian process regression and acquisition function next , and then simulate again to obtain samples.
[0059] And so on, continue to iterate until P target (X) When the number of iterations for five consecutive times is greater than 0.9999, the optimization process is stopped. At this time, it is considered that the sample points of the last iteration are highly similar to the target structure, that is, the free parameter vector that can stably form the target structure is found.
[0060] Figure 3 (b) shows the P during 50 iterations. target (X) Changes in the predicted values. BO can optimize to the stable region of the target structure in a shorter iteration cycle, minimizing the experimental cost and prior knowledge required to find the target structure.
[0061] (5) In the phase boundary drawing module, when the phase boundary of the target structure is drawn by active learning method, a small batch of prior samples are usually required to initialize the sampling. The target structure free parameter optimization module provides initialization samples for active learning by searching for samples of the target structure, and no prior samples are needed.
[0062] Figure 4 (a1) shows the initialization process of active learning. First, the parameter space is divided into a 26×21 grid. Then, the sampling is initialized, including a sample point obtained by reverse optimization and a sample point of a non-target structure (the non-target structure sample point can be obtained by random sampling or selecting P in the reverse search process). target (X) Sample points close to 0), LPA predicts the grid samples in the entire parameter space based on the known sample points by using the similarity measure (such as distance) between the graph and the sample points to obtain a probability distribution vector p = [α, β, ...], and then determines the next most uncertain sample point based on the minimum confidence method, and then obtains the structure at the sample point.
[0063] And so on, the iteration continues until LPA's prediction of samples in the entire parameter space no longer changes.
[0064] Figure 4 (b1) shows the uncertainty scores of all grid samples in the parameter space, with the red mark indicating the location of the next sampling point. Figure 4 (c1) is obtained by recalculating the probability distribution vector of the grid point samples after refining the grid points, and marking the probability distribution vector with the largest target structure probability component and the largest non-target structure probability component differently.
[0065] Figure 4 (a-c2) show the results after the convergence of active learning. At the 43rd iteration, LPA's prediction of samples in the parameter space no longer changes with the iteration process. At this time, the phase boundary of the HST structure is drawn.
[0066] Figure 5 The calculation method of the phase boundary of the inversely designed nanoparticle self-assembly structure is extended to the three-dimensional parameter space and a visualization scheme of the three-dimensional phase diagram is provided. Specifically, we specially mark the samples that have been sampled on the phase boundary and marked as the target phase or non-target phase, such as marking them with the values "1" and "2", and use the Tecplot360 drawing software to construct a facet between the values "1" and "2" to realize the visualization of the three-dimensional phase boundary.
[0067] From the above calculations, it can be seen that compared with the active learning algorithm that traverses the entire parameter space and takes the phase diagram as the target, the efficiency of drawing a single target phase boundary is improved by more than ten times and more than two times, respectively.
[0068] The above-listed implementations are for the convenience of the technical personnel in the field to better understand and use the present invention, and do not fully demonstrate the implementation methods of the present invention. Therefore, the present invention is not limited to the above embodiments, and the improvements and modifications made according to the description of the present invention are within the scope of the claims of the present invention.
Claims
1. A method for calculating the phase boundary of a reversely designed nanoparticle self-assembly structure, characterized in that: It includes: A module for constructing a candidate target structure data set, which constructs a candidate structure data set through empirical structures output by experimental simulation and geometric structures that satisfy mathematical equations; A nanoparticle self-assembly simulation module, which uses material simulation calculation methods to simulate the nanoparticle self-assembly system; A structural similarity prediction module, which uses a convolutional neural network (CNN) as a proxy model to measure the similarity between the target structure and the search structure; A target structure free parameter optimization module, which takes the free parameters of the self-assembly system as the parameter space, specifies a candidate structure as the initial structure of the simulation, and uses the Bayesian optimization (BO) method in the parameter space to optimize the similarity function constructed by the CNN proxy model, and iterates until the fitness function is close to 1.0; A target structure phase boundary drawing module uses an active learning method based on a graph algorithm combined with uncertain sampling to draw phase boundaries.
2. The calculation method according to claim 1, characterized in that: The construction of the candidate structure data set includes: on the one hand, obtaining a large number of coordinate files of stable structures based on simulation experiments; on the other hand, constructing potential candidate new structures based on particle arrangement rules; usually, due to inconsistent structural stability domains, the amount of data for each structure is not consistent. In order to ensure the effect of CNN model training, the data of each structure is copied according to the structure with the largest amount of data.
3. The calculation method according to claim 2, characterized in that: In the nanoparticle self-assembly system, the arrangement of particles satisfies certain rules. The rules can be recognized specifications, such as the point group and crystal system description of the unit cell, or they can be based on simple mathematical geometric equations or even generative models.
4. The calculation method according to claim 2, characterized in that: Designing new candidate structures through mathematical geometric equations helps to discover potential new structures in the parameter space during the search process.
5. The calculation method according to claim 1, characterized in that: The nanoparticle self-assembly system adopts common simulation methods, including Monte Carlo simulation, molecular dynamics simulation, etc.
6. The calculation method according to claim 1, characterized in that: The CNN proxy model is built based on the candidate structure dataset; the three-dimensional coordinates of the nanoparticles are first projected onto a two-dimensional plane, and then the real-space particle coordinates are projected into reciprocal space based on Fourier transform to obtain the diffraction pattern.
7. The calculation method according to claim 1, characterized in that: The free parameters are any system parameters of research significance, such as two or more interaction parameters that describe the thermodynamics of the self-assembly system or the properties of particles, and the parameter vector is a sample point in the parameter space formed by the free parameters.
8. The calculation method according to claim 1, characterized in that: The similarity function optimized by the global optimization algorithm is: Where Ω represents the parameter search space, P target (X) is the predicted probability component of the target structure of the CNN model at the free parameter vector X.
9. The calculation method according to claim 1, characterized in that: The global optimization algorithm is the Bayesian optimization algorithm (BO).
10. The calculation method according to claim 1, characterized in that: The graph algorithm used in active learning is the label propagation algorithm (LPA), and the sampling method is the minimum confidence method (LC). The principle of active learning is mainly: First, the parameter space is divided into a finite grid, where all grid points constitute the sampling set; LPA classifies all sample points in the parameter space according to the labeled samples; specifically, the converged samples obtained by the inverse search are used as the initial sampling points, and the label propagation algorithm is applied to derive the probability distribution of the labeled phases of all unchecked points at position X in the parameter space, and then the uncertainty score of the sample is calculated, and the sample with the largest uncertainty score will be used as the next sampling sample; the above steps are repeated until the phase boundary no longer changes.
11. The calculation method according to claim 9, characterized in that: The initialization of LPA uses the converged samples of the reverse search as initialization samples, without the need for a small batch of empirical sample datasets.
12. The calculation method according to claim 9, characterized in that: The structural markers of LPA are divided into target structures and non-target structures.
13. The calculation method according to claim 9, characterized in that: The structure tagging of LPA uses a trained CNN proxy model for discrimination. The CNN proxy model has two functions: one is to serve as an optimization function in reverse search, and the other is to complete the automatic labeling of samples in active learning.