Method of training an artificial intelligence bond order prediction model and bond order prediction method
By training an artificial intelligence bond number prediction model and utilizing generative adversarial networks and random forest models, the problem of insufficient research on the Si-O-Si bond number of polydimethylsiloxane was solved, achieving rapid and low-cost prediction and improving R&D efficiency.
Patent Information
- Application Number
- CN202411488238.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-10-24
AI Technical Summary
In existing technologies, the number of Si-O-Si bonds in polydimethylsiloxane is insufficient for studying the gas adsorption and separation performance of silicon-based nanofiltration membranes, and the construction of dense membranes requires a lot of computing power and time, resulting in low research and development efficiency.
By training an artificial intelligence bond number prediction model, using generative adversarial networks (GANs) for data augmentation, and combining a random forest model, the number of Si-O-Si bonds in a dimethylsiloxane dense film can be quickly predicted. This includes acquiring noisy data, constructing generated samples, molecular dynamics simulations, and loss function optimization to generate more data that conforms to the distribution of real samples.
It improves the prediction speed of Si-O-Si bond number, reduces computing resource consumption, improves R&D efficiency, and provides an efficient and low-cost prediction solution.
Smart Images

Figure CN119446318B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular, to a method for training an artificial intelligence bond number prediction model and a bond number prediction method. BACKGROUND
[0002] Polydimethylsiloxane (PDMS) membrane, also known as dimethylsiloxane membrane, has excellent performance in gas adsorption separation due to its high gas flux, good chemical stability, flexibility and thermal stability. Its unique Si-O-Si bond structure and hydrophobic methyl functional groups on its side chain endow it with excellent permeability, good hydrophobicity and organic affinity, so that the PDMS membrane can preferentially permeate small organic molecules, effectively improving the selectivity in gas adsorption separation. Therefore, the bond number of the Si-O-Si bond is likely to be a key parameter that determines the permeation characteristics of the nanofiltration membrane.
[0003] However, there is still a lack of systematic research on the bond number of polydimethylsiloxane Si-O-Si bond in the performance of silicon-based nanofiltration membrane gas adsorption separation. This is mainly because the membrane studied is a commercial product, and the exact properties of its selective layer have not been fully revealed. Even so, the bond number of polydimethylsiloxane Si-O-Si bond may be important in gas adsorption separation, but further research is needed to systematically evaluate its specific impact. At present, the bond number of Si-O-Si bond is obtained by statistical analysis of the constructed dense membrane, which requires a large amount of computing power and time. Therefore, if the bond number of Si-O-Si bond of PDMS membrane can be quickly predicted, the research and development efficiency can be improved and the subsequent research can be guided. SUMMARY
[0004] The purpose of the present disclosure is to provide a method for training an artificial intelligence bond number prediction model and a bond number prediction method to train a model that can predict the bond number of Si-O-Si bond of dimethylsiloxane dense membrane based on temperature, pressure, number of siloxane prepolymer structure model and number of dimethylsiloxane monomer structure model, improve the scheduling speed of predicting the bond number of Si-O-Si bond, reduce resource consumption such as computing power and time, and speed up research and development efficiency.
[0005] To achieve the above-mentioned purpose, the first aspect of the present disclosure provides a method for training an artificial intelligence bond number prediction model, comprising:
[0006] Obtain noise data, input the noise data into a preset GAN generator, and obtain generated samples;
[0007] Set a preset temperature, a preset pressure, a first number of siloxane prepolymer structure models, and a second number of dimethylsiloxane monomer structure models;
[0008] apply the preset temperature, the preset pressure, the first quantity and the second quantity in a preset molecular dynamics simulation calculation process to construct a dimethylsiloxane dense membrane structure model, and obtain a bond number of Si-O-Si bonds of the dimethylsiloxane dense membrane structure model;
[0009] The following two steps are repeatedly executed:
[0010] The preset temperature, the preset pressure, the first quantity, the second quantity and the corresponding bond number are combined into a real sample, the real sample and the generated sample are jointly input into the preset GAN discriminator to obtain a first prediction result, a first loss function of the preset GAN discriminator is calculated according to the first prediction result, the parameters of the preset GAN generator are fixed, and the parameters of the preset GAN discriminator are updated according to the first loss function;
[0011] The generated sample is input into the preset GAN discriminator to obtain a second prediction result, a second loss function of the preset GAN generator is calculated according to the second prediction result, the parameters of the preset GAN discriminator are fixed, and the parameters of the preset GAN generator are updated according to the second loss function;
[0012] When the two steps are repeatedly executed for a preset number of rounds, the two steps are stopped, and noise data is input into the trained preset GAN generator to obtain a plurality of groups of generated data;
[0013] The generated data and the real sample are used to train a random forest model to obtain a random forest model for predicting the bond number of Si-O-Si bonds of the dimethylsiloxane dense membrane structure model using temperature, pressure, the first quantity and the second quantity.
[0014] Optionally, the preset GAN generator is a fully connected neural network with 3 layers, the preset GAN discriminator is a fully connected neural network with 3 layers, and the activation function of the preset GAN generator and the preset GAN discriminator is a tanh activation function.
[0015] Optionally, the first loss function is calculated based on the following formula:
[0016]
[0017] wherein, L D is the first loss function, x is the real sample, z is the noise data, D is the preset GAN discriminator, G is the preset GAN generator, E is the expectation, p data is the probability distribution of the real sample, p z is the probability distribution of the noise data.
[0018] Optionally, the second loss function is calculated based on the following formula:
[0019]
[0020] wherein, L G is a second loss function, z is noise data, D is a preset GAN discriminator, G is a preset GAN generator, E is an expectation, p z is a probability distribution of the noise data.
[0021] Optionally, the preset number of rounds is 3000 rounds.
[0022] Optionally, the application of the preset temperature, the preset pressure, the first quantity and the second quantity in the preset molecular dynamics simulation calculation process to construct a dimethylsiloxane dense membrane structure model, and obtaining a bond number of Si-O-Si bond of the dimethylsiloxane dense membrane structure model, comprises:
[0023] A siloxane prepolymer structure model and a dimethylsiloxane monomer structure model are constructed by using Material Studio software, and geometric optimization is performed on the siloxane prepolymer structure model and the dimethylsiloxane monomer structure model respectively, and a pdb format file of the geometrically optimized siloxane prepolymer structure model and dimethylsiloxane monomer structure model is derived;
[0024] The pdb format file is opened in Packmol software, a first quantity of geometrically optimized siloxane prepolymer structure models and a second quantity of geometrically optimized dimethylsiloxane monomer structure models are randomly mixed to establish a target high polymer film initial structure model, and a pdb format file of the target high polymer film initial structure model is derived;
[0025] The pdb format file of the target high polymer film initial structure model is converted into an lmp format file by using Atomsk software, and the lmp file is converted into a data format file by using OVITO software;
[0026] The data format file is opened by using LAMMPS software, a target high polymer film initial structure model loaded is subjected to static minimization processing based on a conjugate gradient algorithm, then an NVT ensemble is selected, a Verlet-Velocity algorithm is used to integrate a motion equation, temperature equilibrium is performed at a first preset temperature for 50 ps, a time step is 0.1 fs, a Nose-Hoover algorithm is used to adjust temperature, and equilibrium simulation is performed at a second preset temperature for 100 ps, a standard relaxation time is 10 fs, and an initial dimethylsiloxane dense membrane structure model is obtained;
[0027] The NPT ensemble is selected, the pressure is set as a first preset pressure, the temperature is set as a third preset temperature, the simulation duration is 1 ns, the volume compression simulation calculation based on the molecular dynamics simulation is performed on the dimethylsiloxane dense membrane structure model, and a dimethylsiloxane dense membrane structure model with a stable structure is obtained.
[0028] In a second aspect of the present disclosure, a bond order prediction method is provided, and the method comprises:
[0029] The method of the first aspect is performed to obtain a trained random forest model for predicting the bond order of the Si-O-Si bond of the dimethylsiloxane dense membrane structure model;
[0030] A preset temperature, a preset pressure, a first quantity of siloxane prepolymer structure models, and a second quantity of dimethylsiloxane monomer structure models are obtained.
[0031] The preset temperature, the preset pressure, the first quantity, and the second quantity are input into the random forest model to obtain a predicted bond order of the Si-O-Si bond of the dimethylsiloxane dense membrane structure model constructed under the conditions of the preset temperature, the preset pressure, the first quantity, and the second quantity.
[0032] In a third aspect of the present disclosure, a computer-readable storage medium is provided, and the medium stores a computer program, which, when executed by a processor, implements the steps of the method of any one of the first aspect or the second aspect.
[0033] In a fourth aspect of the present disclosure, an electronic device is provided, and the electronic device comprises:
[0034] A memory storing a computer program;
[0035] A processor configured to execute the computer program in the memory to implement the steps of the method of any one of the first aspect or the second aspect.
[0036] Through the above technical solution, the preset GAN generator and the preset GAN discriminator of the generative adversarial network are trained to perform data augmentation on the data for constructing the polydimethylsiloxane dense membrane structure model, increase the data quantity, generate more generated data conforming to the real sample distribution, and overcome the challenge of data scarcity. Then, the random forest model is trained using the generated data and the real samples, the rich data can improve the training effect, the trained random forest model can effectively capture the nonlinear relationship between variables, and the bond order of the Si-O-Si bond under different conditions can be accurately predicted, thereby providing an efficient and low-cost solution for predicting the bond order of the Si-O-Si bond of the dimethylsiloxane dense membrane structure model.
[0037] Other features and advantages of the present disclosure will be made apparent from the following detailed description of a specific embodiment. BRIEF DESCRIPTION OF DRAWINGS
[0038] The accompanying drawings are included to provide a further understanding of the present disclosure and constitute a part of the specification, illustrate embodiments of the present disclosure and together with the detailed description help to explain the present disclosure, but do not limit the present disclosure. In the drawings:
[0039] Figure 1 is a flowchart of a method of training an artificial intelligence bond order prediction model according to an exemplary embodiment.
[0040] Figure 2 is a flowchart of step S103 according to an exemplary embodiment.
[0041] Figure 3 is a target polymer film initial structure model diagram according to an exemplary embodiment.
[0042] Figure 4 is a dimethylsiloxane dense film structure model diagram according to an exemplary embodiment.
[0043] Figure 5 is a flowchart of a bond order prediction method according to an exemplary embodiment.
[0044] Figure 6 is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0045] The specific embodiments of the present disclosure described below are intended to be illustrative of the present disclosure and are not to be limiting of the present disclosure. Unless otherwise defined, scientific and technical terms used in connection with the present disclosure shall have the meanings that are commonly understood by those of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and vice versa. The terms "comprises", "comprising", "includes", "including", "has", "having", "contains", "containing", or any other similar phrase are intended to mean "including but not limited to".
[0046] Molecular dynamics (MD) is a technique that combines knowledge from multiple disciplines. It is based on Newtonian mechanics and simulates the dynamic behavior of a molecular system. By extracting samples of different states of the molecular system, it can approximate the configuration integral of the system to evaluate the thermodynamic properties and other macroscopic aspects of the system. Researchers can explore the behavior of complex systems at the atomic and molecular scale with it, and it has been widely used in many fields such as chemistry, physics, and materials science.
[0047] Molecular dynamics (MD) simulations can provide deep insights into the behavior of materials at the microscopic level, but they are computationally expensive and generate limited data sets. As a result, researchers often face inefficiencies and data shortages when dealing with complex multivariate systems. Because the crosslinking process involves complex interactions among multiple variables, the use of classical methods to predict and optimize experimental conditions is inefficient. These complex nonlinear behaviors and interaction effects need to be gradually revealed and verified through extensive experimental work, which is costly and inefficient.
[0048] The rise of artificial intelligence (AI) and machine learning (ML) technologies offers new possibilities for the optimization of such complex multivariate systems. ML algorithms can extract hidden patterns and rules from large amounts of experimental data without explicit programming. This allows researchers to predict and optimize material performance more efficiently with fewer experimental inputs. In particular, in the crosslinking process of polymer materials, ML algorithms can help identify the variation of the number of important chemical bond structures under different conditions, providing a scientific basis for material design. By combining molecular dynamics simulation (MD) and machine learning (ML) technologies, researchers can theoretically simulate crosslinking behavior under different conditions and conduct experimental verification based on this, greatly reducing the number of experiments and research costs.
[0049] Figure 1 is a flowchart of a method for training an artificial intelligence bond number prediction model according to an exemplary embodiment, as shown in Figure 1 , the method comprises:
[0050] S101, obtaining noise data, inputting the noise data into a preset GAN generator to obtain generated samples.
[0051] S102, setting a preset temperature, a preset pressure, a first number of siloxane prepolymer structure models, and a second number of dimethylsiloxane monomer structure models.
[0052] S103, applying the preset temperature, the preset pressure, the first number, and the second number to a preset molecular dynamics simulation calculation process to construct a dimethylsiloxane dense film structure model, and obtaining the bond number of Si-O-Si bonds in the dimethylsiloxane dense film structure model.
[0053] The following two steps are repeated:
[0054] S104, forming real samples by combining the preset temperature, the preset pressure, the first number, the second number, and the corresponding bond numbers, inputting the real samples and the generated samples into a preset GAN discriminator to obtain a first prediction result, calculating a first loss function of the preset GAN discriminator according to the first prediction result, fixing the parameters of the preset GAN generator, and updating the parameters of the preset GAN discriminator according to the first loss function.
[0055] S105, input the generated sample into the preset GAN discriminator to obtain a second prediction result, calculate a second loss function of the preset GAN generator according to the second prediction result, fix the parameters of the preset GAN discriminator, and update the parameters of the preset GAN generator according to the second loss function.
[0056] S106, when the two steps are repeatedly executed for a preset number of rounds, stop repeating the two steps, input the noise data into the trained preset GAN generator to obtain a plurality of groups of generated data.
[0057] S107, training the random forest model with the generated data and the real sample to obtain a random forest model for predicting the bond number of the Si-O-Si bond of the dimethylsiloxane dense membrane structure model using temperature, pressure, the first number and the second number.
[0058] The preset GAN generator and the preset GAN discriminator jointly constitute a generative adversarial network (GAN) for data augmentation.
[0059] In step S101, random noise data conforming to a Gaussian distribution can be obtained, and then the noise data is input into the preset GAN generator, and the preset GAN generator outputs generated samples after calculation. The preset GAN generator can be a fully connected neural network with 3 layers, and the activation function is a tanh activation function.
[0060] After step S101 is executed, step S102 is executed. In step S102, the preset temperature can include a plurality of groups of temperature data, and the preset pressure can include a plurality of groups of pressure data, which are used for corresponding condition setting in the process of constructing a dimethylsiloxane (PDMS) dense membrane by molecular dynamics simulation. The preset temperature, the preset pressure, the first number of the siloxane prepolymer structure model and the second number of the dimethylsiloxane monomer structure model for predicting the molecular dynamics simulation can be determined according to the actual situation, so as to be applied when the molecular dynamics simulation is performed.
[0061] After step S102 is executed, step S103 is executed. In step S103, a PDMS dense membrane structure model is constructed by using a preset molecular dynamics simulation calculation process. For example, a PDMS membrane structure can be constructed by using a chemical modeling software such as Material Studio, Packmol, LAMMPS, and the like, and a stable membrane structure model is obtained by performing appropriate structure optimization and the like. The structure optimization operation used in the process is based on molecular dynamics calculation, and part of the calculation process needs to be run under certain temperature and pressure conditions. Therefore, the temperature and pressure conditions in the preset molecular dynamics simulation calculation process are set to the preset temperature and the preset pressure, and a first number of siloxane prepolymer structure models and a second number of dimethylsiloxane monomer structure models are used to construct the PDMS dense membrane structure model. After the PDMS dense membrane structure model is constructed, the bond number of the Si-O-Si bond of the dense membrane structure model can be counted for subsequent machine learning model training. The Si-O-Si bond refers to the backbone structure of the siloxane prepolymer in the dimethylsiloxane dense membrane, in which the silicon atoms are connected by oxygen bridge bonds. In this structure, the silicon atoms and oxygen atoms on the backbone of the prepolymer form chemical bonds with the silicon atoms and oxygen atoms at the ends of the dimethylsiloxane monomer, thereby constructing the overall cross-linked structure.
[0062] After step S103 is executed, step S104 can be executed. In step S104, the preset temperature, the preset pressure, the first number, the second number, and the bond number of the PDMS dense membrane constructed by using these conditions set in step S102 are taken as real samples, and the real samples and the generated samples obtained in step S101 are taken as training data to input a preset GAN discriminator, to obtain a first prediction result output by the preset GAN discriminator. For example, a number between 0 and 1 can be output to represent the probability that it is true, such as “1” indicating that the predicted sample is true, and “0” indicating that the predicted sample is false. Then, the parameters of the preset GAN generator are fixed, and the parameters of the preset GAN discriminator are trained and updated. The preset GAN discriminator can be a fully connected neural network with 3 layers, and the activation function thereof is a tanh activation function.
[0063] The parameters of the preset GAN generator can be kept unchanged, and a first loss function of the preset GAN discriminator is calculated according to the first prediction result. The first loss function can be calculated based on the following formula:
[0064]
[0065] wherein L D is the first loss function, x is the real sample, z is the noise data, D is the preset GAN discriminator, G is the preset GAN generator, E represents expectation, p data is the probability distribution of the real sample, and pz is a probability distribution of the noise data.
[0066] Specifically, G(z) in the formula is a generated sample generated by the preset GAN generator according to the noise data, D(G(z)) can represent a first preset result of the generated sample by the preset GAN discriminator, and D(x) is a first preset result of the real sample by the preset GAN discriminator.
[0067] The first loss function is calculated according to the first prediction result, for example, a cross-entropy loss function of the above formula is calculated, and then the gradient descent algorithm and the cross-entropy loss function can be used to update the parameters of the preset GAN discriminator. Then the generated sample and the real sample can be input into the preset GAN discriminator after the parameter update, and the above process is repeated to continue updating the parameters of the preset GAN discriminator until a preset condition is reached, such as the absolute value of the first loss function being less than a threshold, or the parameters of the preset GAN discriminator are updated a certain number of times.
[0068] After the step S104 is executed, the step S105 can be executed. In the step S105, the generated sample is input into the preset GAN discriminator to obtain a second prediction result output by the preset GAN discriminator, and then the parameters of the preset GAN discriminator are kept unchanged, and the parameters of the preset GAN generator are trained and updated.
[0069] The parameters of the preset GAN discriminator can be kept unchanged, and a second loss function corresponding to the preset GAN generator is calculated according to the second prediction result. The second loss function can be calculated based on the following formula:
[0070]
[0071] wherein, L G is the second loss function, z is the noise data, D is the preset GAN discriminator, G is the preset GAN generator, E is the expectation, and p z is a probability distribution of the noise data.
[0072] Specifically, G(z) in the formula is a generated sample generated by the preset GAN generator according to the noise data, and D(G(z)) is a second prediction result of the generated sample by the preset GAN discriminator.
[0073] The second loss function is calculated according to the second prediction result, for example, the above formula L GThe cross-entropy loss function can then be used with a gradient descent algorithm to update the parameters of the preset GAN generator. The noise data can then be input into the preset GAN generator with the updated parameters to obtain new generated samples, which are then input into the preset GAN discriminator, and the above process is repeated to continue updating the parameters of the preset GAN generator until a preset condition is reached, such as convergence of the second loss function value or the number of times the preset GAN generator parameters are updated reaches a preset number of rounds.
[0074] After step S105 is executed, the process can return to re-iterate steps S104 and S105. Performing step S104 once and step S105 once can be considered as one round. As the number of training rounds increases, the learning effect of the generative adversarial network formed by the preset GAN generator and the preset GAN discriminator will continue to improve, and the accuracy of data augmentation will also increase. Selecting an appropriate value of N is crucial to the result, where N represents the preset number of rounds for which steps S104 and S105 are repeated. N cannot be too large or too small. When N is constantly increasing, the generated samples generated by the preset GAN generator will become more and more close to the real samples, but this can also lead to overfitting. On the other hand, when the preset number of rounds N is small, the data generated by the preset GAN generator is easily identified by the preset GAN discriminator. Therefore, a suitable value of N must be carefully selected during training, and when the preset number of rounds N = 3000, the optimal form of data augmentation can be achieved, covering the initial data distribution.
[0075] Error analysis of the generative adversarial network. Generally, the errors of the preset GAN discriminator and the preset GAN generator are opposite. The more the preset GAN discriminator can distinguish between true and false, the greater the error of the preset GAN generator. Conversely, if the output of the preset GAN generator can make the preset GAN discriminator unable to distinguish, the error of the preset GAN generator is smaller, while the error of the preset GAN discriminator is larger. After the training starts, the errors of the preset GAN generator and the preset GAN discriminator change more stably. After a certain number of training rounds, the errors converge. Based on this, the optimal value of the preset number of rounds N in the parameter updating process can be determined to be 3000.
[0076] After step S105 is executed, step S106 can be performed. In step S106, if the execution of steps S104 and S105 is repeated for a preset number of rounds, the repeated execution of steps S104 and S105 can be stopped, and the noise data is input into the trained preset GAN generator to obtain a plurality of sets of generated data. At this time, the generated data generated by the preset GAN generator has a distribution close to that of the real samples and is difficult to be identified by the preset GAN discriminator, and can be used for subsequent training of the random forest model. In a possible implementation, after the generative adversarial network is trained, the number of samples in the training set can be expanded from the order of hundreds to the order of thousands. In the prediction of the bonding process of complex substances, other chemical bonds may be incorrectly predicted due to imperfect models or insufficient data. In order to improve the prediction accuracy, more diverse data can be generated through data enhancement techniques (such as generative adversarial network GAN), thereby improving the prediction accuracy.
[0077] After step S106 is executed, step S107 can be performed. In step S107, the generated data and the real samples are used to form a training set, and the random forest model is trained using the training set. The trained random forest model can be used to predict the number of Si-O-Si bonds in the PDMS membrane structure using temperature, pressure, first quantity, and second quantity. The random forest model uses a regression model, and the output value is continuous. 80% of the data in the training set is used to train the random forest model, and 20% of the data in the training set is used to test the training result. During training, multiple sample extractions can be performed on the training set. Each time, the bootstrap sampling method is used to extract training samples from the training set to obtain multiple sets of training samples, and each set of training samples is a subset of the training set. For this random forest model, the number of decision trees can be set to a value between 50 and 100, the maximum number of features for each decision tree can be set to a value between 2 and 4, and the minimum number of leaf nodes can be set to 2. By generating decision trees from each set of training samples, multiple decision trees are obtained, and the average value of the output of each decision tree is calculated to obtain the prediction result of the final random forest model.
[0078] Through the above technical solution, the preset GAN generator and the preset GAN discriminator of the generative adversarial network are trained to perform data enhancement on the data for constructing the polydimethylsiloxane dense membrane structure model, increase the data volume, and generate more generated data conforming to the distribution of the real samples to overcome the challenge of data scarcity. Then, the generated data and the real samples are used to train the random forest model, and the rich data can improve the training effect, so that the trained random forest model can effectively capture the non-linear relationship between variables and accurately predict the number of Si-O-Si bonds under different conditions, providing an efficient and low-cost solution for predicting the number of Si-O-Si bonds in the polydimethylsiloxane dense membrane structure model.
[0079] Optionally, referring to Figure 2 In step S103, the preset temperature, the preset pressure, the first quantity and the second quantity are applied to a preset molecular dynamics simulation calculation process to construct a dimethylsiloxane dense membrane structure model, and a bond number of Si-O-Si bonds of the dimethylsiloxane dense membrane structure model is obtained, including:
[0080] S1031, a siloxane prepolymer structure model and a dimethylsiloxane monomer structure model are constructed by using Material Studio software, and geometric optimization is performed on the siloxane prepolymer structure model and the dimethylsiloxane monomer structure model respectively, and a pdb format file of the geometrically optimized siloxane prepolymer structure model and dimethylsiloxane monomer structure model is derived.
[0081] S1032, the pdb format file is opened in Packmol software, the first quantity of geometrically optimized siloxane prepolymer structure models and the second quantity of geometrically optimized dimethylsiloxane monomer structure models are randomly mixed to establish a target high polymer membrane initial structure model, and a pdb format file of the target high polymer membrane initial structure model is derived.
[0082] S1033, the pdb format file of the target high polymer membrane initial structure model is converted into an lmp format file by using Atomsk software, and the lmp file is converted into a data format file by using OVITO software.
[0083] S1034, the data format file is opened by using LAMMPS software, a target high polymer membrane initial structure model loaded is subjected to static minimization processing based on a conjugate gradient algorithm, then an NVT ensemble is selected, a Verlet-Velocity algorithm is used to integrate a motion equation, a temperature equilibrium is performed at a first preset temperature for 50 ps, a time step is 0.1 fs, a Nose-Hoover algorithm is used to adjust a temperature, and an equilibrium simulation is performed at a second preset temperature for 100 ps, a standard relaxation time is 10 fs, and an initial dimethylsiloxane dense membrane structure model is obtained.
[0084] S1035, an NPT ensemble is selected, a pressure is set as a first preset pressure, a temperature is set as a third preset temperature, a simulation time length is 1 ns, a volume compression simulation calculation based on molecular dynamics simulation is performed on the dimethylsiloxane dense membrane structure model, a dimethylsiloxane dense membrane structure model with a stable structure is obtained, and a bond number of Si-O-Si bonds of the dimethylsiloxane dense membrane structure model is counted.
[0085] In step S1031, dimethylsiloxane monomers and tetraethyl silicate structures can be established in the Materials studio software first, four ethyl groups on the tetraethyl silicate are removed, and two dimethylsiloxane monomers are connected to the silicon and oxygen atoms of the tetraethyl silicate monomer to form a siloxane prepolymer structure, wherein the molecular formula of the dimethylsiloxane is -(Si(CH3)2-O-) n
[0086]
[0087] The molecular formula of the tetraethyl silicate monomer is Si(OEt)4(C8H 20 O4Si), and the structural formula is as follows:
[0088]
[0089] The van der Waals and electrostatic interactions of the dimethylsiloxane monomer structure and the siloxane prepolymer structure can be calculated by using the COMPASSIII force field and the Atom-based and Eward methods in the Forcite module of the Materials studio software, and based on the principle of system energy minimization, the geometric structure of the constructed model is finally optimized.
[0090] Specifically, the Geometry optimization task in the Forcite module can be used, the Smart algorithm is adopted, the Energy is 0.001 kcal / mol, the Force is Displacement is The maximum number of iterations is 500 steps, the force field is COMPASSIII force field, the charge is Forcefield assigned, the charge group is automatically assigned by the force field, the electrostatic term is Ewald summation method, the accuracy is 0.001 kcal / mol, the van der Waals term is Atom-based summation method, the truncation method is Cubic spline, and the cutoff radius is The bond width is Long range correction is used for geometric optimization calculation during the calculation process. After the geometric optimization is completed, the pdb format files of the siloxane prepolymer structure model and the dimethylsiloxane monomer structure model are exported for subsequent calculation by other software.
[0091] After step S1031 is executed, step S1032 can be performed. In step S1032, a second number of geometrically optimized dimethylsiloxane monomers and a first number of newly synthesized siloxane prepolymer molecules which are also geometrically optimized can be selected in the modeling software Packmol, as shown inFigure 3 These molecules are randomly placed in a cubic box with dimensions of 40 x 40 x 40 A to simulate the initial structure of the target polymer film as an amorphous mixture under actual conditions. See the cube composed of lines shown in FIG. 1, which is a three-dimensional box used to represent the core component of the crystal structure, the unit cell. It is the basic structural module that reveals the ordered and periodic arrangement of the crystal interior. This step allows the random distribution of dimethylsiloxane monomers and siloxane prepolymers in the same simulation environment, thereby simulating their interaction and structure formation process. Through the random placement function of Packmol, the initial structure of the material can be accurately established, providing appropriate starting conditions for subsequent molecular dynamics simulation to further explore the changes in its structure and properties. After the initial structure model of the target polymer film is constructed, its pdb format file is exported. Figure 3 After step S1032 is executed, step S1033 can be entered. In step S1033, the pdb file of the initial structure model of the target polymer film is converted into an lmp format using Atomsk software for further processing. Then, the lmp file is converted into a data file recognizable by LAMMPS using OVITO software. In this way, a seamless transition from construction to simulation of the initial model is ensured, making the molecular dynamics simulation proceed smoothly and providing a reliable foundation for subsequent simulation and analysis.
[0092] After step S1033 is executed, step S1034 can be entered. In step S1034, an in format file can be written in a program writing software according to the simulation format requirements and integrated into a script to form a file that can be automatically executed. This method not only simplifies the simulation process and improves efficiency, but also ensures that each step is automatically run according to the preset parameters and conditions, thereby ensuring the accuracy and stability of the simulation.
[0093] In the script, the calculation process can be written, including using the conjugate gradient algorithm (cg) to perform static energy minimization processing on the initial amorphous structure model containing the second number of dimethylsiloxane monomer molecules and the first number of newly synthesized siloxane prepolymer molecules. This step includes fine adjustment of the atomic coordinates of the initial model loaded in LAMMPS to find the most stable structure of the system. The conjugate gradient algorithm is a high-efficiency global optimization technique suitable for unconstrained optimization problems, known for its iterative simplicity and low storage requirements. This optimization process is crucial for ensuring the convergence and accuracy of subsequent molecular dynamics (MD) simulations.
[0094] In the script, the calculation process can be written, including using the conjugate gradient algorithm (cg) to perform static energy minimization processing on the initial amorphous structure model containing the second number of dimethylsiloxane monomer molecules and the first number of newly synthesized siloxane prepolymer molecules. This step includes fine adjustment of the atomic coordinates of the initial model loaded in LAMMPS to find the most stable structure of the system. The conjugate gradient algorithm is a high-efficiency global optimization technique suitable for unconstrained optimization problems, known for its iterative simplicity and low storage requirements. This optimization process is crucial for ensuring the convergence and accuracy of subsequent molecular dynamics (MD) simulations.
[0095] Specifically, the LAMMPS software is used to perform stability analysis and structure optimization on the initial structure model of the target polymer film using the ReaxFF force field. In the simulation process, the NVT ensemble is adopted, the Verlet-Velocity algorithm is used to integrate the motion equation, the time step is 0.1 fs, the Nose-Hoover algorithm is used to adjust the temperature, and the standard relaxation time is 10 fs. First, a temperature balance of 50 ps is performed at a first preset temperature, and the time step is 0.1 fs; then, a balance simulation of 100 ps is further performed at a second preset temperature to further stabilize the system. In the simulation process, the data file and trajectory file of the system are output through the LAMMPS command to record and analyze the dynamic evolution of the system. By executing an automated script, a new siloxane prepolymer structure is successfully constructed by repeatedly performing MD simulation at different temperatures under the NVT ensemble, the prepolymer is connected to the terminal silicon atom and oxygen atom in the dimethylsiloxane monomer structure through two oxygen active sites and two silicon active sites, a stable three-dimensional network structure is formed, and an initial dimethylsiloxane dense film structure model is obtained. The first preset temperature can be a temperature between 500K and 800K, and the second preset temperature can be a temperature between 298K and 330K.
[0096] After step S1034 is executed, step S1035 can be performed. In step S1035, the Material Studio software can be used to perform volume compression simulation calculation on the initial dimethylsiloxane dense film structure under the NPT ensemble, set the pressure and temperature to the first preset pressure and the third preset temperature, and the simulation time is 1 ns, and then a dimethylsiloxane dense film structure model with a stable structure as shown in Figure 4 is obtained, and the calculation process is based on molecular dynamics simulation calculation. The first preset pressure can be a pressure between 1kPa and 100kPa, and the third preset temperature can be a temperature between 300K and 800K.
[0097] The dimethylsiloxane dense film structure model constructed by the above scheme has good structural stability, so that the structure of the model can approach the actual physical state under the periodic boundary condition. This method optimizes the model through relaxation (energy optimization) and compression (adjusting the structure density) to make it as close as possible to the global energy minimum at the end of the simulation, thereby improving the accuracy of the simulation.
[0098] Figure 5 is a flowchart of a bond order prediction method according to an example embodiment, as shown in Figure 5 , the method comprises:
[0099] S501, execute the method as described in steps S101 to S107 to obtain a trained random forest model for predicting the bond number of Si-O-Si bond of the dense PDMS membrane structure model.
[0100] S502, obtain a preset temperature, a preset pressure, a first quantity of siloxane prepolymer structure model, and a second quantity of dimethylsiloxane monomer structure model.
[0101] S503, input the preset temperature, the preset pressure, the first quantity, and the second quantity into the random forest model to obtain the predicted bond number of Si-O-Si bond of the dense PDMS membrane structure model constructed under the conditions of the preset temperature, the preset pressure, the first quantity, and the second quantity.
[0102] In step S501, the method corresponding to steps S101 to S107 is executed to train a random forest model that can predict the bond number of Si-O-Si bond of the dense PDMS membrane structure model using temperature, pressure, the first quantity, and the second quantity.
[0103] After step S501 is executed, step S502 is executed. In step S502, one or more sets of calculation data can be set, each set of calculation data including a preset temperature, a preset pressure, a first quantity, and a second quantity as the condition parameters for simulating the construction of the PDMS dense membrane.
[0104] After step S502 is executed, step S503 is executed. In step S503, the preset temperature, the preset pressure, the first quantity, and the second quantity are input into the random forest model to obtain the predicted bond number of Si-O-Si bond, which is the predicted bond number of the PDMS dense membrane constructed based on the molecular dynamics simulation using the input temperature, pressure, and quantity condition parameters. The PDMS dense membrane is not actually constructed, which can save computing resources.
[0105] By this technical solution, the trained random forest model can efficiently estimate the bond number of Si-O-Si bond of the dense PDMS membrane structure under various conditions, so that the time-consuming molecular dynamics simulation does not need to be performed every time. This method reduces the dependence on molecular dynamics simulation and saves computing resources and time cost, and is especially suitable for situations where experimental parameters need to be frequently changed.
[0106] Figure 6 is a block diagram of an electronic device according to an exemplary embodiment. As shown in Figure 5As shown, the electronic device 700 can include a processor 701, a memory 702. The electronic device 700 can further include one or more of a multimedia component 703, an input / output (I / O) interface 704, and a communication component 705.
[0107] The processor 701 is configured to control overall operations of the electronic device 700 to complete all or part of the steps in the above-described method of training an artificial intelligence bond order prediction model and the method of bond order prediction. The memory 702 is configured to store various types of data to support operations of the electronic device 700, which can include, for example, instructions for operating any application or method on the electronic device 700, and application-related data, such as contact data, sent and received messages, pictures, audio, video, and the like. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 703 can include a screen and an audio component. The screen can be, for example, a touch screen, and the audio component is configured to output and / or input audio signals. For example, the audio component can include a microphone configured to receive external audio signals. The received audio signals can be further stored in the memory 702 or transmitted through the communication component 705. The audio component further includes at least one speaker configured to output audio signals. The I / O interface 704 provides an interface between the processor 701 and other interface modules, which can be a keyboard, a mouse, a button, and the like. The buttons can be virtual buttons or physical buttons. The communication component 705 is configured to perform wired or wireless communication between the electronic device 700 and other devices. The wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, and the like, or a combination of one or more of them, is not limited herein. Therefore, the corresponding communication component 705 can include a Wi-Fi module, a Bluetooth module, an NFC module, and the like.
[0108] In an exemplary embodiment, the electronic device 700 can be implemented by one or more Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), controller, microcontroller, microprocessor, or other electronic elements for performing the above-described method of training an artificial intelligence bond order prediction model and the bond order prediction method.
[0109] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implements the steps of the above-described method of training an artificial intelligence bond order prediction model and the bond order prediction method. For example, the computer-readable storage medium can be the above-described memory 702 including program instructions, which can be executed by the processor 701 of the electronic device 700 to complete the above-described method of training an artificial intelligence bond order prediction model and the bond order prediction method.
[0110] The preferred embodiments of the present disclosure are described in detail above with reference to the accompanying drawings, but the present disclosure is not limited to the specific details in the above-described embodiments. Within the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all belong to the protection scope of the present disclosure.
[0111] In addition, it should be noted that each specific technical feature described in the above-described specific embodiments can be combined in any appropriate manner without contradiction. In order to avoid unnecessary repetition, various possible combinations are not described again in the present disclosure.
[0112] Furthermore, any combination of various different embodiments of the present disclosure can also be made, as long as it does not deviate from the idea of the present disclosure, and it should also be considered as disclosed in the present disclosure.
Claims
1. A method for training an artificial intelligence bond number prediction model, characterized in that, The method includes: Acquire noise data and input the noise data into a preset GAN generator to obtain generated samples; Set the preset temperature, preset pressure, first number of siloxane prepolymer structure models, and second number of dimethylsiloxane monomer structure models; The preset temperature, preset pressure, first quantity, and second quantity are applied to a preset molecular dynamics simulation calculation process to construct a dense film structure model of dimethylsiloxane, and the number of Si-O-Si bonds in the dense film structure model of dimethylsiloxane is obtained. Repeat the following two steps: The preset temperature, preset pressure, first quantity, second quantity and corresponding number of bonds are used to form real samples. The real samples and generated samples are input into the preset GAN discriminator to obtain the first prediction result. The first loss function of the preset GAN discriminator is calculated based on the first prediction result. The parameters of the preset GAN generator are fixed. The parameters of the preset GAN discriminator are updated according to the first loss function. The generated sample is input into the preset GAN discriminator to obtain the second prediction result. The second loss function of the preset GAN generator is calculated based on the second prediction result. The parameters of the preset GAN discriminator are fixed, and the parameters of the preset GAN generator are updated according to the second loss function. When the two steps are repeated for a preset number of rounds, the two steps are stopped, and the noisy data is input into the trained preset GAN generator to obtain multiple sets of generated data. By training a random forest model using generated data and real samples, a random forest model is obtained that predicts the number of Si-O-Si bonds in the dense film structure model of dimethylsiloxane using temperature, pressure, the first quantity, and the second quantity.
2. The method according to claim 1, characterized in that, The preset GAN generator is a fully connected neural network with 3 layers, the preset GAN discriminator is a fully connected neural network with 3 layers, and the activation function of the preset GAN generator and the preset GAN discriminator is the tanh activation function.
3. The method according to claim 2, characterized in that, The first loss function is calculated based on the following formula: Among them, L D Here, x is the first loss function, z is the real sample, D is the preset GAN discriminator, G is the preset GAN generator, E is the expectation, and p is the expected value. data p represents the probability distribution of the real sample. z This represents the probability distribution of the noise data.
4. The method according to claim 2, characterized in that, The second loss function is calculated based on the following formula: Among them, L G This is the second loss function, where z is the noisy data, D is the preset GAN discriminator, G is the preset GAN generator, E is the expectation, and p is the expected value. z This represents the probability distribution of the noise data.
5. The method according to claim 2, characterized in that, The preset number of rounds is 3000 rounds.
6. The method according to claim 1, characterized in that, The application of the preset temperature, preset pressure, first quantity, and second quantity in a preset molecular dynamics simulation calculation process to construct a dense film structure model of dimethylsiloxane, and obtaining the number of Si-O-Si bonds in the dense film structure model of dimethylsiloxane, includes: Using Material Studio software, structural models of siloxane prepolymers and dimethylsiloxane monomers were constructed. Geometric optimization was performed on the siloxane prepolymer and dimethylsiloxane monomer structural models, and pdb format files of the geometrically optimized siloxane prepolymer and dimethylsiloxane monomer structural models were exported. Open the pdb format file in Packmol software, randomly mix the first number of geometrically optimized siloxane prepolymer structure models and the second number of geometrically optimized dimethylsiloxane monomer structure models to establish the initial structure model of the target polymer membrane, and export the pdb format file of the initial structure model of the target polymer membrane. The pdb format file of the initial structure model of the target polymer membrane was converted into an lmp format file using Atomsk software, and then the lmp file was converted into a data format file using OVITO software. The data format file was opened using LAMMPS software. The initial structural model of the target polymer membrane was statically minimized based on the conjugate gradient algorithm. Then, the NVT ensemble was selected, and the Verlet-Velocity algorithm was used to integrate the equation of motion. Temperature equilibrium was carried out for 50 ps at the first preset temperature with a time step of 0.1 fs. The Nose-Hoover algorithm was then used to adjust the temperature, and equilibrium simulation was carried out for 100 ps at the second preset temperature with a standard relaxation time of 10 fs to obtain the initial dimethylsiloxane dense membrane structural model. Select the NPT ensemble, set the pressure to the first preset pressure, the temperature to the third preset temperature, and the simulation duration to 1 ns. Perform volume compression simulation calculations based on molecular dynamics simulation on the dimethylsiloxane dense film structure model to obtain a dimethylsiloxane dense film structure model with a stable structure. Then, count the number of Si-O-Si bonds in the dimethylsiloxane dense film structure model.
7. A method for predicting the number of bonds, characterized in that, The method includes: The method described in claim 1 is performed to obtain a trained random forest model for predicting the number of Si-O-Si bonds in a dense film structure model of dimethylsiloxane. Obtain a first number of pre-set temperature, pre-set pressure, siloxane prepolymer structure models, and a second number of dimethylsiloxane monomer structure models; The preset temperature, preset pressure, first quantity, and second quantity are input into the random forest model to obtain the predicted number of Si-O-Si bonds in the dimethylsiloxane dense film structure model constructed under the conditions of the preset temperature, preset pressure, first quantity, and second quantity.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-7.
9. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Method for predicting organic matter PDMS film-air distribution coefficient through quantitative structure-activity relationship model
CN113722988A
Method and device for predicting synthesis accessibility of compound and training method thereof
CN118016186A