A method of optimizing a training set of machine learning potential functions
By integrating Boltzmann energy distribution theory with farthest-point sampling techniques, the training set for machine learning potential functions is optimized, solving the problem of poor sampling results in existing technologies. This enables efficient and accurate dataset construction, improving the accuracy and efficiency of molecular dynamics simulations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BOHAI UNIV
- Filing Date
- 2024-10-31
- Publication Date
- 2026-04-10
AI Technical Summary
Existing machine learning potential function training set sampling methods ignore subtle differences in real-world physical contexts, resulting in poor training performance. Furthermore, increasing the size of the training set is costly and has limited effectiveness.
By integrating Boltzmann energy distribution theory with farthest point sampling technology, and through probability distribution calculation and farthest point sampling method, the most representative and information-rich subset in the dataset is automatically identified and extracted.
This improves the representativeness and information content of the training set, enhances the accuracy and speed of the machine learning potential function, and reduces the complexity and cost of data collection.
Smart Images

Figure CN119479843B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of force field development, in particular to a method for optimizing a training set of a machine learning potential function. BACKGROUND
[0002] Molecular dynamics simulation, as a powerful tool to reveal the microscopic behavior of matter and its properties, has shown great potential and significant achievements in many scientific research fields. However, the accuracy of this technology has been limited by force field models, and the limitations of traditional empirical potentials often constrain the accuracy of simulation results. With the rise of machine learning potential functions, the field of molecular dynamics simulation has undergone a revolutionary change. By training machine learning models to predict atomic interactions, machine learning potential functions not only break through the accuracy bottleneck of traditional empirical potentials, but also achieve high accuracy comparable to ab initio molecular dynamics, while maintaining the computational efficiency of molecular dynamics. This breakthrough greatly enhances the practical value of molecular dynamics simulation, enabling researchers to explore the dynamic behavior of molecules and materials with unprecedented accuracy and speed, thereby accelerating the development of fields such as materials science, biochemistry, and drug design. The training core of machine learning potential functions lies in constructing a high-quality training set, and the key lies in how to accurately capture and generalize to a wide range of physical configurations using limited and representative configurations.
[0003] Currently, many training set sampling methods ignore the subtle differences in real physical situations, resulting in poor performance of the trained machine learning potential functions in actual applications. In addition, although increasing the size of the training set aims to improve performance, this approach often comes with high costs and limited effectiveness. SUMMARY
[0004] To overcome the shortcomings of the prior art, the purpose of the present application is to provide a method for optimizing a training set of a machine learning potential function, which solves the problem of poor performance of traditional data sampling techniques by fusing Boltzmann energy distribution theory and farthest point sampling technology, and realizes automatic identification and extraction of the most representative and information-rich subset in the data set.
[0005] To achieve the above-mentioned purpose, the present application provides the following solutions:
[0006] A method for optimizing a training set of a machine learning potential function, comprising:
[0007] constructing a machine learning potential function;
[0008] performing molecular dynamics simulation using the machine learning potential function to obtain an original database;
[0009] performing single-point energy calculation on the original database using first-principle software to obtain configuration data;
[0010] Target parameter analysis and data rejection are performed on the configuration data to obtain preprocessed data; the target parameter analysis includes energy analysis, force analysis and bit force distribution analysis;
[0011] Atomic average energy calculation is performed on the preprocessed data to obtain atomic average energy data;
[0012] Probability distribution calculation is performed on the atomic average energy data by using Boltzmann energy distribution to obtain probability distribution data;
[0013] The overall configuration sampling number is determined according to the probability distribution data;
[0014] Interval division is performed on the atomic average energy data to obtain a plurality of energy intervals, and configuration calculation is performed on the atomic average energy data of each energy interval according to the probability distribution data and the overall configuration sampling number to obtain the in-interval configuration sampling number;
[0015] The configurations in different energy intervals are marked according to the in-interval configuration sampling number, and Cartesian distance descriptors of the configurations in different energy intervals are calculated by using the farthest point sampling technology to obtain difference data;
[0016] According to the difference data, the farthest point sampling method is used to select the configuration with the largest Cartesian distance descriptor distance in each energy interval to obtain a training set, and the configurations in each energy interval except the training set are integrated into a test set.
[0017] Preferably, the calculation formula of the probability distribution data is: wherein, wherein, P(E) is the probability distribution data; E is the average particle energy of the atomic average energy data; k is the Boltzmann constant; T is the temperature; Z is the partition function; E i is the average particle energy of the i-th energy interval.
[0018] Preferably, the construction technology of the machine learning potential function includes ab initio molecular dynamics, stretching and compression simulation and density functional theory single-point energy calculation method.
[0019] The present application discloses the following technical effects:
[0020] The present application provides a method for optimizing machine learning potential function training set, which solves the problem of poor effect of traditional data sampling technology by fusing Boltzmann energy distribution theory and farthest point sampling technology, and realizes automatic identification and extraction of the most representative subset with the most abundant information in the data set. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described below only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained based on these drawings without creative labor.
[0022] Figure 1 The optimization machine learning potential function training set flowchart provided for the embodiments of the present application;
[0023] Figure 2 The implementation flowchart provided for the embodiments of the present application;
[0024] Figure 3 The energy force potential force distribution diagram provided for the embodiments of the present application, Figure 3 (a) is the energy distribution diagram in the data set, Figure 3 (b) is the force distribution diagram in the data set, Figure 3 (c) is the potential force distribution diagram;
[0025] Figure 4 The screening result diagram provided for the embodiments of the present application, Figure 4 (a) is the distribution statistical diagram of the selected AlN-wz configuration in the total configuration space; Figure 4 (b) is the distribution statistical diagram of the selected AlN-rs configuration in the total configuration space
[0026] Figure 5 The fitting principle diagram provided for the embodiments of the present application;
[0027] Figure 6 The loss comparison diagram of different methods for constructing training set provided for the embodiments of the present application, Figure 6 (a) is the comparison diagram of the total loss function, Figure 6 (b) is the root mean square error comparison diagram of the energy of the training set and the test set of the three methods respectively, Figure 6 (c) is the root mean square error comparison diagram of the force of the training set and the test set of the three methods respectively, Figure 6 (d) is the root mean square error comparison diagram of the potential force of the training set and the test set of the three methods respectively;
[0028] Figure 7 The energy and volume comparison diagram provided for the embodiments of the present application;
[0029] Figure 8 The phase transition diagram provided for the embodiments of the present application. DETAILED DESCRIPTION
[0030] With reference to the accompanying drawings: clearly and completely describe the technical solutions in the embodiments of the present application, apparently, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.
[0031] The purpose of the present application is to provide a method for optimizing machine learning potential function training set, by fusing Boltzmann energy distribution theory and farthest point sampling technology, solving the problem of poor effect of traditional data sampling technology, realizing automatic identification and extraction of the most representative subset and the most information-rich subset in the data set.
[0032] In order to make the above-mentioned purposes, characteristics and advantages of the present application more apparent and easy to understand, the present application will be further described in detail below in combination with the drawings and specific embodiments.
[0033] Figure 1 The present application provides a flowchart of optimizing machine learning potential function training set, Figure 2 The present application provides an implementation flowchart, as shown in Figure 1 and Figure 2 The present application provides a method for optimizing machine learning potential function training set, comprising:
[0034] Step 100: constructing machine learning potential function;
[0035] Step 200: using machine learning potential function for molecular dynamics simulation to obtain original database;
[0036] Step 300: using first principle software to calculate single-point energy of the original database to obtain configuration data;
[0037] Step 400: performing target parameter analysis and data rejection on the configuration data to obtain preprocessed data; the target parameter analysis includes energy analysis, force analysis and potential distribution analysis;
[0038] Step 500: performing atomic average energy calculation on the preprocessed data to obtain atomic average energy data;
[0039] Step 600: using Boltzmann energy distribution to calculate probability distribution of the atomic average energy data to obtain probability distribution data;
[0040] Step 700: determining the overall configuration sampling number according to the probability distribution data;
[0041] Step 800: interval division is performed on the atomic average energy data to obtain a plurality of energy intervals, and atomic average energy data in each energy interval is calculated according to the probability distribution data and the number of configuration samples in the whole configuration to obtain the number of configuration samples in the interval;
[0042] Step 900: configurations in different energy intervals are marked according to the number of configuration samples in the interval, and Cartesian distance descriptors of the configurations in different energy intervals are calculated by using the farthest point sampling technique to obtain difference data;
[0043] Step 1000: according to the difference data, the farthest point sampling method is used to select the configuration with the largest Cartesian distance descriptor distance in each energy interval to obtain a training set, and configurations other than the training set in each energy interval are integrated into a test set.
[0044] Specifically, the calculation formula of the probability distribution data is: Wherein, Wherein, P(E) is the probability distribution data; E is the average particle energy of the atomic average energy data; k is the Boltzmann constant; T is the temperature; Z is the partition function; E i is the average particle energy of the i-th energy interval.
[0045] Preferably, the construction technology of the machine learning potential function includes ab initio molecular dynamics, stretching and compression simulation, and density functional theory single-point energy calculation method.
[0046] Optionally, the embodiment provides a method for constructing a data set (molecular dynamics part): molecular dynamics simulation is used to obtain configurations of various stretching and compression models, for example, a box is enlarged or reduced by a certain proportion to explore the configuration space of molecules under different density and pressure conditions, or the positions of all atoms are perturbed to simulate the configurations that the molecules may reach after a plurality of molecular dynamics evolutions, and then high-precision density functional theory single-point energy calculation is used to calculate the configuration energy, force and potential information of the configurations, so as to be added to the data set.
[0047] Further, the configuration data obtained by high-precision single-point energy calculation using first-principle software is sorted and uniformly managed, and the data is stored in a file for subsequent processing.
[0048] Reference Figure 3 (a) to Figure 3 (c), Figure 3 (a) is an energy distribution diagram in the data set, wherein the horizontal coordinate is the average energy of each atom, and the vertical coordinate is the frequency; Figure 3 (b) is a force distribution diagram in the data set, wherein the horizontal coordinate is the average interaction force of each atom, and the vertical coordinate is the frequency;Figure 3 (c) is a force distribution map, where the abscissa is the average force of each atom, and the ordinate is the frequency; analyze the energy, force, and force distribution of the configuration. By observing the distribution of these physical quantities, identify those configurations with abnormally high energy or force. Based on professional knowledge and experience, screen out unreasonable configurations and consider deleting them from the training set to avoid affecting the training effect.
[0049] Specifically, the screened data set is subjected to accurate calculation of the average energy of the atoms. The principle of Boltzmann energy distribution is applied to quantitatively evaluate the probability of each configuration appearing in the current configuration space. This provides a probability basis for subsequent data set optimization, ensuring the statistical representativeness of the samples in the training set.
[0050] Further, based on the predetermined number of samples, the configuration space is divided into detailed regions. Through scientific calculation, the number of configurations to be extracted in each region is determined to achieve uniform distribution and comprehensive coverage of the samples.
[0051] Preferably, in each divided region, the farthest point sampling technique is used to reduce the data dimension and extract key features. Subsequently, the farthest point sampling method is used to carefully select configurations that are farthest apart in the feature space, ensuring that the selected samples can fully reflect the diversity of the configuration space. Finally, these selected configurations are defined as the training set, and the remaining configurations are used as the test set, completing the optimization selection of the training set.
[0052] Specifically, in this embodiment, a machine learning potential function for AlN is developed. First, initial configurations are obtained through ab initio molecular dynamics, stretching and compression simulations, etc., and then high-precision ab initio molecular dynamics is used to accurately calculate these configurations to obtain energy, force, and force information, which is summarized in a file. Taking AlN-wz and AlN-rs as examples, ab initio molecular dynamics simulation is performed, and scripts are used to perturb and scale the initial configurations to build a preliminary machine learning potential function. Subsequently, molecular dynamics simulation is performed using the potential energy to simulate stretching strain and other simulations for the two configurations, obtaining a large number of configurations, and then performing high-precision single-point energy calculation using VASP to establish a configuration database.
[0053] Optionally, after determining the quantity, the configurations in the database need to be regionally divided according to the average energy of the atoms. The probability of each configuration in a region is obtained by accumulating the probability of each atom in that region, which will be used as the probability of sampling in that interval. Based on the required number of configurations, the number of configurations to be sampled in each energy interval is calculated.
[0054] Specifically, in the configuration interval that needs to be sampled, the farthest point sampling technique is applied, the Cartesian distance descriptor distance between configurations is analyzed, and the farthest point sampling is performed to select several configurations with the farthest distance. This process is applied to all divided configuration spaces, and the selected configurations are aggregated into the training set, and the configurations not selected are put into the test set. The screening results are shown in Figure 4 (a) and Figure 4 (b). Figure 4 (a) is a histogram of the distribution of the selected configurations of AlN-wz in the total configuration space, where the abscissa is the average energy of each atom, and the ordinate is the frequency; Figure 4 (b) is a histogram of the distribution of the selected configurations of AlN-rs in the total configuration space, where the abscissa is the average energy of each atom, and the ordinate is the frequency.
[0055] Specifically, the final machine learning potential developed by the training set optimized by the method of the present example is used as the force field of the molecular dynamics simulation of AlN. The calculation shows that:
[0056] First, the method of the present embodiment can achieve more scientific division of the data set when dividing the database. For neural fitting, more scientific data distribution is important. Using traditional manual screening of the training set or simply selecting configurations with larger differences will cause the fitting result to deviate to a more unreliable direction, as shown in Figure 5 It can be seen that if the data on both ends of the straight line is uniformly distributed, the most suitable straight line can be fitted, but if the data on one side is more, the fitted straight line will deviate downward. In nanoscale simulation, a slight deviation of force will have a serious consequence. If only manual selection of data is used in ab initio molecular dynamics calculation, the above situation will inevitably occur. Using the method of the present embodiment will make the sampling of the data set closer to the natural rule.
[0057] Second, after multiple rounds of stretching, perturbation, and strain addition, the data set may be very large. After various manual screening, the data set becomes very complex while being rich. This may seriously slow down the speed of researchers developing machine learning potential functions. Using the method, manual division of the training set is not needed. Only the data obtained by various methods needs to be put into a database, and the method is directly used for screening, which greatly improves the speed of developers.
[0058] Third, using the method can improve the accuracy and breadth of training machine learning potential functions. As shown in Figure 6 (a) to Figure 6 (d), Figure 6Loss comparison diagram of different methods provided by the embodiment of the present application for constructing a training set, Figure 6 (a) is a comparison diagram of a total loss function, wherein the abscissa is the number of iterations, and the ordinate is the loss; Figure 6 (b) is a comparison diagram of the root mean square error of energy of the training set and the test set of each of the three methods, wherein the abscissa is the number of iterations, and the ordinate is the root mean square error of energy; Figure 6 (c) is a comparison diagram of the root mean square error of force of the training set and the test set of each of the three methods, wherein the abscissa is the number of iterations, and the ordinate is the root mean square error of force; Figure 6 (d) is a comparison diagram of the root mean square error of potential energy of the training set and the test set of each of the three methods, wherein the abscissa is the number of iterations, and the ordinate is the root mean square error of potential energy. The accuracy of the machine learning potential function trained by the present method is compared with that of other methods. The other methods are as follows: first, the training set and the test set are the same without distinguishing between the training set and the test set; second, the farthest point sampling technique is used to obtain the training set by sampling the farthest points together. It can be seen that the total loss of the present method is greatly reduced compared with the other two methods, and the accuracy of the energy of the test set and the training set is the highest. Although the method using the farthest point sampling technique has a smaller error in the energy of the test set, it is because the data difficult to fit is put into the training set, and the fitting of the data of the training set is not good. The same is true for the fitting of force. The machine learning potential function trained by using the farthest point sampling technique has a poor fitting for the training set, but the fitting using the present method has a high accuracy for both the training set and the test set.
[0059] Fourth, the calculation results of the properties of AlN-wz and AlN-rs using the machine learning potential function fitted by the present method are as shown in Figure 7 , wherein the abscissa is the average volume of atoms, and the ordinate is the average energy of atoms. The relationship between energy and volume is calculated by molecular dynamics simulation and ab initio molecular dynamics for two configurations respectively. It can be seen that the results calculated by the machine learning potential function are well fitted with the results calculated by DFT.
[0060] Fifth, the machine learning potential function fitted by the present method has a good description of phase transition, as shown in Figure 8 , which well shows the transition between AlN-rs and AlN-wz. It can be seen that the machine learning potential function constructed in the example can well describe the phase transition, wherein the dark part represents the atoms of the AlN-rs structure, and the light atoms appearing after stretching represent the atoms of the AlN-wz structure.
[0061] Further, the method of the embodiment improves the precision of machine learning potential function training by scientifically screening a database, and has significant advantages and application potential compared with traditional manual screening.
[0062] The beneficial effects of the present application are as follows:
[0063] The present application solves the problem of poor effect of traditional data sampling technology by fusing the Boltzmann energy distribution theory and the farthest point sampling technology, realizes automatic identification and extraction of the most representative subset and the subset with the most abundant information in the data center, and improves the speed and reliability of data acquisition.
[0064] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments.
[0065] The principles and implementation modes of the present application are described by applying specific examples herein, and the above description of the embodiments is only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, the specific implementation modes and application ranges will be changed according to the idea of the present application. In conclusion, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A method of optimizing a machine learning potential function training set, comprising: The application relates to a method for constructing a machine learning potential function. The method comprises the following steps: a molecular dynamics simulation is performed by using the machine learning potential function to obtain an original database; The original database refers to configurations obtained by using molecular dynamics simulation to obtain various stretching and compression models, and specifically comprises the following steps: a box is enlarged or reduced by a certain proportion to explore configurations of molecules under different density and pressure conditions, or a certain perturbation is performed on positions of all atoms to simulate configurations of molecules after multiple molecular dynamics evolutions; single-point energy calculation is performed on the original database by using first-principle software to obtain configuration data; the configuration data refers to energy, force and potential force information of the configuration; target parameter analysis and data rejection are performed on the configuration to obtain pretreatment data; the target parameter analysis comprises energy analysis, force analysis and potential force distribution analysis; energy, force and potential force distribution of the configuration are analyzed, abnormal high configurations of energy or force are identified by observing distribution of the physical quantities, unreasonable configurations are screened based on professional knowledge and experience, and the unreasonable configurations are deleted; atomic average energy calculation is performed on the pretreatment data to obtain atomic average energy data; probability distribution calculation is performed on the atomic average energy data by using a Boltzmann energy distribution to obtain probability distribution data; the overall configuration sampling quantity is determined according to the probability distribution data; interval division is performed on the atomic average energy data to obtain a plurality of energy intervals, and configuration calculation is performed on the atomic average energy data in each energy interval according to the probability distribution data and the overall configuration sampling quantity to obtain an in-interval configuration sampling quantity; configurations in different energy intervals are marked according to the in-interval configuration sampling quantity, and Cartesian distance descriptors of the configurations in different energy intervals are calculated by using a farthest point sampling technology to obtain difference data; the farthest point sampling method is used to select configurations with the largest Cartesian distance descriptors in each energy interval according to the difference data to obtain a training set, and configurations in each energy interval except the training set are integrated into a test set.
2. The method of claim 1, wherein, The calculation formula of the probability distribution data is: ; wherein, ; wherein, is the probability distribution data; E is the average particle energy of the atomic average energy data; k is the Boltzmann constant; T is the temperature; Z is the partition function; is the average particle energy of the i-th energy interval.
3. The method of claim 1, wherein, The construction technology of the machine learning potential function comprises ab initio molecular dynamics, stretching and compression simulation and a density functional theory single-point energy calculation method.
Citation Information
Patent Citations
Machine learning potential energy surface model construction method
CN117273169A
Computer-aided neural network force field training routines for molecular dynamics computer simulations
DE102022210046A1