Selection program, selection device, and selection method
By selecting intermediate structures with minimal energy difference from relaxed structures during DFT calculations as additional training data, the technique enhances the estimation accuracy of machine learning models for energy estimation, addressing the limitations of high computational costs in DFT calculations.
Patent Information
- Application Number
- JP2023212615
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-30
AI Technical Summary
The high computational cost of DFT calculations limits the generation of a large number of training data for DFT surrogate models, which in turn affects the estimation accuracy of these models.
A selection technique that identifies intermediate structures during DFT calculations, where the energy difference with the relaxed structure is smaller than a predetermined value, and uses these structures as additional training data for machine learning models.
This approach increases the number of training data for machine learning models, improving their estimation accuracy for energy from molecular structures without significantly increasing computational costs.
Smart Images

Figure 2025096733000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a selection technique for selecting training data of a machine learning model.
Background Art
[0002] Structure optimization is a technique for optimizing the molecular structure of a substance. In structure optimization, density functional theory (DFT) calculations may be used. DFT calculation is one of the first-principles quantum chemical calculations that approximately calculates the electron density of a molecule, and can efficiently calculate the electronic properties of a molecule.
[0003] Regarding DFT calculations, a training device for training a model of NNP (Neural Network Potential) is known (see, for example, Patent Document 1 and Patent Document 2). A method of ranking the energies of molecular crystals using DFT calculations is also known (see, for example, Patent Document 3).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0005] By performing structure optimization using DFT calculations, a relaxed structure can be obtained from the initial structure of the molecules of a substance via a plurality of intermediate structures. The relaxed structure is the optimized molecular structure of the substance. In structure optimization, the positions of each atom are repeatedly adjusted so that the total energy of the molecular structure is minimized and the molecular structure is stabilized. In the relaxed structure where the total energy is minimized, all atoms are arranged at equilibrium positions.
[0006] However, the computational cost of DFT calculations is very high, and DFT calculations take a long time. Therefore, as an alternative means with a lower computational cost and higher speed than DFT calculations, a DFT surrogate model may be used. The DFT surrogate model is a trained model generated by machine learning and is used to estimate the total energy from the molecular structure of a substance.
[0007] Since the estimation accuracy of the DFT surrogate model depends greatly on the number of training data, when the DFT surrogate model is trained using a small number of training data, the estimation accuracy becomes considerably lower than that of DFT calculations. Therefore, in order to improve the estimation accuracy, it is desirable to train the DFT surrogate model using a large number of training data.
[0008] However, the training data of the DFT surrogate model is generated by DFT calculations, and a combination of the relaxed structure obtained by DFT calculations and the total energy of the relaxed structure is used as the training data. Therefore, if the DFT calculations are repeated the same number of times as the number of training data in order to generate a large number of training data, the computational cost for generating the training data increases.
[0009] Note that such a problem occurs not only in DFT calculations but also when generating training data using various numerical calculations.
[0010] In one aspect, the present invention aims to increase the training data of a machine learning model for estimating energy from the structure of a substance.
Means for Solving the Problems
[0011] In one solution, the selection program causes a computer to execute the following processes.
[0012] The computer obtains the relaxed structure of the substance from the initial structure of the substance by numerical calculation. Among the plurality of intermediate structures of the substance obtained in the calculation process of obtaining the relaxed structure, the computer selects, as training data for training the machine learning model, an intermediate structure in which the difference between the energy of the intermediate structure and the energy of the relaxed structure is smaller than a predetermined value. The machine learning model estimates the energy of a predetermined structure from a predetermined structure of the substance.
Advantages of the Invention
[0013] According to one aspect, it is possible to increase the training data of the machine learning model for estimating the energy from the structure of the substance.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Modes for Carrying Out the Invention
[0015] Hereinafter, embodiments will be described in detail with reference to the drawings.
[0016] Figure 1 shows a functional configuration example of the selection device according to the embodiment. The selection device 101 in Figure 1 includes a calculation unit 111 and a selection unit 112.
[0017] Figure 2 is a flowchart showing an example of the selection process performed by the selection device 101 in Figure 1. First, the calculation unit 111 obtains the relaxed structure of the substance from the initial structure of the substance by numerical calculation (step 201).
[0018] Next, the selection unit 112 selects, as training data for training the machine learning model, an intermediate structure among a plurality of intermediate structures of the substance obtained in the calculation process of obtaining the relaxed structure, where the difference between the energy of the intermediate structure and the energy of the relaxed structure is smaller than a predetermined value (step 202). The machine learning model estimates the energy of a predetermined structure from a predetermined structure of the substance.
[0019] According to the selection device 101 in Figure 1, the training data for the machine learning model that estimates the energy from the structure of the substance can be increased.
[0020] Figure 3 shows a functional configuration example of the training device corresponding to the selection device 101 in Figure 1. The training device 301 in Figure 3 includes a calculation unit 311, a selection unit 312, a training unit 313, and a storage unit 314. The calculation unit 311 and the selection unit 312 respectively correspond to the calculation unit 111 and the selection unit 112 in Figure 1.
[0021] The storage unit 314 stores an initial data set 321. The initial data set 321 includes the initial data of the molecules of N substances (N is an integer greater than or equal to 1), and each initial data represents the initial structure of the molecule.
[0022] The calculation unit 311 performs structure optimization using DFT calculation, obtains a relaxed structure via a plurality of intermediate structures from each initial structure represented by each initial data included in the initial data set 321, generates structure information 322, and stores it in the storage unit 314. The structure information 322 includes, for each of the N pieces of initial data, a combination of each intermediate structure and the total energy of each intermediate structure, and a combination of the relaxed structure and the total energy of the relaxed structure. As the total energy of each structure, for example, the sum of the potential energies of the atomic system is used.
[0023] By performing structure optimization using DFT calculation, the relaxed structure of the molecule can be accurately obtained, and a plurality of intermediate structures leading to the relaxed structure can be obtained.
[0024] FIG. 4 shows an example of the change in the total energy in structure optimization. The horizontal axis represents the optimization step, and the vertical axis represents the total energy (electron volts) of the molecular structure. At each step, an intermediate structure and the total energy of that intermediate structure are calculated. The total energy decreases rapidly from steps 0 to 30 and then decreases gradually.
[0025] The selection unit 312 selects, as training data, a combination of each relaxed structure included in the structure information 322 and the total energy of each relaxed structure. Further, for each relaxed structure included in the structure information 322, the selection unit 312 selects, as training data, a combination of one or more intermediate structures having a total energy close to the total energy of the relaxed structure among the plurality of intermediate structures and the total energy of that intermediate structure. Then, the selection unit 312 generates a training data set 323 including the selected training data and stores it in the storage unit 314.
[0026] As a result, in addition to the relaxed structure, one or more intermediate structures are selected as training data, so the number of training data obtained by one DFT calculation increases and the training data is expanded.
[0027] As an intermediate structure having a total energy close to the total energy of the relaxed structure, for example, an intermediate structure in which the difference between the total energy of the intermediate structure and the total energy of the relaxed structure is smaller than a predetermined value is selected. As such an intermediate structure, at the time of structure optimization, the last intermediate structure obtained immediately before the relaxed structure may be selected. In addition, when there is no intermediate structure in which the difference is smaller than the predetermined value, only the relaxed structure is selected as the training data.
[0028] The training unit 313 performs machine learning for training the machine learning model before learning using the training data set 323, generates an estimation model 324, and stores it in the storage unit 314. The estimation model 324 is a learned model and estimates the total energy of the molecular structure from the molecular structure of the substance to be estimated. As the estimation model 324, a neural network, a linear regression model, a random forest, etc. are used.
[0029] By training the machine learning model using the training data set 323 with the added intermediate structure, the estimation accuracy of the generated estimation model 324 is improved compared to the case of using the training data set containing only the relaxed structure. The estimation model 324 is used to estimate the total energy of the molecular structure in various services such as material discovery, catalyst screening, and drug discovery.
[0030] FIG. 5 is a flowchart showing an example of the training process performed by the training apparatus 301 in FIG. 3. First, the calculation unit 311 obtains a relaxed structure from the initial structure represented by each initial data included in the initial data set 321 by performing structure optimization using DFT calculation (step 501), and generates structure information 322 (step 502).
[0031] Next, the selection unit 312 selects, as training data, the combination of each relaxation structure included in the structure information 322 and the total energy of each relaxation structure (step 503). Next, for each relaxation structure included in the structure information 322, the selection unit 312 selects, as training data, the combination of one or more intermediate structures having a total energy close to the total energy of the relaxation structure and the total energy of the intermediate structure (step 504). Then, the selection unit 312 generates a training data set 323 including the selected training data (step 505).
[0032] Next, the training unit 313 generates an estimation model 324 by training a machine learning model using the training data set 323 (step 506). Then, the training unit 313 calculates the estimation error of the estimation model 324 using the verification data (step 507).
[0033] Next, the training unit 313 compares the number of times of training executed with the number of epochs (step 508). If the number of times of training executed is less than the number of epochs (step 508, NO), the training device 301 repeats the processing after step 506. In this case, in step 506, the training unit 313 further trains the already trained estimation model 324.
[0034] When the number of times of training executed reaches the number of epochs (step 508, YES), the training unit 313 performs the processing of step 509, and the training device 301 ends the processing. In step 509, the training unit 313 selects, as the optimal estimation model 324, the estimation model 324 having the smallest estimation error among the estimation models 324 generated in each epoch.
[0035] Next, a specific example of the training process will be described. In the specific example of the training process, as the estimation model 324, the DFT surrogate model PaiNN of the Open Catalyst Project is used, and as the estimation error, the Mean Absolute Error (MAE) is used. The learning rate is 0.00001, and the number of epochs is 100. In this case, the estimation error ER of the estimation model 324 is calculated by the following formula.
[0036] ER=(1 / K)Σ|y(i)-yp(i)| (1)
[0037] K is an integer greater than or equal to 1 representing the number of verification data. Each verification data represents a molecular structure. y(i) (i = 1 to K) represents the correct value of the total energy of the molecular structure represented by the i-th verification data, and yp(i) represents the estimated value of the total energy estimated by the estimation model 324 from the i-th verification data. |y(i)-yp(i)| represents the absolute value of y(i)-yp(i), and Σ represents the sum over i = 1 to K.
[0038] Figure 6 shows an example of the number of training data. The vertical axis represents the number of training data. Rectangle 601 represents the number N of initial data included in the initial data set 321. N represented by rectangle 601 is 1780. Rectangle 602 represents the number of relaxed structures included in the structural information 322 generated by DFT calculation from the initial data set 321. The number of relaxed structures represented by rectangle 602 is the same as N, which is 1780.
[0039] Rectangle 603 represents the number of training data included in the training data set 323. In this example, among the intermediate structures included in the structural information 322, the intermediate structure obtained immediately before the relaxed structure is selected as the training data. Among the intermediate structures immediately before the relaxed structure, since the number of intermediate structures having a total energy close to the total energy of the relaxed structure is 1757, 1780 relaxed structures and 1757 intermediate structures are selected as the training data. Therefore, the number of training data represented by rectangle 603 is 3537.
[0040] Figure 7 shows an example of the estimation error ER. The vertical axis represents the estimation error ER. In this example, the number K of verification data is 30. Rectangle 701 represents the estimation error ER when the estimation model 324 is trained using only the 1780 relaxation structures represented by rectangle 602 in Figure 6 as training data. The ER represented by rectangle 701 is 0.5. Rectangle 702 represents the estimation error ER when the estimation model 324 is trained using the 3537 training data represented by rectangle 603 in Figure 6. The ER represented by rectangle 702 is 0.44.
[0041] Figure 8 shows an example of the change in the estimation error ER. The horizontal axis represents the epoch, and the vertical axis represents the estimation error ER. The broken line 801 represents the change in the estimation error ER when the estimation model 324 is trained using only the 1780 relaxation structures represented by rectangle 602 in Figure 6 as training data.
[0042] The solid line 802 represents the change in the estimation error ER when the estimation model 324 is trained using the 3537 training data represented by rectangle 603 in Figure 6. The ER represented by rectangle 702 in Figure 7 corresponds to the minimum value of the ER represented by the broken line 802.
[0043] From Figures 7 and 8, it can be seen that by adding the intermediate structure to the training data set 323 and training the estimation model 324, the estimation error ER becomes smaller than when training using only relaxation structures.
[0044] The configuration of the selection device 101 in Figure 1 is only an example, and some components may be omitted or changed according to the use or conditions of the selection device 101.
[0045] The configuration of the training device 301 in Figure 3 is only an example, and some components may be omitted or changed according to the use or conditions of the training device 301. For example, when the training of the machine learning model is performed by another device, the training unit 313 can be omitted.
[0046] The flowcharts of FIGS. 2 and 5 are merely examples, and some processes may be omitted or modified according to the configuration or conditions of the selection device 101 or the training device 301. For example, when the training of the machine learning model is performed by another device, the processes of steps 506 to 509 in FIG. 5 can be omitted. In step 501 of FIG. 5, the calculation unit 311 may perform structure optimization using another numerical calculation instead of the DFT calculation.
[0047] The change in the total energy shown in FIG. 4 is merely an example, and the total energy changes according to the molecule to be calculated. The number of training data shown in FIG. 6 and the estimation errors shown in FIGS. 7 and 8 are merely examples, and the number of training data and the estimation errors change according to the molecule to be calculated and the calculation conditions.
[0048] Equation (1) is merely an example, and the training device 301 may calculate other errors such as the root mean square error as the estimation error of the estimation model 324.
[0049] FIG. 9 shows an example of the hardware configuration of an information processing device (computer) used as the selection device 101 in FIG. 1 and the training device 301 in FIG. 3. The information processing device in FIG. 9 includes a CPU (Central Processing Unit) 901, a memory 902, an input device 903, an output device 904, an auxiliary storage device 905, a media drive device 906, and a network connection device 907. These components are hardware and are connected to each other by a bus 908.
[0050] The memory 902 is, for example, a semiconductor memory such as a ROM (Read Only Memory) or a RAM (Random Access Memory), and stores programs and data used for processing. The memory 902 may operate as the storage unit 314 in FIG. 3.
[0051] The CPU 901 (processor) operates as, for example, the calculation unit 111 and the selection unit 112 in FIG. 1 by executing a program using the memory 902. The CPU 901 also operates as the calculation unit 311, the selection unit 312, and the training unit 313 in FIG. 3 by executing a program using the memory 902.
[0052] The input device 903 is, for example, a keyboard, a pointing device, etc., and is used for inputting instructions or information from a user or an operator. The output device 904 is, for example, a display device, a printer, etc., and is used for making inquiries or giving instructions to the user or the operator and outputting the processing results. The processing results may be the training data set 323 or the optimal estimation model 324.
[0053] The auxiliary storage device 905 is, for example, a magnetic disk device, an optical disk device, a magneto-optical disk device, a tape device, etc. The auxiliary storage device 905 may be a hard disk drive or an SSD (Solid State Drive). The information processing device can store programs and data in the auxiliary storage device 905 and load them into the memory 902 for use. The auxiliary storage device 905 may operate as the storage unit 314 in FIG. 3.
[0054] The medium drive device 906 drives the portable recording medium 909 and accesses the recorded content. The portable recording medium 909 is a memory device, a flexible disk, an optical disk, a magneto-optical disk, etc. The portable recording medium 909 may be a CD-ROM (Compact Disk Read Only Memory), a DVD (Digital Versatile Disk), a USB (Universal Serial Bus) memory, etc. A user or an operator can store programs and data in the portable recording medium 909 and load them into the memory 902 for use.
[0055] Thus, the computer-readable recording medium for storing the programs and data used in the processing is a physical (non-transitory) recording medium such as the memory 902, the auxiliary storage device 905, or the portable recording medium 909.
[0056] The network connection device 907 is a communication circuit that is connected to a communication network such as a WAN (Wide Area Network) or a LAN (Local Area Network) and performs data conversion associated with communication. The information processing apparatus can receive programs and data from an external device via the network connection device 907, load them into the memory 902, and use them.
[0057] Note that the information processing apparatus does not necessarily need to include all the components in FIG. 9, and it is also possible to omit some components according to the use or conditions of the information processing apparatus. For example, when an interface with a user or an operator is not required, the input device 903 and the output device 904 may be omitted. When the portable recording medium 909 or the communication network is not used, the medium drive device 906 or the network connection device 907 may be omitted.
[0058] Although the disclosed embodiments and their advantages have been described in detail, those skilled in the art will be able to make various changes, additions, and omissions without departing from the scope of the invention clearly described in the claims.
[0059] Regarding the embodiments described with reference to FIGS. 1 to 9, the following additional remarks are further disclosed. (Additional Remark 1) Obtain the relaxed structure of the substance from the initial structure of the substance by numerical calculation, Among the plurality of intermediate structures of the substance obtained in the calculation process of obtaining the relaxed structure, select, as training data for training a machine learning model that estimates the energy of a predetermined structure from the predetermined structure of the substance, an intermediate structure in which the difference between the energy of the intermediate structure and the energy of the relaxed structure is smaller than a predetermined value. A selection program for causing a computer to execute the processing. (Appendix 2) Using the combination of the relaxation structure and the energy of the relaxation structure, and the combination of the selected intermediate structure and the energy of the selected intermediate structure as the training data, further causing the computer to execute a process of training the machine learning model, the selection program according to Appendix 1, characterized in that. (Appendix 3) The numerical calculation is a density functional theory calculation, and the selection program according to Appendix 1 or 2, characterized in that. (Appendix 4) A calculation unit that obtains the relaxation structure of the substance from the initial structure of the substance by numerical calculation, Among the plurality of intermediate structures of the substance obtained in the calculation process of obtaining the relaxation structure, an intermediate structure in which the difference between the energy of the intermediate structure and the energy of the relaxation structure is smaller than a predetermined value is selected as training data for training a machine learning model for estimating the energy of the predetermined structure from the predetermined structure of the substance. A selection unit, A selection device, characterized in that it comprises. (Appendix 5) The selection device according to Appendix 4, further comprising a training unit that uses the combination of the relaxation structure and the energy of the relaxation structure, and the combination of the selected intermediate structure and the energy of the selected intermediate structure as the training data to train the machine learning model, characterized in that. (Appendix 6) The numerical calculation is a density functional theory calculation, and the selection device according to Appendix 4 or 5, characterized in that. (Appendix 7) Obtaining the relaxation structure of the substance from the initial structure of the substance by numerical calculation, Among the plurality of intermediate structures of the substance obtained in the calculation process of obtaining the relaxation structure, an intermediate structure in which the difference between the energy of the intermediate structure and the energy of the relaxation structure is smaller than a predetermined value is selected as training data for training a machine learning model for estimating the energy of the predetermined structure from the predetermined structure of the substance. A selection method, characterized in that a computer executes the process. (Appendix 8) Using the combination of the relaxation structure and the energy of the relaxation structure, and the combination of the selected intermediate structure and the energy of the selected intermediate structure as the training data, the process of training the machine learning model is further executed by the computer. The selection method according to Supplementary Note 7 is characterized by this. (Supplementary Note 9) The numerical calculation is a density functional theory calculation. The selection method according to Supplementary Note 7 or 8 is characterized by this.
Explanation of Signs
[0060] 101 Selection device 111, 311 Calculation unit 112, 312 Selection unit 301 Training device 313 Training unit 314 Storage unit 321 Initial data set 322 Structure information 323 Training data set 324 Estimation model 601~603, 701, 702 Rectangle 801, 802 Broken line 901 CPU 902 Memory 903 Input device 904 Output device 905 Auxiliary storage device 906 Media drive device 907 Network connection device 908 Bus 909 Portable recording medium
Claims
1. Obtaining the relaxed structure of a substance from its initial structure by numerical calculation, selecting, as training data for training a machine learning model that estimates the energy of a predetermined structure from a predetermined structure of the substance, an intermediate structure among a plurality of intermediate structures of the substance obtained in the calculation process of obtaining the relaxed structure, where the difference between the energy of the intermediate structure and the energy of the relaxed structure is smaller than a predetermined value; A selection program for causing a computer to execute the process.
2. The selection program according to claim 1, further causing the computer to execute a process of training the machine learning model using, as the training data, a combination of the relaxed structure and the energy of the relaxed structure and a combination of the selected intermediate structure and the energy of the selected intermediate structure.
3. The selection program according to claim 1 or 2, wherein the numerical calculation is density functional theory calculation.
4. A calculation unit that obtains the relaxed structure of a substance from its initial structure by numerical calculation, a selection unit that selects, as training data for training a machine learning model that estimates the energy of a predetermined structure from a predetermined structure of the substance, an intermediate structure among a plurality of intermediate structures of the substance obtained in the calculation process of obtaining the relaxed structure, where the difference between the energy of the intermediate structure and the energy of the relaxed structure is smaller than a predetermined value; A selection device comprising the above.
5. Obtaining the relaxed structure of a substance from its initial structure by numerical calculation, selecting, as training data for training a machine learning model that estimates the energy of a predetermined structure from a predetermined structure of the substance, an intermediate structure among a plurality of intermediate structures of the substance obtained in the calculation process of obtaining the relaxed structure, where the difference between the energy of the intermediate structure and the energy of the relaxed structure is smaller than a predetermined value; A selection method characterized in that a computer executes the process.
Citation Information
Patent Citations
Method for energy ranking of molecular crystals using dft calculations and empirical van der waals potentials
US20070185695A1
Estimation device, training device, estimation method, training method, program, and non-transitory computer readable medium
WO2022260177A1
Training device, training method, program, and inference device
WO2022260179A1