Selection Program for Machine Learning Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The high calculation cost and time required for Density Functional Theory (DFT) calculations limit the generation of a large number of training data for machine learning models, which in turn affects the estimation accuracy of these models.
Innovation Solution
A selection program is used to obtain a relaxed structure from an initial structure through numerical calculation and select intermediate structures with energy differences less than a predetermined value from the calculation process, expanding the training data set for machine learning models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If DFT calculation is performed to obtain training data for machine learning models, then the estimation accuracy of the models is improved, but the calculation cost and time increase significantly
Solution Approach 1:
The patent extracts only the necessary intermediate structure data from the DFT calculation process that meets specific energy difference criteria, rather than using all intermediate structures. This selective extraction reduces the amount of training data needed while maintaining model accuracy, thereby reducing calculation time and cost.
Solution Approach 2:
The patent introduces an energy difference parameter (comparing intermediate structures to relaxed structures) as a selection criterion. By changing the parameter for selecting training data from using all intermediate structures to using only those with energy differences below a threshold, it optimizes the balance between training data quality and calculation efficiency.
2Measurement precision
If DFT calculation is performed to obtain training data for machine learning models, then the estimation accuracy of the models is improved, but the calculation cost increases
Solution Approach 1:
The patent extracts only the necessary intermediate structure data from the DFT calculation process that meets specific energy difference criteria, rather than using all intermediate structures. This selective extraction reduces the amount of training data needed while maintaining model accuracy, thereby reducing calculation time and cost.
Solution Approach 2:
The patent introduces an energy difference parameter (comparing intermediate structures to relaxed structures) as a selection criterion. By changing the parameter for selecting training data from using all intermediate structures to using only those with energy differences below a threshold, it optimizes the balance between training data quality and calculation efficiency.
3Measurement precision
If a large number of training data are generated through DFT calculation, then the estimation accuracy of machine learning models is improved, but the calculation time and cost increase
Solution Approach 1:
The patent extracts only the necessary intermediate structure data from the DFT calculation process that meets specific energy difference criteria, rather than using all intermediate structures. This selective extraction reduces the amount of training data needed while maintaining model accuracy, thereby reducing calculation time and cost.
Solution Approach 2:
The patent introduces an energy difference parameter (comparing intermediate structures to relaxed structures) as a selection criterion. By changing the parameter for selecting training data from using all intermediate structures to using only those with energy differences below a threshold, it optimizes the balance between training data quality and calculation efficiency.
Data Source
AI summary
A non-transitory computer-readable recording medium stores a selection program for causing a computer to execute processing including: obtaining a relaxed structure of a substance from an initial structure of the substance by numerical calculation; and selecting an intermediate structure of which a difference between energy of the intermediate structure and energy of the relaxed structure is less than a predetermined value, from among a plurality of the intermediate structures of the substance, obtained in a calculation process to obtain the relaxed structure, as training data used to train a machine learning model that estimates energy of a predetermined structure from the predetermined structure of the substance.


