ML Force Field Training With Off-Equilibrium Data Sampling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning force fields (MLFF) models are inaccurate when predicting material properties under real-world temperature and pressure conditions, as they are trained on near-equilibrium databases, failing to realistically simulate material behaviors.
Innovation Solution
The MatterSim MLFF model is trained using an off-equilibrium dataset generated by sampling chemical systems based on uncertainty values, narrowing the search space through subsets with high uncertainty, and performing ab initio simulations to enhance training efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If MLFF models are trained on near-equilibrium databases, then training efficiency is improved, but prediction accuracy under real-world temperature and pressure conditions deteriorates
Solution Approach 1:
The patent changes the training parameters by transitioning from near-equilibrium conditions to off-equilibrium conditions with varied temperature and pressure parameters. This allows the model to learn material behaviors under realistic operating conditions while maintaining training efficiency through targeted sampling of high-uncertainty regions in the parameter space.
Solution Approach 2:
The patent introduces dynamic training data generation where chemical systems are sampled at multiple temperature and pressure points to create off-equilibrium datasets. This dynamic approach enables the model to adapt to varying real-world conditions rather than being static to equilibrium states only.
2Reliability
If ab initio simulations are performed on all chemical systems, then prediction accuracy is improved, but computational time increases
Solution Approach 1:
The patent applies partial action by performing ab initio simulations only on a subset of chemical systems with high uncertainty values rather than all systems. This selective approach maintains prediction accuracy for critical regions while significantly reducing overall computational time through targeted sampling.
Solution Approach 2:
The patent implements feedback through uncertainty quantification that guides subsequent sampling decisions. Systems with higher prediction uncertainty are prioritized for ab initio simulation, creating a feedback loop that efficiently allocates computational resources to improve accuracy where most needed.
Data Source
AI summary
A computing system including one or more processing devices configured to obtain sets of ground-state force field data associated with equilibrium chemical systems. The one or more processing devices are further configured to compute ground-state uncertainty values of the sets of ground-state force field data and select a first subset of the equilibrium chemical systems that have respective ground-state uncertainty values above a ground-state uncertainty threshold. The one or more processing devices are further configured to compute off-equilibrium chemical systems by modifying a respective temperature and/or pressure of each equilibrium chemical system included in the first subset. The one or more processing devices are further configured to compute respective ab initio simulations of a second subset of the off-equilibrium chemical systems to obtain off-equilibrium force field data. The one or more processing devices are further configured to train a machine learning force fields model with the off-equilibrium force field data.


