Transformer Model for NADES Stability Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The development of deep eutectic solvents (DES) and natural deep eutectic solvents (NADES) is empirically driven, leading to time-consuming, labor-intensive, and expensive bench-top trials, with current computational tools requiring large databases and specialized knowledge to predict their stability and properties effectively.
Innovation Solution
A computerized system using a transformer-based neural network model that predicts the stability of eutectic mixtures through simplified molecular-input line-entry system (SMILES) representations, allowing for the use of small datasets and reducing training time, model complexity, and computational cost, by pre-training with unlabeled general chemical data and fine-tuning with labeled NADES data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If thermodynamic modeling (PC-SAFT) or atomistic modeling (DFT) is used to predict DES formation, then the formation mechanism can be explained, but specialized knowledge is required and statistically validated predictions of new mixtures are not yet achievable
Solution Approach 1:
The patent replaces complex thermodynamic and atomistic modeling approaches with a machine learning model that uses molecular descriptors and SMILES representations. This substitution eliminates the need for specialized knowledge in thermodynamic modeling while achieving statistically validated predictions of new DES mixtures, directly resolving the contradiction between prediction reliability and model complexity.
Solution Approach 2:
The patent creates a computational model that learns from existing DES data and generates predictions for new mixtures without requiring physical experimentation. This copying approach allows the system to achieve statistically validated predictions while avoiding the complexity of traditional modeling methods, as the model captures patterns from training data rather than requiring deep theoretical understanding.
2Reliability
If machine learning approaches with deep neural networks are used to predict solvent properties, then prediction capability is improved, but a substantial volume of data is required to train the model
Solution Approach 1:
The patent uses a simplified machine learning model that requires only a moderate volume of training data rather than the substantial data volumes needed by deep neural networks. By applying partial action (using a simpler model architecture), the system achieves adequate prediction accuracy without requiring millions of data entries, thus resolving the contradiction between prediction reliability and data volume requirements.
Solution Approach 2:
The patent changes the parameter of model complexity by using a simpler machine learning architecture instead of deep neural networks. This parameter change allows the system to achieve satisfactory prediction accuracy with significantly reduced data requirements, as the simpler model can be effectively trained on smaller datasets while still capturing the essential patterns in DES formation.
3Reliability
If bench-top trials of new mixtures are conducted to develop DES, then experimental validation is achieved, but the process is time-consuming, labor-intensive, and expensive
Solution Approach 1:
The patent performs preliminary computational screening and prediction of DES mixture stability before conducting any bench-top trials. By using the machine learning model to pre-identify promising candidates, the system reduces the number of experimental trials needed, thereby maintaining experimental validation reliability while significantly improving development speed and reducing costs.
Solution Approach 2:
The patent introduces a computational prediction model as an intermediary between theoretical DES design and experimental validation. This intermediary layer filters and prioritizes candidate mixtures before they undergo bench-top trials, reducing the overall number of experiments required while maintaining the reliability of experimental validation for the most promising candidates.
Data Source
AI summary
A computerized system for generating digital representations of natural deep eutectic solvents comprising: a training dataset having a plurality of natural deep eutectic solvents; a set of non-stable natural deep eutectic solvents generated by random variation of the number of components, random variation of the individual chemical component, random variation of the stoichiometric coefficient for each component and any combination thereof; a set of non-transitory computer readable instructions, that when executed by a process are adapted to: receive a training dataset having a set of compounds using the simplified molecular-input line-entry system and having a designation of stable (e.g., 1) or not stable (e.g., 0), pre-training a language model according to the training dataset, fine tuning the language model according to a subset of labeled DES data, applying a classifier, and, providing results in a textual format.


