Neural Network Prediction Model Using Molecular Descriptors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for predicting physical properties and yields of materials are inaccurate and require numerous experiments, with pre-training models struggling to maintain accuracy with small changes in material structure.
Innovation Solution
A method involving the training of a prediction model that includes obtaining molecular descriptors, pre-training a neural network using PCA to reduce dimensionality, and adjusting the network to match a target task using labeled training data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If pre-training models are used to predict physical properties and yields, then the prediction process can be automated, but the accuracy significantly changes even with small changes in material structure
Solution Approach 1:
The patent applies pre-training on large-scale molecular descriptor data before fine-tuning on specific task data. This preliminary action allows the model to learn general molecular structure-property relationships from abundant unlabeled data, then adapt to specific prediction tasks with limited labeled data, improving both automation and accuracy
Solution Approach 2:
The patent transforms molecular structures into molecular descriptors (parameter transformation) and uses dimensionality reduction techniques to change the parameter space. This allows the model to work with stabilized numerical representations rather than raw structural data, reducing sensitivity to small structural changes while maintaining prediction accuracy
2Quantity of substance
If numerous experiments are performed to predict physical properties and yields, then more training data can be obtained, but the process requires great number of human resources and time
Solution Approach 1:
The patent uses molecular descriptors as numerical copies or representations of molecular structures. Instead of performing physical experiments on actual molecules, the model works with computed descriptor vectors that capture essential molecular properties, enabling massive data generation without physical experimentation
Solution Approach 2:
The patent replaces the mechanical experiment system with a computational system. Molecular structures are converted to descriptors through computational methods, and predictions are made through neural network computations, substituting physical experimentation with information processing
3Loss of information
If molecular descriptors with high dimensionality are used for pre-training, then more molecular information can be captured, but the complexity of the model increases
Solution Approach 1:
The patent extracts the most important information from high-dimensional molecular descriptors using dimensionality reduction techniques like PCA. This extraction process identifies and retains the principal components that capture the majority of variance in the data, removing redundant dimensions and simplifying the model while preserving essential molecular information
Data Source
AI summary
Provided is a method of training a prediction model, the method including obtaining molecular descriptors of molecules based on a molecular database, pre-training a pre-training neural network based on the molecular descriptors, and adjusting the pre-training neural network such that the pre-training neural network matches a target task, by applying a training data set labeled corresponding to the target task to the pre-trained pre-training neural network.


