Chemical Formulation Graphs for Accurate Attribute Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models for chemical formulations face challenges due to sparse and unbalanced data structures, making it difficult to accurately predict chemical product attributes without physically producing the products, which is time-consuming and expensive.
Innovation Solution
Represent chemical formulations as digital formulation graphs, using graph neural networks (GNNs) to process ingredient information, enabling more accurate predictions and recommendations of substitute products without actual production.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional machine learning models use sparse formulation data and ingredient descriptors, then model training is simpler, but prediction accuracy of chemical product attributes deteriorates
Solution Approach 1:
The patent transforms the flat, sparse formulation data into a multi-dimensional hierarchical structure using formulation graphs. These graphs organize ingredients and their properties across multiple levels (raw materials, intermediate products, final formulations), adding dimensional depth to the data. This hierarchical organization enables the GNN to capture complex relationships and interactions that were invisible in traditional sparse matrices, thereby improving prediction accuracy without requiring simpler data structures.
Solution Approach 2:
The patent creates a composite data structure by integrating multiple types of information (ingredient identities, formulations, descriptors, and relationships) into a unified formulation graph. This composite structure combines heterogeneous data sources into a coherent framework that the GNN can process effectively, allowing the model to leverage diverse information sources simultaneously to improve attribute prediction accuracy.
2Reliability
If physical production of chemical products is performed to obtain attribute data, then training data quality improves, but time and cost increase
Solution Approach 1:
The patent applies preliminary action by pre-processing and organizing formulation data into graph structures before model training. The formulation graphs are constructed in advance with all ingredient relationships, descriptors, and hierarchical connections established beforehand. This pre-organization enables the GNN to efficiently learn from structured data without requiring time-consuming physical product production, thereby maintaining training data quality while significantly reducing development time.
Solution Approach 2:
The patent creates digital copies of formulation data and ingredient properties in the form of formulation graphs. Instead of physically producing chemical products to obtain attribute data, the system uses digital representations that capture the essential characteristics and relationships. These digital copies serve as proxies for physical products, enabling virtual training and prediction that eliminates the need for actual product production while maintaining data reliability.
3Measurement precision
If ingredient descriptors are collected to improve model predictions, then prediction accuracy improves, but data collection cost and complexity increase
Solution Approach 1:
The patent creates a universal formulation graph framework that can accommodate multiple types of ingredient descriptors and formulation data through a single unified structure. The graph architecture is designed to be multi-functional, handling various data types (chemical properties, physical properties, safety data, etc.) within the same framework. This universality eliminates the need for separate data collection processes for different descriptor types, reducing overall collection complexity while maintaining comprehensive data for accurate predictions.
Solution Approach 2:
The formulation graph acts as an intermediary structure that mediates between raw ingredient data and the machine learning model. Instead of directly collecting and processing numerous individual ingredient descriptors, the graph structure organizes and transforms this data into a coherent format that the GNN can efficiently process. This intermediary representation simplifies data collection by providing a structured template and reduces the burden of manual descriptor gathering while preserving prediction accuracy.
Data Source
AI summary
Chemical formulations for chemical products can be represented by digital formulation graphs for use in machine learning models. The digital formulation graphs can be input to graph-based algorithms such as graph neural networks to produce a feature vector, which is a denser description of the chemical product than the digital formulation graph. The feature vector can be input to a supervised machine learning model to predict one or more attribute values of the chemical product that would be produced by the formulation without actually having to go through the production process. The feature vector can be input to an unsupervised machine learning model trained to compare chemical products based on feature vectors of the chemical products. The unsupervised machine learning model can recommend a substitute chemical product based on the comparison.


