Machine Learning Solubility Estimation via Molecular Descriptors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for determining solubility of solutes in solvents are cumbersome and impractical for repeated experiments across various combinations, making it difficult to identify solutes and solvents with desired solubility properties.

Innovation Solution

A system and method utilizing machine learning models trained on chemical structures and solubility parameters to generate descriptors, which are then used to calculate solubility parameters, enabling quick and accurate estimation of solubility through the use of zero-dimensional, one-dimensional, or three-dimensional descriptors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If repeated experiments are conducted to detect solubility for various combinations of solutes and solvents, then accurate solubility data can be obtained, but the time and resource consumption increases significantly

Engineering Contradiction:
Improvesolubility data accuracyVSAvoidexperiment time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training a machine learning model using experimental solubility data from multiple solutes and solvents before actual solubility estimation is needed. The model is trained in advance on a dataset containing chemical structures and corresponding solubility parameters, so that when new solubility predictions are required, the system can quickly process new queries without conducting new experiments. This pre-computation approach resolves the contradiction by preparing the predictive capability beforehand, eliminating the need for time-consuming repeated experiments while maintaining accurate solubility data through the trained model.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies copying by using a machine learning model to replicate the relationship between chemical structures and solubility parameters that was originally established through experimental data. Instead of physically repeating experiments to obtain solubility data for new solute-solvent combinations, the system creates a computational copy of the experimental knowledge embedded in the trained model. This computational copy can rapidly predict solubility for any new combination by processing chemical structure inputs, thereby avoiding the need for actual repeated experiments while maintaining measurement precision through the model's learned patterns.

Inventive Principle:
Principle #26Copying

2Productivity

If machine learning models are used to estimate solubility, then the estimation speed increases, but the complexity of the system increases

Engineering Contradiction:
Improvesolubility estimation speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies the intermediary principle by introducing a machine learning model as a mediator between chemical structure data and solubility parameters. The model acts as an intermediate computational layer that processes chemical structure inputs (represented as molecular graphs or descriptors) and transforms them into predicted solubility outputs. This intermediary approach resolves the contradiction by automating the complex relationship analysis between molecular structure and solubility properties, thereby increasing productivity through rapid predictions while managing system complexity through the use of established machine learning frameworks and pre-trained models rather than requiring complex custom-built experimental apparatus.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20220415450A1System and method for estimating solubility
Publication Date: 2022.12.29 SAMSUNG ELECTRONICS CO LTD
  • US20220415450A1 patent drawing
  • US20220415450A1 patent drawing
  • US20220415450A1 patent drawing

AI summary

A method of estimating solubility includes obtaining input data representing a chemical structure of a target material; generating at least one descriptor based on the input data; obtaining at least one solubility parameter by providing the at least one descriptor to a machine learning model trained based on chemical structures and sample solubility parameters of sample materials; and calculating the solubility based on the at least one solubility parameter, wherein the at least one descriptor includes at least one of a zero-dimensional descriptor, a one-dimensional descriptor, a two-dimensional descriptor, or a three-dimensional descriptor, each representing the chemical structure of the target material.