Evolutionary algorithm for inverse design of polyurethane formulations
The integration of a genetic algorithm with machine learning methods addresses the inefficiencies in chemical formulation design by generating optimized formulations with novel combinations, thereby reducing costs and time through iterative evolution.
Patent Information
- Application Number
- PCT/US2025/018322
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-22
- Filing Date
- 2025-03-04
- Publication Date
- 2025-09-25
AI Technical Summary
Existing chemical formulation design methods are labor-intensive and costly due to the exponential increase in combinations and experiments needed to map the chemical parameter space, leading to large libraries that require extensive validation, often resulting in significant time and material costs.
A genetic algorithm (GA) combined with machine learning (ML) is used to generate prospective chemical formulations by defining target properties, applying a fitness function, and iteratively evolving formulations through crossover and mutation operations, reducing the need for extensive experimental validation.
This approach generates optimized formulations with novel component combinations efficiently, increasing the relevance of predicted formulations and reducing the time and cost associated with experimental validation.
Smart Images

Figure IMGF000010_0001 
Figure IMGF000010_0002 
Figure IMGF000011_0001
Abstract
Description
[0001] EVOLUTIONARY ALGORITHM FOR INVERSE DESIGN OF POLYURETHANE FORMULATIONS
[0002] Field
[0003] Embodiments relate to computer-based methods for generating prospective chemical formulations given an input of one or more desired target product properties. Methods may also include validating a prospective formulation and modifying a chemical process accordingly.
[0004] Introduction
[0005] The design of chemical formulations across industries often involves combining various chemical components to generate a product having one or more target product properties through labor intensive methodologies that can involve serial and / or parallel high throughput experimentation. However, as the scale and complexity of the chemical formulations increase, the number of combinations and experiments needed to map the chemical parameter space increases exponentially. This has led to the development of computer-based approaches that utilize machine learning models trained using volumes of historical formulation data that generate predicted target properties from an input prospective formulation. The use of machine learning (ML) modeling approaches can enable users to design and customize prospective formulations,
[0006] To gain full utility of ML modeling approaches, users may design and customize prospective formulations that can then be surveyed as meeting desired target properties using a “forward” ML model (predicting material properties from a formulation). On the other hand, “inverse” ML modeling may be used to generate new prospective formulations that meet various application-specific constraints, and may include novel and unique combinations of components predicted to meet the target product properties. However, the number of possible formulation combinations can be virtually infinite due to the permutation and combination of components, and are often unconstrained by field-specific and empirical knowledge from the field. This can result in the generation of large libraries that require validation through various experiments to verify that all design criteria are satisfied, resulting in significant costs in time and materials.
[0007] Summary
[0008] In one aspect, methods may include generating a prospective formulation library of chemical formulations, including: (a) defining one or more target product properties; (b) preparing an initial formulation library curated from historical data comprising one or more historical product properties relevant to the prospective formulation library; (c) applying a fitness function to determine difference between the one or more target product properties and the one or more predicted product properties, and discarding formulations from the initial formulation library that do not meet the fitness function; (d) using a genetic algorithm based methodology to develop an evolved formulation library from the initial formulation library including: (i) selecting one or more parent formulations from the initial formulation library having a desired closeness to the target product properties according to the fitness function; (ii) converting each of the one or more parent formulations to a bit vector; (iii) generating one or more offspring formulations from the one or more parent formulation bit vectors using mutation and / or crossover operations; and (iv) collecting the one or more offspring formulations to create the evolved formulation library; (e) generating descriptors for the one or more offspring formulations in the evolved formulation library; (f) analyzing the evolved formulation library using a trained machine learning module to calculate one or more predicted product properties; (g) applying the fitness function to the evolved formulation library to determine difference between the one or more target product properties and the one or more predicted product properties; and (h) based on the determined difference in (g); (i) collecting a subset of the evolved formulation library to generate the prospective formulation library having a desired closeness to the target product properties; or (ii) repeat steps (d) to (g) until the one or more target product properties desired for the prospective formulation library or stopping criteria are met.
[0009] Brief Description of The Drawings
[0010] Figs. 1 and 2 are flow diagram illustrating an overall method of using genetic algorithm methodologies to develop one or more prospective chemical formulations.
[0011] Fig. 3 is a flow diagram illustrating a method of training a machine learning module to estimate one or more target product properties for a prospective chemical formulation.
[0012] Fig. 4 is a flow diagram illustrating a method to integrate a trained machine learning module capable of generating one or more prospective product properties into a genetic algorithm methodology.
[0013] Fig. 5 illustrates a schematic drawing of a cloud-based server cluster in accordance with one example of the present disclosure.
[0014] Detailed Description
[0015] Methods disclosed herein enable the generation of hypothetical formulation libraries, which are generated using genetic algorithm (GA) methodologies that mimic natural evolutionary processes. In addition, GA methodologies convert formulated systems into bit vectors that are then manipulated systematically to produce novel formulations within the desired constraints. GA methodologies may be combined with ML-based validation techniques that may provide optimized formulations, reducing the time and cost for experimental validation. In some cases, methods may include inverse design approaches to generating a hypothetical formulation library containing one or more prospective chemical formulations derived from input of one or more desired target product properties (effectively in real time). Methods disclosed herein may also be combined with user interface strategies that enable management of large volumes of formulations and facilitate search based on desired property ranges.
[0016] Methods disclosed herein leverage historical formulation data to train ML models for novel formulation design, including forward design models (predicting material properties from composition) and inverse design models (predicting material composition from desired properties) to generate predicted formulation libraries of increased size. However, while inverse modeling has the ability to generate large predicted formulation libraries, the constituent chemical components and properties tend to mirror those in the historical data sets used to train the ML module. The low variance among predicted formulations then limits the appearance of novel component combinations, particularly new formulations having target product properties, components, and / or concentration ranges that are sparsely represented in the training dataset.
[0017] Methods disclosed herein generate hypothetical formulation libraries using GA methodologies incorporating directed evolution techniques of crossover and / or mutation to modify an initial formulation library. Such directed evolution techniques can generate high levels of with higher probabilities of accessing undersampled regions of parameter space. Methods disclosed herein may increase: (i) the relevancy of prospective formulation libraries generated from smaller numbers of historical formulation data; and (ii) the number of predicted formulations having unique component combinations meeting the target product properties by some defined threshold.
[0018] As used herein, a “formulation” refers to combination of components (e.g., polymer- forming compositions, reactant mixtures, blends, etc.) for a particular application (e.g., polymer formation, chemical treatment, article construction).
[0019] As used herein, a “formulation library” or “library” is a collection of two or more formulations. Libraries may include data tagged to each formulation including names of formulation components, component classes (e.g., monomer, surfactant, catalyst), component subclasses (e.g., polyol, anionic surfactant, gelling catalyst), concentration limits, and the like.
[0020] As used herein, “component” refers to chemical species including, but not limited to, monomers, prepolymers, catalysts, additives, and the like.
[0021] As used herein, “component class” refers to classes of component types categorized by one or more similar properties (e.g., chemical structure, function, molecular weight, polarity, etc.). Component class construction may be dependent on specific application, which may be derived from historical data and / or general knowledge. In some cases, classes may be constructed based on existing types of raw materials, or subdivided into classes by chemical type. For example, chemical formulation components for a polyurethane foam may be divided into classes for polyol, isocyanate, catalysts, surfactants, blowing agents, additives, and the like. As used herein, “process conditions” (e.g., atmospheric pressure, variable pressure foaming, relative humidity, overpacking percentage, etc.) refers to an expression describing the process conditions affecting the properties.
[0022] As used herein, “descriptor” (alternatively “specific descriptor”) refers to an expression describing a correlation within a chemical system (e.g., polymer system) that can provide additional information and / or generalizations regarding system behaviors. Examples of descriptors for polymer systems (e.g., polyurethane-specific descriptors, polyurethane foamspecific descriptors) include monomer mw, water content, catalytic activity, polymer chain entanglement, and the like. Descriptors herein may be calculated from component information and concentration using various physics-based approaches and models of target product properties.
[0023] As used herein, a “forward ML model” (or “ML model”) is a machine learning (ML) model trained on historical formulation libraries, which are used to generate predicted properties given a formulation input. Suitable ML models may have varied architectures and algorithms to predict different target product properties, and may include modules that combine multiple models or average outputs from multiple models. ML models may be incorporated into a method, for example, to calculate the mean absolute difference in the fitness function.
[0024] As used herein, “historical formulation data” includes component names, component quantities, target product properties obtained from prior research and experimental results; and optionally descriptors.
[0025] As used herein, “target product property” refers to a property associated with a unique chemical formulation (e.g., chemo-rheology, foam density, hardness, modulus, etc.) that is selected based on desired user input for a given product application.
[0026] As used herein, “variable parameter” refers to a feature of a machine learning module and / or model that can be varied during training (e.g., chemical species and / or class, concentration, descriptor, etc.).
[0027] Methods disclosed herein train ML modules by genetic algorithm (GA) to generate novel predicted chemical formulations having novel component combinations over that present in an initial training data set. GA methodologies utilize processes analogous to evolutionary processes in biological systems that select candidates (e.g., formulations, chemical components, and component classes) based on defined criteria, which are then used to produce novel “evolved” formulations by crossover and mutation operations. Methods are described below with respect to formulations for producing polymer systems, such as polyurethanes and polyurethane foams; however, it is envisioned that the techniques may be applied to other polymer systems and materials. Methods disclosed herein may utilize inverse modeling in which an initial formulation library is curated from historical data, input into a ML model to predict one or more target product properties. The predicted target product properties are then input into a fitness function to select a subset of the initial formulations weighted by their closeness to the desired target product properties. The subset of initial formulations is then subjected to GA methodologies analogous to directed evolution in which formulations are subject to mutation and crossover to generate an evolved formulation library. The process of inputting the library into the ML model, selecting for fitness, and apply GA is iterated until a final formulation library predicted to meet the desired target product properties is developed.
[0028] With respect to FIG. 1, an example method 100 for generating an initial formulation library for manipulation by GA is shown. At 102, a user defines one or more target product properties and optional constraints for a desired product formulation. The number and type of target product properties may be based on the system targeted for formulation (e.g., polyurethanes, polyurethane foams) and the particular end-use application (e.g., bedding and cushioning, insulation, sound absorption, energy efficiency applications). Multiple target properties may be selected and a subset of target properties may be prioritized (e.g., density over compression force deflection) to direct formulation selection and library creation. Target product properties may include a single desired value or range of desired values having set endpoints or defined by a percentage around a desired value (e.g., within 5% of a target or in a defined range).
[0029] Target product properties are then used to develop a “parameter space” that defines formulations and formulation libraries. Parameter space criteria include at least a fisting of component names and acceptable concentration ranges (i.e., an upper and lower concentration limit) for a formulation and acceptable operations (e.g., mutation .crossover) that may be performed. In some cases, parameter space may also include general and context-specific descriptors, component classes (e.g., surfactant), and component subclasses (e.g., anionic surfactant).
[0030] Parameter space used to define the formulation and / or library bounds may also include optional user-supplied constraints. As used herein, “constraints” refer to additional controls that can be imposed by a user on prospective formulations generated by an inverse modeling methodology. Constraints may enhance the expected accuracy of hypothetical formulations based on historical data, while also limiting bandwidth consumption attributed to determination of target properties that lie outside of the ranges of the historical data and that are likely to have high error rates and low relevance. Constraints include restrictions on component type and / or number, component class size, concentration limits, physical specifications, (e.g., molecular weight, functionality, polydispersity, ionic charge, and the like), physics-based descriptors, and the like. For example, the number of components in a class (or subclass) can be listed as a constraint (e.g., 1 to 6 for a polyol class) when developing new formulations with user-defined properties, and formulations generated that violate defined constraints may be discarded (or appropriately categorized for later retrieval).
[0031] At 104, parameter space and constraints defined at 102 are then used to search historical formulation database for formulations having historical product properties matching (or approaching within a defined threshold) the target product properties and optional constraints. Thresholds may be as a single desired value or range of desired values having set endpoints or defined by a percentage around a desired value (e.g., within 5% of a target or in a defined range). For example, if a target product property is 40, a threshold may be set as a range (e.g., 30 to 50) or a percentage around the target product property (e.g., 2%, 5%, 10%).
[0032] Historical product properties may be derived from experimental data or predicted from historical data using a trained machine learning module. For example, the historical data may include the relevant target product properties as an associated value or, in some cases, may be a descriptor or property that is predicted using ML module or relevant function.
[0033] At 106, an initial formulation library is generated from formulations retrieved from the historical formulation library in 104 and that satisfy a given parameter space and constraints. In some cases, the initial formulation library may be defined in two ways: (1) a random selection of formulations from historical formulation data within the parameter space; or (2) a weighted selection of formulations from historical formulation data defined by prioritization of one or more target product properties.
[0034] At 108, one or more descriptors are generated for the initial formulation library, and used in conjunction with a ML module to generate the predicted target product properties. Descriptors for the formulations may be generated by converting component physical properties and concentrations (e.g., wt%) to descriptors (e.g., OH-number, NCO, functionality, etc.) using suitable physics-based models / tools known in the art. Descriptors disclosed herein are computed from the properties of the individual components in the formulation, such as by using a physics model suited for the particular chemical application (e.g., polyurethane production). Descriptors may contain data regarding formulation components, component concentrations and ratios, such as ratios of reactants (e.g., the ratio of the isocyanate component and polyol component for polyurethane systems), product generating reactions among components, and properties result from various chemical interactions (e.g., functionality for isocyanates or reactive species, crosslinking, blowing agent reactivity). Suitable descriptors also include those detailing mechanical properties, such as vapor heat capacity, foam density, Young’s modulus, rheology properties, heat transfer properties, and the like.
[0035] ML modules used to predict target product properties given a formulation input (forward model) utilize known physics relationships and chemical formulation data to predict various target product properties, such as density, hardness, compression force deflection(CFDs), compression strength, tear strength, gel time, Lambda, and other properties. ML modules and training methods used to generate the modules are discussed in greater detail below with respect to FIG. 3. In some cases, predicted product properties may also be correlated with the formulations in the historical and / or initial formulation libraries, stored, and / or indexed for later access.
[0036] At 110, the predicted target product properties for the initial formulation library are analyzed using a fitness function that calculates the absolute difference (A) between desired and predicted product property values. As used herein, “fitness function” refers to the absolute percentage difference between target product properties and the predicted product properties for a formulation output, as calculated by a suitable machine learning module. In some cases, the fitness function may be calculated with consideration to multiple target product properties, which may also be ranked in terms of priority for a subset of one or more target product properties.
[0037] Methods disclosed herein may seek to minimize the fitness function, or satisfy the fitness function within a selected threshold. The initial formulation library is then analyzed for fitness and formulations that do not satisfy the fitness function are discarded or indexed and stored for later retrieval.
[0038] At 112, the initial formulation library selected using the fitness function is compared with “stopping criteria” to determine whether the requirements for the prospective formulations have been satisfied. In some cases, if the predicted and desired target product properties are met, the methodology is completed and the prospective formulation library may be moved to the next stage of the process, such as to experimental validation. On the other hand, if stopping criteria are not satisfied, the initial formulation library may proceed to the GA methodology to begin the iterative process of directed evolution to generate an evolved formulation library that may meet the target product properties.
[0039] With respect to FIG. 2, the overall method 100 continues from 112 to the GA methodology 200 (indicated by a dashed box). GA methodologies are a type of directed evolution strategy that uses a set of desired target product properties as the “environment” that prospective formulations are adapted through iterative manipulation and screening for closeness to desired target product properties (i.e., satisfy the fitness function). The GA methodology begins by using the output of the fitness function to select a subset of the initial formulation library for directed evolution by (a) mutation and / or (b) crossover operations to generate an evolved formulation library.
[0040] At 202, a subset of formulations from initial formulation library are selected for crossover or mutation based on user-defined constraints (e.g., fitness function threshold). During selection, a subset of individual formulations are tagged for manipulation as “parents” to create the next generation. Selection may be visualized as a weighted roulette wheel, where the “fittest” formulations according to the particular fitness function have the highest probability of being chosen. Selection methods may include selecting the highest scoring formulations according to a fitness function, or selecting poor scoring formulation to propagate formulation features (e.g., “genes”) that may enhance overall performance through the GA process.
[0041] At 204, initial formulation libraries are subjected to directed evolution operations (e.g., mutation and crossover). Crossover involves crossing one or more components and / or component classes between two or more “parent” formulations, creating “offspring” formulations with potentially novel combinations of components. Mutation involves manipulating components within a formulation by random or directed manipulations of components, component classes, concentrations, and the like.
[0042] Prior to evolution, selected formulations may be converted to bit vectors for manipulation. The components of formulations within the library are provided as a choice and assigned a binary code and given ranges, which the GA methodology uses during crossover and mutation operations. In one example, a binary code is developed for a component class within a formulation having 5 potential choices. The number of bits for the choices are calculated and assigned as the smallest power of two that is greater than or equal to the number of choices. For a component class with 5 choices, at least 3 bits are required to represent each choice (e.g., 23=8 and is >5). One possible encoding is shown in Table 1.
[0043] For a range of values specified in the space file, the number of bits required represent the range are given by Eq. 1.
[0044] Number of bits = \log2(Range) (1) which is rounded to the nearest integer. The original range is then mapped according to Eq. 2.
[0045] [xmin, x^] to the range [0, 2number°fbits- 1] For an individual formulation from the initial formulation library, binary representations of choices for all variables are concatenated into a single binary bit vector. Each part of the bit vector corresponds to a variable in the “space”, and the bits within each part encode the selected choice for that variable. The length of this bit vector is the sum of the lengths of the binary representations of choices for all variables.
[0046] Crossover aims to combine parts of the bit vectors representing each parent, potentially promoting exploitation novel regions of the parameter space and creating diverse offspring. An example of a crossover operation is shown in Tables 2 and 3, where bits from the parent formulations (Table 2) are recombined to form novel offspring formulations (Table 3). A crossover of the bits / genes of the selected parents exchanges two components within the same component class (e.g., Polyol, etc.) and their respective quantities between the two parents.
[0047] Table 2: Example of parent formulations prior to crossover operation.
[0048] Table 3: Example of parent formulations after crossover operation.
[0049] The mutation operator introduces diversity through random or directed changes in a parent formulation (prior to crossover) or offspring formulation (after crossover). In some cases, one or more bits of a formulation are selected and assigned values randomly within the allowable range. Mutations may be a change in the component, component class, or component concentration, and may also include randomized addition of components along with their quantities to an parent formulation. For avoidance of doubt, GA methodologies may include one or more (or all) evolution operations performing only crossover or only mutation.
[0050] For mutation, random substitutions are made for at least one component within the component classes of the starter formulations (e.g., the parent formulation bit vector or the chile formulation bit vector) using class equivalents, removing duplicate formulations, and generating new hypothetical formulations. For example, if a parent formulation uses components selected from classes a+b+c, hypothetical offspring formulations may include a+b+d, a+e+c, f+b+c, etc. (where alternate letters represent functional class equivalents). In some cases, combinatorics may also substitute components into one or more classes, such as a+d+e, f+d+c, f+b+c, etc.
[0051] During mutation operations, offspring formulations are assigned concentration ratios in proportion to that associated with the historical data ranges to produce an evolved formulation library. In some cases, methods may also include varying the weight amount / ratio extend beyond the historical values by some amount, such as more or less than 5%, 10%, 15%, and the like.
[0052] Mutation may also include linear combinations of multiple classes. In another example, randomization may involve linear combinations of component classes, such as the case in which the total amount of components in the classes for polyols and isocyanate are equal to 4, combinations may include 1 polyol and 3 isocyanates, 3 polyol and 1 isocyanates, 2 polyol and 2 isocyanates, etc.
[0053] Table 4 shows an additional example in which a parent formulation is modified by a mutation operation by changing values of multiple components, which includes the elimination of a component (Additive B).
[0054] Table 4: Example of mutation operation.
[0055] The GA methodology may be iterated multiple times, in which the process of selection, crossover, and mutation is repeated on the evolved formulation library until stopping criteria are met.
[0056] At 206, an evolved formulation library is generated by collecting the offspring formulations from the crossover and / or mutation operations, applying the fitness function and constraints, and removing rejected and redundant formulations. In some cases, formulations not meeting the defined constraints are discarded and the evolved library is replenished with new formulations generated through crossover and mutation of the initial subset formulation library. In this way, the surviving formulations serve as the template / parent formulations for successive generations.
[0057] At 208, as with 108, relevant descriptors are generated for the evolved formulation library and a ML module is used to generate predicted target product property values. In some cases, the evolved formulations may be ranked based on the absolute percentage error (APE) relative to the mean of the target property range(s) or a specific value(s). At 210, the absolute difference between desired and predicted target product property values are calculated for the evolved formulation library.
[0058] At 212, the stopping criteria are assessed and the process is terminated or reiterated from 202 until evolved formulation library is within acceptable threshold of target product property values.
[0059] The process of evaluation, selection, crossover and mutation, and evaluation (202 to 212) forms one generation in the execution of GA methodology 200. The generation of new evolved formulations ends upon reaching a threshold of closeness (A) to the desired target product properties, or other metric is met, such as an arbitrarily selected number of user-defined generations (e.g., 50 to 100), a lack of improvement over 10 generations, or other criteria.
[0060] At 108 and 208, formulation libraries may be input into a machine learning module to determine one or more predicted target product properties, including prediction intervals. Machine learning modules may include any suitable machine learning model trained to determine one or more target product properties. Suitable machine learning modules may include artificial neural networks such as deep neural networks (DNNs), symbolic regression, recurrent neural networks (RNNs) that include long short-term memory (LSTM) networks or Gated Recurrent Unit (GRU) networks, decision trees, random forests, boosted trees such as gradient boosted trees (XGBoost), linear regression, partial least squares regression, support vector machines, multilayer perceptron (MLP), autoencoders (e.g., denoising autoencoders such as stacked denoising autoencoders), Bayesian networks, support vector machines (SVMs), hidden Markov models (HMNIs), and the like. Commercially available software packages may include JMP software, Microsoft AzureML, SAP data analysis tools, soft independent modeling by class analogy (SIMCA) by Sartorius, and the like.
[0061] Machine learning architectures may also utilize deep learning in which a neural network is generated with multiple layers. These layers extract successively higher order features from a training data set. For chemical formulations, examples of layers containing lower order features may include general classifications of component type, while layers including higher order features in the network may include details dependent on functional groups, ionization state, charge, and the like.
[0062] With particular respect to FIG. 3, machine learning modules disclosed herein may be trained using a training set composed of historical data and / or hypothetical data for a given chemical application (e.g., polyurethane foams). At 302, one or more training data sets are constructed from one or more variable parameters including formulation components, descriptors, process conditions, and composition properties.
[0063] Training data sets may also include one or more descriptors generated by converting component physical properties and concentrations (e.g., wt%) to descriptors (e.g., OH-number, NCO, functionality, etc.) using suitable physics-based models / tools known in the art. Descriptors disclosed herein are computed from the properties of the individual components in the formulation, such as by using a physics model suited for the particular chemical application (e.g., polyurethane compositions). Descriptors may contain data regarding formulation components, component concentrations and ratios, such as ratios of polyurethane reactants (e.g., the ratio of the isocyanate component and polyol component), product generating reactions among components, and properties result from various chemical interactions (e.g., functionality for isocyanates or reactive species, crosslinking, blowing agent reactivity). Suitable descriptors also include those detailing mechanical properties, such as vapor heat capacity, foam density, Young’s modulus, rheology properties, heat transfer properties, and the like.
[0064] At 304, feature selection is performed on the training data set constructed in 302. During feature selection, a subset of the variable parameters identified in the training set are identified as “driving” variables affecting the targeted formulation property. Feature selection may then involve the exclusion of irrelevant, noisy, and redundant features from the training data set.
[0065] Feature selection techniques may include one or more of descriptor feature selection; removing constraining features; correlation testing methods such as Pearson, Spearman, Kendall, and the like; analysis of variance (ANOVA) univariate testing; mean absolute different testing; LI or Least absolute shrinkage and selection operator (Lasso) regularization; multivariate analysis (baseline); and the like.
[0066] At 306, machine learning model architectures are surveyed by training one or more machine learning models with the driving variables established from 304 as input. The generated machine learning models from the surveyed architectures are then compared and rated for accuracy, which is then used to select one or more model architectures used in subsequent stages. For example, a machine learning module may output one or more predicated performance properties from a prospective chemical formulation. In some cases, more than one trained machine learning model may be combined into a machine learning module, where the output is the result of constituent machine learning model having higher accuracy for the selected target product property and / or is the result of averaging the output of one or more machine learning models.
[0067] At 308, the method further includes training and validating multiple models with a testing data set containing the variable data and target product property data, and then selecting an appropriate model based on desired model criteria such as best fit according to error calculation techniques such as R2, mean associated percent error (MAPE), root mean squared error (RMSE), and the like. The testing data set may contain chemical formulation information and descriptor information that is similar in structure to the training data set, however, it usually contains sample information that is minimally duplicative to the training data in order to provide an adequate test of the ability of the machine learning module to handle new information, as opposed to storage and retrieval of the training data set values.
[0068] Underfitting is a scenario where a model cannot capture the relationship between the input and output variables. In the case of R2-based methods, the underfitted models tend to have undesirably low training R2. The training R2threshold may be set at a desired accuracy to filter out underfit models, such as greater than 0.70, 0.85, 0.88, or 0.9. Further, overfit models may occur in which the model is too closely aligned to the training data, and the learned representation cannot predict the validation data accurately. In some cases, error calculation techniques may also be used to select for generalizable models while filtering out overfitted models. For example, validation percentage RMSE threshold of less than 40%, 35%, or 30% may be used to predict target product properties.
[0069] In some cases, parameters for the model are optimized by minimizing the mean squared error (MSE) between the actual value yt and the predicted value y; , as shown in equation (I) below: where n is the total number of data points, y^ is the actual value of the target product property measured, i is the sample data point, and yt is the model predicted value of the target product property.
[0070] MSE is a metric used to measure the average of the squares of the difference between the predicted values and the actual values. RMSE means root mean squared error that is a metric to measure the difference between the predicted values and the actual values. Percentage RMSE is the square root of MSE normalized by the population mean to a dimensionless number expressed as a percentage. The percentage RMSE can be calculated as below: 100 (II) where n, yt, i and ytare as defined above in equation (I) above, and yF is the mean value for the target product property measured. The lower the percentage RMSE, the better the model fits a dataset.
[0071] R2means the coefficient of discrimination, which is a commonly used performance metric for regression that calculates the proportion of variance explained by a regression model. R2normally ranges from 0 to 1 and can be calculated according to equation (III) below: where y i and ytare as defined above in equation (I) above, and yi is the mean value. The higher the R2, the better the model fits a dataset.
[0072] In one embodiment, the method can provide predictions for target product properties of a chemical composition or product, as indicated by “training R2> 0.70” as calculated according to the equation (III) above, and test percentage RMSE less than 30% (< 30%) as calculated according to the equation (II) above. In addition, validation percentage RMSE is <30% as calculated according to the equation (II) above. “Training R2” refers to R2for the training dataset, “validation percentage RMSE” refers to the percentage RMSE for the validation dataset, and “test percentage RMSE” refers to the percentage RMSE for the test dataset.
[0073] At 310, prediction intervals are generated for the created models using bootstrapping (bagging) or other methodologiesto provide a measure of the probable error for target product property predictions. For example, prediction intervals may be represented as range of variance around the associated mean value. During bootstrapping, samples are drawn from the training data and input into a model, and the results are combined by averaging for regression and simple voting for classification, to obtain the overall prediction. Other suitable prediction interval techniques include those described in open-source libraries such including the model agnostic prediction interval estimator (MAPIE), including conformal prediction, naive methods, split methods, jackknife methods, CV methods, ensemble batch prediction intervals (EnbPI) methods, and the like.
[0074] At 312, model optimization is used to identify driving variables to determine whether the obtained model accuracy is increased by removing or adding additional variables. Optimization methods may evaluate how significant each selected variable is in determining the prediction of the model outputs. The subset of variable having the greatest impact on model accuracy can then be used to determine if training process 300 is complete (e.g., prediction accuracy acceptable), or should be reiterated by repeating 304 to 312. In some cases, iterations may be repeated multiple times such as 2 to 3 or more.
[0075] Optimization methods may evaluate the selected variables on the target product properties output from the models and processed using interpreting and explanatory software, such as SHapley Additive ex-Planation (SHAP), Local Interpretable Model-Agnostic Explanations (LIME), and the like. For example, SHAP analysis assigns each input variable an importance value for a particular prediction by comparing a model’s output with and without a specific variable. SHAP values are then computed using a sum that represents the impact of each input variable added to the model averaged over all possible orderings of variable being introduced. Positive SHAP values indicate that higher feature values lead, on average, to higher predicted values, while negative SHAP values indicate that lower feature values lead, on average, to lower predicted values.
[0076] At 314, the optimized and trained machine learning module is deployed and used to generate predicted target product properties for a prospective formulation.
[0077] FIG. 4 illustrates a method of using a trained machine learning module to generate one or more target product properties from a prospective formulation, such as employed in 108 and 208. At 402, select prospective formulation(s) are input into a trained machine learning module to calculate one or more predicted target product properties at 404.
[0078] At 406, the predicted properties may be used to determine whether the evolved formulation libraries meet desired performance criteria, or whether further modifications to the prospective formulations. Performance criteria may vary depending on the polymer system targeted for formulation (e.g., polyurethanes, polyurethane foams); the particular application (e.g., bedding and cushioning, insulation, sound absorption, energy efficiency applications, and the like); and the presence of other factors such as the need to minimize costs, carbon footprint, use of hazardous substances, and the like. Evolved formulation libraries obtained from 406 may be advanced as candidate formulations for experimental validation or to evaluate an existing process for producing products or articles.
[0079] Computer systems disclosed herein may coordinate the input, selection, and directed evolution of formulation libraries, including the integration of a ML module for generating predicted target product properties at various steps. Computer systems may include a processor and data storage, where the data storage has stored thereon computer-executable instructions that, when executed by the processor, cause the computing device to carry out functions. Computing devices may include a client device (e.g., a device actively operated by a user), a server device (e.g., a device that provides computational services to client devices), or some other type of computational platform. Some server devices can operate as client devices from time to time in order to perform particular operations, and some client devices can incorporate server features.
[0080] The processor useful in the present disclosure can be one or more of any type of computer processing element, such as a central processing unit (CPU), a co-processor (e.g., a mathematics, graphics, neural network, or encryption co-processor), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a network processor, and / or a form of integrated circuit or controller that performs processor operations.
[0081] The data storage can include one or more data storage arrays that include one or more drive array controllers configured to manage read and write access to groups of hard disk drives and / or solid-state drives.
[0082] Computing devices may be deployed to support a clustered architecture. The exact physical location, connectivity, and configuration of these computing devices can be unknown and / or unimportant to client devices. Accordingly, the computing devices can be referred to as “cloudbased” devices that can be housed at various remote data center locations, such as a cloud-based server cluster. Desirably, the computing device is a cloud-based server cluster and inputting the concentration data to the model via a web-based user interface where users can get access.
[0083] Fig. 5 depicts a schematic drawing of a cloud-based server cluster 500 in accordance with one example of the present disclosure. Desirably, operations of a computing device can be distributed between server devices 502, data storage 504, and routers 506, all of which can be connected by local cluster network 308. The amount of server devices 502, data storage 504, and routers 506 in the server cluster 500 can depend on the computing task(s) and / or applications assigned to the server cluster 500. For example, the server devices 502 can be configured to perform various computing tasks of the computing device. Thus, computing tasks can be distributed among one or more of the server devices 502. As an example, the data storage 304 can store any form of database, such as a structured query language (SQL) database . Furthermore, any databases in the data storage 504 can be monolithic or distributed across multiple physical devices. The routers 506 can include networking equipment configured to provide internal and external communications for the server cluster 500. For example, the routers 506 can include one or more packet-switching and / or routing devices (including switches and / or gateways) configured to provide (i) network communications between the server devices 502 and the data storage 504 via the cluster network 508, and / or (ii) network communications between the server cluster 500 and other devices via the communication link 510 to the network 512. The server devices 502 can be configured to transmit data to and receive data from cluster data storage 504. Moreover, the server devices 502 can have the capability of executing various types of computerized scripting languages, such as Perl, Python, PHP Hypertext Preprocessor (PHP), Active Server Pages (ASP), or JavaScript. Computer program code written in these languages can facilitate the providing of web pages to client devices, as well as client device interaction with the web pages.
[0084] Examples
[0085] The following examples are provided to illustrate the embodiments of the invention, but are not intended to limit the scope thereof. Table 1 provides the materials used in the following examples.
[0086] Example 1: Use of GA methodology to generate prospective polyurethane foam formulations
[0087] In this example, a GA methodology was used to produce an evolved formulation library containing formulations producing a PU foam having the foam properties shown in Table 1. Genetic algorithm implementations were prepared in ChemML5. Bit vectors encode components and their quantities in parts per hundred polyol, which function as “genes” during directed evolution.
[0088] Example 2: Design of insulating polyurethane foam formulations Using the same GA methodology, we also targeted formulations with different foam properties (table below). This indicates that this inverse design methodology can be generalized for different market segments with polyurethane business.
[0089] Based on the input target properties, 5 formulations meeting the criteria were generated for later potential validation. The components are shown in pphp.
[0090] While the foregoing is directed to exemplary embodiments, other and further embodiments may be devised without departing from the basic scope thereof, and the scope thereof is determined by the claims that follow.
Claims
Claims1. A method of generating a prospective formulation library of chemical formulations, comprising:(a) defining one or more target product properties for the prospective formulation library ;(b) preparing an initial formulation library curated from historical data comprising one or more historical product properties;(c) applying a fitness function to determine difference between the one or more historical product properties and the one or more target product properties, and discarding formulations from the initial formulation library that do not meet the fitness function;(d) using a genetic algorithm based methodology to develop an evolved formulation library from the initial formulation library comprising: i. selecting one or more parent formulations from the initial formulation library having a desired closeness to the target product properties according to the fitness function; ii. converting each of the one or more parent formulations to a bit vector; iii. generating one or more offspring formulations from the one or more parent formulation bit vectors using mutation and / or crossover operations; and iv. collecting the one or more offspring formulations to create the evolved formulation library;(e) generating descriptors for the one or more offspring formulations in the evolved formulation library;(f) analyzing the evolved formulation library using a trained machine learning module to calculate one or more predicted product properties;(g) applying the fitness function to the evolved formulation library to determine difference between the one or more target product properties and the one or more predicted product properties; and(h) based on the determined difference in (g): i. collecting a subset of the evolved formulation library to generate the prospective formulation library having a desired closeness to the target product properties; orii. repeat steps (d) to (g) until the one or more target product properties desired for the prospective formulation library or stopping criteria are met.
2. The method of claim 1, further comprising selecting one or more formulations from the prospective formulation library for experimental validation.
3. The method of claim 2, further comprising modifying a chemical process using the validated formulation library.
4. The method of claim 1, wherein the chemical formulations are polyurethane formulations.
5. The method of claim 1, wherein the descriptors are specific to polyurethanes or polyurethane foams.
6. The method of claim 1, wherein the one or more historical product properties are derived from experimental data or predicted from historical data using a trained machine learning module.
7. The method of claim 1 , offspring formulations from the one or more parent formulation bit vectors using mutation and crossover operations.
8. The method of claim 1 , wherein the stopping criteria is selected from a group consi sting of: reaching a threshold of closeness (A) to the desired target product properties, an arbitrarily selected number of generations, and a lack of improvement over 10 generations.
9. The method of claim 1, wherein crossover comprises exchanging one or more components and / or component classes between the parent or offspring formulation bit vectors.
10. The method of claim 1, wherein mutation comprises a change in at least one component, component class, component concentration, or randomized addition.
Citation Information
Patent Citations
Simulation guided inverse design for material formulations
US11837333B1