Evolutionary algorithm for inverse design of polyurethane formulations
Patent Information
- Application Number
- CN202580018233.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-22
- Filing Date
- 2025-03-04
- Publication Date
- 2026-09-29
AI Technical Summary
这可能导致生成大型库,这些大型库需要通过各种实验进行验证以确认所有设计标准均被满足,从而导致时间和材料的显著成本
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The implementation scheme relates to a computer-based method for generating a desired chemical formulation given inputs of one or more desired target product properties. The method may also include validating the desired formulation and modifying the chemical process accordingly. Background Technology
[0002] The design of cross-industry chemical formulations typically involves combining various chemical components through labor-intensive methodologies to generate products with one or more target product properties. These labor-intensive methodologies can involve high-throughput experiments conducted sequentially and / or in parallel. However, as the scale and complexity of chemical formulations increase, the number of combinations and experiments required to map the chemical parameter space increases exponentially. This has led to the development of computer-based methods that utilize machine learning models trained on large amounts of historical formulation data, generating predicted target properties from input desired formulations. Using machine learning (ML) modeling methods enables users to design and customize desired formulations.
[0003] To achieve full practicality with ML modeling methods, users can design and customize desired formulations, and then use "forward" ML models (which predict material properties based on the formulation) to investigate whether these desired formulations meet the target properties. On the other hand, "reverse" ML modeling can be used to generate new desired formulations that meet constraints for various specific applications, and can include novel and unique combinations of components predicted to meet target product properties. However, due to the permutations and combinations of components, the number of possible formulation combinations can be virtually infinite and is often unconstrained by domain-specific and empirical knowledge from the field. This can lead to the generation of large libraries that require validation through various experiments to confirm that all design criteria are met, resulting in significant time and material costs. Summary of the Invention
[0004] In one aspect, the method may include generating a prospective formulary library of chemical formulations, comprising: (a) defining one or more target product characteristics; (b) preparing an initial formulary library from which historical data of one or more historical product characteristics associated with the prospective formulary library are compiled; (c) applying a fitness function to determine the difference between the one or more target product characteristics and the one or more predicted product characteristics, and discarding formulas from the initial formulary library that do not satisfy the fitness function; (d) developing an evolutionary formulary library from the initial formulary library using a genetic algorithm-based methodology, comprising: (i) selecting one or more parent formulas from the initial formulary library that have a desired closeness to the target product characteristics according to the fitness function; (ii) converting each of the one or more parent formulas into a bit vector; and (iii) using mutation and / or crossover operations to... The one or more parent recipe bit vectors generate one or more offspring recipes; and (iv) collect the one or more offspring recipes to create the evolutionary recipe library; (e) generate descriptors for the one or more offspring recipes in the evolutionary recipe library; (f) use a trained machine learning module to analyze the evolutionary recipe library to compute one or more predicted product characteristics; (g) apply the fitness function to the evolutionary recipe library to determine the difference between the one or more target product characteristics and the one or more predicted product characteristics; and (h) based on the difference determined in (g); (i) collect a subset of the evolutionary recipe library to generate the expected recipe library with a desired closeness to the target product characteristics; or (ii) repeat steps (d) to (g) until one or more target product characteristics or stopping criteria are met as desired by the expected recipe library. Attached Figure Description
[0005] Figure 1 and Figure 2 This is a flowchart illustrating a general approach to developing one or more desired chemical formulations using a genetic algorithm methodology.
[0006] Figure 3 This is a flowchart illustrating a method for training a machine learning module to estimate the properties of one or more target products of a desired chemical formulation.
[0007] Figure 4 This is a flowchart illustrating a method for integrating a trained machine learning module capable of generating one or more desired product characteristics into a genetic algorithm methodology.
[0008] Figure 5 A schematic diagram illustrating a cloud-based server cluster according to an example of this disclosure. Detailed Implementation
[0009] The method disclosed in this paper enables the generation of a library of hypothetical formulations using a genetic algorithm (GA) methodology that simulates natural evolution. Furthermore, the GA methodology transforms the formulation system into bit vectors, which are then systematically manipulated to generate novel formulations within desired constraints. The GA methodology can be combined with ML-based verification techniques that can provide optimized formulations, thereby reducing the time and cost of experimental verification. In some cases, the method may include a reverse design approach to generate a library of hypothetical formulations containing one or more expected chemical formulations derived from inputs of one or more desired target product properties (effectively in real time). The method disclosed in this paper can also be combined with user interface strategies that enable the management of large numbers of formulations and facilitate searches based on a range of desired properties.
[0010] The method disclosed in this paper utilizes historical formulation data to train ML models for novel formulation design, including forward design models (predicting material properties based on composition) and backward design models (predicting material composition based on desired properties), to generate a predictive formulation library of increasing size. However, while backward modeling has the ability to generate a large predictive formulation library, the constituent chemical components and properties tend to reflect those in the historical dataset used to train the ML module. The low variance in the predicted formulations then limits the emergence of novel component combinations, particularly new formulations with target product properties, components, and / or concentration ranges that are sparsely represented in the training dataset.
[0011] The method disclosed in this paper uses GA methodologies to generate a hypothetical formula library, which incorporates directed evolution techniques of crossover and / or mutation to modify the initial formula library. This directed evolution technique generates a higher probability of accessing undersampled regions of the parameter space. The method disclosed in this paper can increase: (i) the relevance of the expected formula library generated from a smaller amount of historical formula data; and (ii) the number of predicted formulas with unique combinations of components that satisfy the characteristics of the target product reaches a defined threshold.
[0012] As used herein, “formulation” refers to a combination of components (e.g., polymer forming compositions, reactant mixtures, blends, etc.) for a particular application (e.g., polymer formation, chemical processing, article construction).
[0013] As used herein, a "formula library" or "recipe library" is a collection of two or more formulations. A library may include data tagged to each formulation, including the name of the formulation components, component category (e.g., monomer, surfactant, catalyst), component type (e.g., polyol, anionic surfactant, gelling catalyst), concentration limits, etc.
[0014] As used in this article, “component” refers to chemical types, including but not limited to monomers, prepolymers, catalysts, additives, etc.
[0015] As used herein, a “component category” refers to a class of component types classified by one or more similar characteristics (e.g., chemical structure, function, molecular weight, polarity, etc.). Component category construction can be application-specific, and may be derived from historical data and / or general knowledge. In some cases, categories can be constructed based on existing types of raw materials or subdivided by chemical type. For example, the chemical formulation components of polyurethane foam can be categorized into polyols, isocyanates, catalysts, surfactants, blowing agents, additives, etc.
[0016] As used in this article, “process conditions” (e.g., atmospheric pressure, variable pressure foaming, relative humidity, overfill percentage, etc.) refers to the expression describing the process conditions that affect the characteristics.
[0017] As used herein, a “descriptor” (or “specific descriptor”) refers to an expression that describes the relevance within a chemical system (e.g., a polymer system) and can provide additional information and / or generalization about the system’s behavior. Examples of descriptors for polymer systems (e.g., polyurethane-specific descriptors, polyurethane foam-specific descriptors) include monomer molecular weight, water content, catalytic activity, polymer chain entanglement, etc. The descriptors described herein can be calculated from component information and concentrations using various physics-based methods and models of target product properties.
[0018] As used in this paper, a “forward ML model” (or “ML model”) is a machine learning (ML) model trained on a historical recipe library to generate predictive properties given a recipe input. A suitable ML model can have varied architectures and algorithms to predict different target product properties and can include modules that combine multiple models or the average output from multiple models. ML models can be incorporated into methods, for example, for calculating the mean absolute difference in a fitness function.
[0019] As used in this article, “historical formulation data” includes component names, component quantities, target product characteristics, and optional descriptors obtained from previous studies and experimental results.
[0020] As used herein, “target product characteristics” refers to characteristics (e.g., chemorheology, foam density, hardness, modulus, etc.) associated with a unique chemical formulation based on desired user input selection for a given product application.
[0021] As used in this article, “variable parameters” refers to features of the machine learning module and / or model that can change during training (e.g., chemical species and / or category, concentration, descriptor, etc.).
[0022] The method disclosed in this paper trains an ML module using a genetic algorithm (GA) to generate novel predictive chemical formulations with new combinations of components beyond those present in the initial training dataset. The GA methodology utilizes a process similar to evolution in biological systems, selecting candidates (e.g., formulations, chemical components, and component categories) based on defined criteria. These candidates are then used to generate novel “evolutionary” formulations through crossover and mutation operations. Methods for formulating polymer systems such as polyurethane and polyurethane foam are described below; however, it is envisioned that these techniques can be applied to other polymer systems and materials.
[0023] The method disclosed in this paper utilizes backpropagation modeling, in which an initial formula library is compiled from historical data and input into an ML model to predict one or more target product properties. The predicted target product properties are then input into a fitness function to select a subset of initial formulas, weighted by their proximity to the desired target product properties. This subset of initial formulas is then subjected to a GA methodology similar to directed evolution, where formulas undergo mutation and crossover to generate an evolving formula library. The process of iteratively inputting the library into the ML model, selecting fitness, and applying GA continues until a final formula library predicted to satisfy the desired target product properties is developed.
[0024] about Figure 1 Example method 100 for generating an initial formula library manipulated by GA is shown. At 102, the user defines one or more target product properties and optional constraints for the desired product formulation. The number and type of target product properties can be based on the system the formulation targets (e.g., polyurethane, polyurethane foam) and the specific end-use application (e.g., bedding and padding, insulation, sound absorption, energy efficiency applications). Multiple target properties can be selected, and subsets of target properties can be prioritized (e.g., density takes precedence over compressive force deflection) to guide formulation selection and library creation. Target product properties can include a single desired value or range of desired values with a set endpoint or defined by a percentage around the desired value (e.g., within 5% of the target or within a defined range).
[0025] Then, the target product characteristics are used to develop a “parameter space” that defines formulations and formulation libraries. Parameter space criteria include at least the component names and acceptable concentration ranges for the formulation (i.e., upper and lower concentration limits), as well as a list of acceptable operations that can be performed (e.g., mutation, crossover). In some cases, the parameter space may also include general and context-specific descriptors, component categories (e.g., surfactants), and molecular classes (e.g., anionic surfactants).
[0026] The parameter space used to define recipe and / or library boundaries may also include optional user-provided constraints. As used herein, a “constraint” refers to additional control that can be imposed by the user on the expected recipe generated by the reverse modeling methodology. Constraints can enhance the predicted accuracy of the hypothetical recipe based on historical data, while also limiting bandwidth consumption attributable to identifying target characteristics that are outside the range of historical data and are likely to have a high error rate and low relevance.
[0027] Constraints include restrictions on component type and / or quantity, component class size, concentration limits, physical specifications (e.g., molecular weight, functionality, polydispersity, ionic charge, etc.), physics-based descriptors, etc. For example, when developing new formulations with user-defined properties, the number of components in a class (or subclass) can be listed as a constraint (e.g., 1 to 6 for polyols), and formulations that violate the defined constraints can be discarded (or appropriately categorized for later retrieval).
[0028] At point 104, the parameter space and constraints defined at point 102 are then used to search the historical recipe database for recipes with historical product characteristics that match (or are close to) the target product characteristic and optional constraints. The threshold can be a single desired value or a range of desired values with set endpoints, or it can be defined as a percentage around the desired value (e.g., within 5% of the target or within a defined range). For example, if the target product characteristic is 40, the threshold can be set as a range around the target product characteristic (e.g., 30 to 50) or a percentage (e.g., 2%, 5%, 10%).
[0029] Historical product characteristics can be derived from experimental data or predicted from historical data using trained machine learning modules. For example, historical data can include relevant target product characteristics as associated values, or in some cases, descriptors or characteristics predicted using ML modules or correlation functions.
[0030] At point 106, an initial formula library is generated from formulas retrieved from the historical formula library at point 104 that satisfy the given parameter space and constraints. In some cases, the initial formula library can be defined in two ways: (1) formulas are randomly selected from historical formula data within the parameter space; or (2) formulas are selected in a weighted manner from historical formula data defined by priority sorting based on one or more target product characteristics.
[0031] At point 108, one or more descriptors are generated for the initial formulation library and used in conjunction with the ML module to generate predicted target product properties. The formulation descriptors can be generated by converting the physical properties and concentrations (e.g., weight %) of components into descriptors (e.g., OH number, NCO, functionality, etc.) using suitable physics-based models / tools known in the art. The descriptors disclosed herein are calculated from the properties of the individual components in the formulation, such as by using a physics model suitable for a specific chemical application (e.g., polyurethane production). Descriptors may contain data on formulation components, component concentrations and ratios (such as the ratio of reactants (e.g., the ratio of isocyanate to polyol components in a polyurethane system)), product-forming reactions between components, and properties resulting from various chemical interactions (e.g., functionality of isocyanates or reactive species, crosslinking, foaming agent reactivity). Suitable descriptors also include those that detail mechanical properties (such as vapor heat capacity, foam density, Young's modulus, rheological properties, heat transfer properties, etc.).
[0032] The ML module for predicting target product properties given a formulation input (forward model) utilizes known physical relationships and chemical formulation data to predict various target product properties, such as density, hardness, compressive force deflection (CFD), compressive strength, tear strength, gel time, λ, and other properties. See below for reference. Figure 3 The ML module and the training methods used to generate the module are discussed in more detail. In some cases, predicted product properties can also be associated with, stored, and / or indexed with recipes from historical and / or initial recipe libraries for later access.
[0033] At point 110, a fitness function is used to analyze the predicted target product characteristics of the initial formulation library, which calculates the absolute difference (Δ) between the expected and predicted product characteristic values. As used herein, a “fitness function” refers to the absolute percentage difference between the target product characteristic and the predicted product characteristic of the formulation output, as calculated by a suitable machine learning module. In some cases, multiple target product characteristics may be considered when calculating the fitness function, and these target product characteristics may also be prioritized according to the priority of a subset of one or more target product characteristics.
[0034] The method disclosed in this paper seeks to minimize the fitness function, or to satisfy the fitness function within a selected threshold. The fitness of the initial recipe library is then analyzed, and recipes that do not satisfy the fitness function are discarded or indexed and stored for later retrieval.
[0035] At point 112, the initial formulation library selected using the fitness function is compared with the "stopping criteria" to determine whether the requirements of the expected formulation have been met. In some cases, if the predicted and desired target product characteristics are met, the methodology is complete and the expected formulation library can be moved to the next stage of the process, such as experimental validation. On the other hand, if the stopping criteria are not met, the initial formulation library can proceed to the GA methodology to begin an iterative process of directed evolution to generate an evolutionary formulation library that can meet the target product characteristics.
[0036] about Figure 2 Overall Methodology 100 continues from 112 to GA Methodology 200 (indicated by the dashed box). GA Methodology is a directed evolutionary strategy that uses a set of desired target product characteristics as an "environment" to adapt the desired recipe to approximate the desired target product characteristics (i.e., satisfying the fitness function) through iterative manipulation and selection. GA Methodology begins by selecting a subset of the initial recipe library using the output of the fitness function to generate an evolutionary recipe library through directed evolution via (a) mutation and / or (b) crossover operations.
[0037] At point 202, a subset of recipes is selected from the initial recipe library for crossover or mutation based on user-defined constraints (e.g., a fitness function threshold). During selection, subsets of individual recipes are labeled as “parents” for manipulation to create the next generation. Selection can be visualized as a weighted roulette wheel, where the “most suitable” recipe according to a specific fitness function has the highest probability of being selected. Selection methods may include selecting the highest-scoring recipe based on the fitness function, or selecting poorly-scoring recipes to propagate recipe features (e.g., “genes”) that can enhance overall performance through the GA process.
[0038] At position 204, the initial formulation library undergoes directed evolutionary operations (e.g., mutation and crossover). Crossover involves crossing one or more components and / or component classes between two or more "parent" formulations to create "child" formulations with potentially novel combinations of components. Mutation involves manipulating the components within a formulation by randomly or directed manipulation of components, component classes, concentrations, etc.
[0039] Before evolution, the selected recipes can be converted into bit vectors for manipulation. Recipe components within the library are provided as choices and assigned binary codes and given ranges, which the GA methodology uses during crossover and mutation operations. In one embodiment, binary codes are developed for component categories within a recipe with 5 potential choices. The number of bits for each choice is calculated and assigned as a power of 2 greater than or equal to the number of choices. For a component category with 5 choices, at least 3 bits are needed to represent each choice (e.g., 2^32 bits). 3 =8 and >5). One possible encoding is shown in Table 1.
[0040]
[0041] For a range of values specified in a spatial file, the number of bits required to represent that range is given by Equation 1.
[0042] Number of bits = [log2 (range)](1)
[0043] Round it to the nearest integer. Then map the original range according to Equation 2.
[0044] [x min , x max [0, 2] 位数 – 1]
[0045] For a single recipe from the initial recipe library, the binary representations of all variable choices are concatenated into a single binary bit vector. Each part of the bit vector corresponds to a variable in the "space," and the bits within each part encode the choice made for that variable. The length of this bit vector is the sum of the lengths of the binary representations of all variable choices.
[0046] Crossover aims to combine portions of the bit vectors representing each parent, potentially facilitating the utilization of novel regions of the parameter space and the creation of distinct offspring. Examples of crossover operations are shown in Tables 2 and 3, where bits from the parent formulation (Table 2) are recombinated to form novel offspring formulations (Table 3). The crossover of selected parent bits / genes involves the exchange of two components within the same component class (e.g., polyols, etc.) and their corresponding amounts between the two parents.
[0047]
[0048]
[0049] Mutation operators introduce diversity by randomly or purposefully altering the parent recipe (before crossover) or the offspring recipe (after crossover). In some cases, one or more bits of the recipe are randomly selected and assigned values within permissible limits. Mutations can be changes in components, component classes, or component concentrations, and may also involve randomly adding a component along with its amount to the parent recipe. To avoid ambiguity, GA methodologies may include one or more (or all) evolutionary operations, such as performing only crossover or only mutation.
[0050] For mutations, at least one component within the component category of the initial formulation (e.g., the parent formulation bit vector or the child formulation bit vector) is randomly replaced using a class equivalence, removing duplicate formulations and generating a new hypothetical formulation. For example, if the parent formulation uses components selected from category a+b+c, then the hypothetical child formulation may include a+b+d, a+e+c, f+b+c, etc. (where the substitution letter represents a functional class equivalence). In some cases, combinatorial theory may also substitute components into one or more categories (such as a+d+e, f+d+c, f+b+c, etc.).
[0051] During the mutation process, concentration ratios proportional to historical data ranges are assigned to progeny formulations to generate an evolutionary formulation library. In some cases, the method may also include altering the amount / ratio of weights beyond historical values by a certain amount, such as greater than or less than 5%, 10%, 15%, etc.
[0052] Mutations can also include linear combinations of multiple categories. For example, randomization can involve linear combinations of component categories, such as when the total amount of components in the categories of polyols and isocyanates is equal to 4. Combinations can include 1 polyol and 3 isocyanates, 3 polyols and 1 isocyanate, 2 polyols and 2 isocyanates, etc.
[0053] Table 4 shows additional embodiments in which the parent formulation is modified by a mutation operation that changes the values of multiple components, including the elimination of component (additive B).
[0054]
[0055] The GA methodology can be iterated multiple times, in which the processes of selection, crossover, and mutation are repeated on the evolutionary recipe library until the stopping criteria are met.
[0056] At point 206, an evolutionary recipe library is generated by collecting progeny recipes from crossover and / or mutation operations, applying fitness functions and constraints, and removing rejected and redundant recipes. In some cases, recipes that do not satisfy the defined constraints are discarded, and the evolutionary library is supplemented with new recipes generated through crossover and mutation of an initial subset of the recipe library. In this way, surviving recipes serve as templates / parent recipes for successive generations.
[0057] At position 208, similar to position 108, relevant descriptors are generated for the evolutionary formulation library, and the ML module is used to generate predicted target product property values. In some cases, evolutionary formulations can be ranked based on the average or absolute percentage error (APE) relative to the target property range or a specific value.
[0058] At position 210, the absolute difference between the expected target product characteristic value and the predicted target product characteristic value is calculated for the evolutionary formula library.
[0059] At point 212, the stopping criteria are evaluated, and the process is terminated or repeated from point 202 until the evolutionary formulation library is within an acceptable threshold of the target product characteristic values.
[0060] The processes of evaluation, selection, crossover, and mutation, along with assessment (202 to 212), are carried out in the execution of the GA methodology 200. The generation of a new evolutionary recipe ends when a proximity threshold (Δ) to the desired target product characteristics is reached, or when other metrics are met, such as any user-defined generation chosen (e.g., 50 to 100), no improvement for 10 consecutive generations, or other criteria.
[0061] At positions 108 and 208, the recipe library can be input into the machine learning module to determine one or more target product properties, including prediction intervals. The machine learning module can include any suitable machine learning model trained to determine one or more target product properties. Suitable machine learning modules can include artificial neural networks, such as deep neural networks (DNNs), symbolic regression, recurrent neural networks (RNNs) including long short-term memory (LSTM) networks or gated recurrent unit (GRU) networks, decision trees, random forests, boosting trees such as gradient boosting trees (XGBoost), linear regression, partial least squares regression, support vector machines, multilayer perceptrons (MLPs), autoencoders (e.g., denoising autoencoders such as stacked denoising autoencoders), Bayesian networks, support vector machines (SVMs), hidden Markov models (HMNIs), etc. Commercially available software packages can include JMP software, Microsoft AzureML, SAP data analysis tools, Sartorius's Simulation-like Soft Independent Modeling (SIMCA), etc.
[0062] Machine learning architectures can also leverage deep learning, where neural networks with multiple layers are generated. These layers sequentially extract higher-order features from the training dataset. For chemical formulations, examples of layers containing lower-order features could include general classifications of component types, while layers in the network that include higher-order features could include details dependent on functional groups, ionization states, charges, etc.
[0063] Specifically about Figure 3 The machine learning module disclosed herein can be trained using a training set consisting of historical and / or hypothetical data for a given chemical application (e.g., polyurethane foam). At 302, one or more training datasets are constructed based on one or more variable parameters, including formulation components, descriptors, process conditions, and composition properties.
[0064] The training dataset may also include one or more descriptors generated by converting the physical properties and concentrations (e.g., weight %) of components into descriptors (e.g., OH number, NCO, functionality, etc.) using suitable physics-based models / tools known in the art. The descriptors disclosed herein are calculated from the properties of individual components in a formulation, such as using a physical model suitable for a specific chemical application (e.g., polyurethane composition). Descriptors may contain data on formulation components, component concentrations, and ratios such as the ratio of polyurethane reactants (e.g., the ratio of isocyanate components to polyol components), product-forming reactions between components, and properties resulting from various chemical interactions (e.g., functionality of isocyanates or reactive species, crosslinking, foaming agent reactivity). Suitable descriptors also include those that detail mechanical properties (e.g., vapor heat capacity, foam density, Young's modulus, rheological properties, heat transfer properties, etc.).
[0065] At 304, feature selection is performed on the training dataset constructed in 302. During feature selection, a subset of the variable parameters identified in the training set are identified as "driving" variables that influence the properties of the target recipe. Feature selection may then involve excluding irrelevant, noisy, and redundant features from the training dataset.
[0066] Feature selection techniques may include one or more of the following: descriptor feature selection; constraint feature removal; correlation testing methods, such as Pearson, Spearman, Kendall, etc.; univariate tests of analysis of variance (ANOVA); mean absolute difference tests; L1 or minimum absolute contraction and selection operator (Lasso) regularization; multivariate analysis (baseline); etc.
[0067] At point 306, a machine learning model architecture is surveyed by training one or more machine learning models using the driving variables established from point 304 as input. The machine learning models generated from the surveyed architectures are then compared and their accuracy is rated. This accuracy rating is then used to select one or more model architectures to be used in subsequent stages. For example, the machine learning module can output one or more predicted performance characteristics from a predicted chemical formulation. In some cases, more than one trained machine learning model can be combined into a machine learning module, where the output is the result of the constituent machine learning models that have higher accuracy for the selected target product characteristics and / or the result of averaging the outputs of one or more machine learning models.
[0068] At point 308, the method further includes: training and validating multiple models using a test dataset containing variable data and target product characteristic data, and then basing the results on expected model criteria (such as those based on error calculation techniques such as R). 2The best fit (mean percentage error (MAPE), root mean square error (RMSE), etc.) is used to select an appropriate model. The test dataset may contain chemical formula information and descriptor information that are structurally similar to the training dataset; however, it usually contains sample information with the lowest degree of repetition with the training data in order to provide a sufficient test of the machine learning module's ability to handle new information, rather than storing and retrieving training dataset values.
[0069] Underfitting occurs when a model fails to capture the relationship between input and output variables. In R-based systems... 2 In the case of this method, underfitting models tend to have undesirably low training R-values. 2 Training R 2 Thresholds can be set to filter out underfitting models based on desired accuracy, such as greater than 0.70, 0.85, 0.88, or 0.9. Conversely, overfitting models may occur, where the model is too closely aligned with the training data and the learned representation cannot accurately predict the validation data. In some cases, error calculation techniques can also be used to select generalizable models while filtering out overfitting models. For example, validation RMSE percentage thresholds of less than 40%, 35%, or 30% can be used to predict target product characteristics.
[0070] In some cases, by minimizing the actual value Compared with the predicted value The mean squared error (MSE) between the two sides is used to optimize the model parameters, as shown in the following equation (I):
[0071]
[0072] in It is the total number of data points. It is the actual value of the measured target product characteristic. These are sample data points, and These are the model predictions of the target product's characteristics.
[0073] MSE is a measure of the average of the squared differences between predicted and actual values. RMSE stands for Root Mean Square Error, which is also used to measure the difference between predicted and actual values. RMSE percentage is the square root of MSE, normalized to the population mean and expressed as a dimensionless percentage. RMSE percentage can be calculated as follows:
[0074]
[0075] in As defined in equation (I) above, and It is the average value of the measured target product characteristics. The lower the RMSE percentage, the better the model fits the dataset.
[0076] R 2 Rdiscriminant represents the discriminant coefficient, a commonly used performance metric for regression, used to calculate the proportion of variance explained by the regression model. Typically, Rdiscriminant... 2 The range is from 0 to 1, and it can be calculated according to the following equation (III):
[0077]
[0078] in As defined in equation (I) above, and It is the average value. R 2 The higher the value, the better the model fits the dataset.
[0079] In one implementation, the method can provide predictions of target product properties of a chemical composition or product, such as the “training R” calculated according to equation (III) above. 2 >0.70” and the test RMSE percentage calculated according to Equation (II) above is less than 30% (<30%) as indicated. Additionally, the validation RMSE percentage is <30%, as calculated according to Equation (II) above. “Training R 2 "Refers to the R of the training dataset" 2 "Validation RMSE percentage" refers to the RMSE percentage of the validation dataset, and "Test RMSE percentage" refers to the RMSE percentage of the test dataset.
[0080] At point 310, bootstrapping (bagging) or other methodologies are used to generate prediction intervals for the created model to provide a measure of the possible error in predicting the target product characteristics. For example, a prediction interval can be represented as the range of variance around the associated mean. During bootstrapping, samples are drawn from the training data and fed into the model, and the results are combined using an average for regression and simple voting for classification to obtain an overall prediction. Other suitable prediction interval techniques include those described in open-source libraries, such as Model-Agnostic Prediction Interval Estimator (MAPIE), which includes conformal prediction, naive methods, segmentation methods, folding knife methods, CV methods, Ensemble Batch Prediction Intervals (EnbPI) methods, etc.
[0081] At point 312, model optimization is used to identify driving variables to determine whether removing or adding additional variables improves the obtained model accuracy. The optimization method evaluates the importance of each selected variable in determining the predictions of the model's output. Then, a subset of the variables with the greatest impact on model accuracy can be used to determine whether training process 300 is complete (e.g., prediction accuracy is acceptable) or whether it should be iterated by repeating steps 304 through 312. In some cases, the iterations may be repeated multiple times, such as two to three or more times.
[0082] Optimization methods can evaluate selected variables based on the characteristics of the target product derived from the model output and processed using interpretation software such as Shapley Additive Ex-Planation (SHAP) and Local Interpretable Model-Agnostic Explanations (LIME). For example, SHAP analysis assigns a specific predictive importance value to each input variable by comparing the model output with and without the specific variable. The SHAP value is then calculated using a summation (which represents the average effect of each input variable added to the model across all possible rankings of the introduced variable). A positive SHAP value indicates that higher eigenvalues, on average, lead to higher predicted values, while a negative SHAP value indicates that lower eigenvalues, on average, lead to lower predicted values.
[0083] At point 314, an optimized and trained machine learning module is deployed, and this optimized and trained machine learning module is used to generate predicted target product characteristics of the desired formulation.
[0084] Figure 4 An example is illustrated of a method for generating one or more target product properties from a desired recipe using a trained machine learning module, such as the method employed in 108 and 208. At 402, the desired recipe is selected as input to the trained machine learning module so that one or more predicted target product properties are computed at 404.
[0085] At point 406, the predicted characteristics can be used to determine whether the evolutionary formulation library meets the desired performance criteria, or whether further modifications to the intended formulation are necessary. Performance criteria can vary depending on the polymer system the formulation targets (e.g., polyurethane, polyurethane foam); specific applications (e.g., bedding and padding, insulation, sound absorption, energy efficiency applications, etc.); and the presence of other factors such as the need to minimize cost, carbon footprint, and use of hazardous substances. The evolutionary formulation library obtained from point 406 can be advanced as candidate formulations for experimental validation or evaluation of existing processes used to produce products or articles.
[0086] The computer system disclosed herein can coordinate the input, selection, and directed evolution of a recipe library, including the integration of ML modules for generating predicted target product properties at each step. The computer system may include a processor and a data storage device, wherein the data storage device stores computer-executable instructions that, when executed by the processor, cause the computing device to perform functions. The computing device may include a client device (e.g., a device actively operated by a user), a server device (e.g., a device providing computing services to client devices), or some other type of computing platform. Some server devices may, from time to time, operate as client devices to perform specific operations, and some client devices may incorporate server features.
[0087] The processors that can be used in this disclosure can be one or more of any type of computer processing element, such as a central processing unit (CPU), a coprocessor (e.g., a math, graphics, neural network, or cryptographic coprocessor), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a network processor, and / or an integrated circuit or controller that performs processor operations.
[0088] The data storage device may include one or more data storage arrays, which in turn include one or more drive array controllers configured to manage read and write access to hard disk drives and / or solid-state drive groups.
[0089] Computing devices can be deployed to support cluster architectures. The exact physical location, connectivity, and configuration of these computing devices may be unknown and / or irrelevant to client devices. Therefore, computing devices can be referred to as "cloud-based" devices, which can be housed in various remote data center locations, such as cloud-based server clusters. Ideally, the computing device is a cloud-based server cluster, and the input of concentration data into the model is done via a web-based user interface accessible to the user.
[0090] Figure 5A schematic diagram depicts a cloud-based server cluster 500 according to an example of this disclosure. Ideally, the operation of computing devices can be distributed among server devices 502, data storage devices 504, and routers 506, all of which can be connected via a local cluster network 308. The number of server devices 502, data storage devices 504, and routers 506 in the server cluster 500 can depend on the computing tasks and / or applications assigned to the server cluster 500. For example, server devices 502 can be configured to perform various computing tasks. Therefore, computing tasks can be distributed among one or more server devices 502. For example, data storage device 304 can store databases of any form, such as Structured Query Language (SQL) databases. Furthermore, any database in data storage device 504 can be monolithic or distributed across multiple physical devices. Router 506 may include a networking device configured to provide internal and external communication to the server cluster 500. For example, router 506 may include one or more packet switching devices and / or routing devices (including switches and / or gateways) configured to (i) provide network communication between server device 502 and data storage device 504 via cluster network 508, and / or (ii) provide network communication between server cluster 500 and other devices via communication link 510 to network 512. Server device 502 may be configured to transmit data to and receive data from cluster data storage device 504. Furthermore, server device 502 may have the ability to execute various types of computerized scripting languages, such as Perl, Python, PHP (Hypertext Preprocessor), Active Server Pages (ASP), or JavaScript. Computer program code written in these languages can facilitate the delivery of web pages to client devices and the interaction between client devices and web pages.
[0091] Example
[0092] The following examples are provided to illustrate embodiments of the invention, but are not intended to limit the scope of the invention. Table 1 provides the materials used in the following examples.
[0093] Example 1: Generating the desired polyurethane foam formulation using the GA methodology
[0094] In this embodiment, the GA methodology is used to generate an evolutionary formulation library containing formulations that produce PU foams with the foam properties shown in Table 1. The genetic algorithm implementation is prepared in ChemML5. Bit vectors encode components and their amounts per hundred parts of polyol, which act as “genes” during directed evolution.
[0095]
[0096]
[0097]
[0098] Example 2: Design of Thermal Insulation Polyurethane Foam Formulation
[0099] Using the same GA methodology, we also targeted formulations with different foam properties (see table below). This demonstrates that this reverse design methodology can be extended to different segments of the polyurethane business.
[0100]
[0101] Based on the input target characteristics, five formulations that meet the criteria are generated for potential future validation. Components are displayed in pphp.
[0102]
[0103]
[0104] While the foregoing relates to exemplary embodiments, other and additional embodiments may be devised without departing from the basic scope of the invention, the scope of which is defined by the appended claims.
Claims
1. A method for generating a desired library of chemical formulations, comprising: (a) Define one or more target product characteristics for the intended formulation library; (b) Prepare an initial formula library from historical data that includes the characteristics of one or more historical products; (c) Apply a fitness function to determine the difference between the characteristics of the one or more historical products and the characteristics of the one or more target products, and discard recipes that do not satisfy the fitness function from the initial recipe library; (d) Developing an evolutionary recipe library from the initial recipe library using a genetic algorithm-based methodology, including: i. Select one or more parent formulations from the initial formulation library that have a desired similarity to the characteristics of the target product according to the fitness function; ii. Convert each of the one or more parent recipes into a bit vector; iii. Generate one or more offspring recipes from the one or more parent recipe bit vectors using mutation and / or crossover operations; and iv. Collect one or more offspring recipes to create the evolutionary recipe library; (e) Generate descriptors for one or more offspring recipes in the evolutionary recipe library; (f) Using a trained machine learning module to analyze the evolutionary recipe library to calculate one or more predictive product properties; (g) Applying the fitness function to the evolutionary recipe library to determine the differences between the characteristics of the one or more target products and the characteristics of the one or more predicted products; and (h) Based on the differences identified in (g): i. Collect a subset of the evolutionary formulation library to generate the expected formulation library that has a desired similarity to the characteristics of the target product; or ii. Repeat steps (d) to (g) until the desired one or more target product characteristics or stopping criteria of the expected formulation library are met.
2. The method according to claim 1, further comprising selecting one or more formulations from the expected formulation library for experimental verification.
3. The method of claim 2 further includes using a validated formula library to modify the chemical process.
4. The method according to claim 1, wherein the chemical formulation is a polyurethane formulation.
5. The method of claim 1, wherein the descriptor is specific to polyurethane or polyurethane foam.
6. The method of claim 1, wherein the one or more historical product characteristics are derived from experimental data or predicted from historical data using a trained machine learning module.
7. The method of claim 1, wherein a child recipe is generated from one or more parent recipe bit vectors using mutation and crossover operations.
8. The method of claim 1, wherein the stopping criterion is selected from the group consisting of: reaching a proximity threshold (Δ) to the desired target product characteristics, any selected number of generations, and no improvement for 10 consecutive generations.
9. The method of claim 1, wherein crossing comprises exchanging one or more components and / or component categories between the parent or child formulation bit vectors.
10. The method of claim 1, wherein the mutation comprises a change in at least one component, component category, component concentration, or random addition.