Digital coating formulation
Patent Information
- Application Number
- PCT/US2026/015445
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-26
- Filing Date
- 2026-02-17
- Publication Date
- 2026-09-03
Smart Images

Figure US2026015445_03092026_PF_FP_ABST
Abstract
Description
1079-127US01 / US231091DIGITAL COATING FORMULATIONTECHNICAL FIELD
[0001] This disclosure relates to a computer-implemented method for predicting coating composition performance characteristics, or defining a coating composition formulation, or defining a manufacturing process to meet desired coating composition characteristics. This disclosure further relates to systems for performing the computer-implemented method, coating composition formulations, and manufacturing processes defined according to the computer-implemented method.BACKGROUND
[0002] A formulation of a coating composition refers to the components of the composition, the proportions of the components in the composition, and as well as manufacturing processes and processing conditions. Paints and stains are one variety of coating compositions used for architectural applications. The components of a coating composition generally include a binder, a solvent, and additives, with or without additional pigment or filler. Pigments provide color and / or hide to the coating. Pigment volume concentration (PVC) is the percentage of the volume concentration of the pigments in the coating as compared to all solids of the paint. A binder, also known as a resin, is a natural or synthetic polymer that upon curing, coalesces to hold the pigments, fillers, and other additives in the paint in a solid film. A solvent is the carrier for the coating composition components. Example solvents include water, oil, or organic solvents, with or without additional organic solvents and combinations thereof. Additives are used to change the properties of a coating composition. Example types of additives include curing agents, coalescents, extenders, defoamers, driers, ultraviolet (UV) stabilizers, and so on. Typically distributed in liquid or powdered form, coating compositions are applied to a previously coated or uncoated substrate such as wood, drywall, vinyl siding, metal, or concrete by brush, roller, or spray, and cured. Curing can occur by evaporation of the solvent and coalescence, by exposure to UV, electron-beam or other radiation source, by application of heat, or a combination thereof.
[0003] Coating compositions are formulated to provide desirable mechanical and chemical properties in both the liquid state and the cured state. Taking paints as an example, the properties of paints in the liquid state include viscosity, VOC level, smell, and color. The properties of paints in the cured state may include hardness, flexibility, durability (e.g., resistance to scratching or1079-127US01 / US231092scuffing), and resistance to various types of stains (e.g., wine stains, mustard stains, etc.). In addition, paints and other coating compositions may provide chemical resistance in the cured state, to oxidation, corrosion, or resistance to acid or base attack, or resistance of fungal defacement. The chemical compositions of the coating composition components and their proportions affect these properties. Desired coating properties vary by the end-use.
[0004] Coating compositions, however, are complex mixtures of components. The amount and type of each component can impact the contribution of other components to the coating composition properties.
[0005] Each coating composition formulation involves a large number (e.g., hundreds) of variables describing component types and their proportions. Since each different combination of the variables is a different coating composition formulation, the number of potential coating composition formulations is very large. Because the number of potential coating composition formulations is very large, the process of identifying a coating composition formulation having a desired set of characteristics is a time consuming and expensive process. The process typically involves a user (e.g., a technician, scientist, chemist, etc.) identifying a set of promising coating composition formulations that may produce the desired set of characteristics. Some initial focusing based on the user’s experience is typically involved in this identification process. After identifying the set of coating composition formulations, experiments are run in which these coating composition formulations are prepared and tested. Experimental data characterizing the mechanical, chemical, and application properties of the resulting coating compositions from the experiments is recorded. The experimental data may include lab measured data, such as tensile data, blocking data, washability data, Martindale data (i.e., measures of durability and wear resistance), and surface energy data. The experimental data may also include calculated data, such as polymer Hansen solubility parameters and coalescent Hansen solubility parameters. Users use the recorded experimental data to refine the coating composition formulations, and the process repeats.
[0006] Since coating composition formulations involve large numbers of variables, it is impractical to run experiments that test variations on all of the variables. Rather, specific key variables are identified, and experiments are run that test variations on these key variables while keeping the other variables fixed. Statistical analyses are performed to identify the key variables.1079-127US01 / US231093For instance, variables may be removed from consideration based on their p-values, R2 values, variance inflation factors (VIFs), and domain knowledge.
[0007] The variables and the resulting experimental data may be stored in a table in which each column corresponds to a different variable and each row corresponds to a different experiment. Because relatively few experiments are performed compared to the number of potential variables, the table may be characterized as “wide” but not “long.” Accelerating the process of identifying promising coating composition formulations is desirable in terms of cost efficiency and business agility.
[0008] There are several challenges associated with using machine learning (ML) models to identify promising coating composition formulations. For example, neural network models are powerful predictive tools but training neural network models typically requires numerous training examples. As mentioned above, the experimental data has relatively few rows available for use as training examples. Other modeling techniques, such as linear regression models, that can be trained using fewer training examples provide poor results or are very difficult because the experimental data has a large number of features. Moreover, some ML models operate like “black boxes,” and it is difficult to understand how the ML models generated their outputs and why specific formulation changes cause the observed changes in coating performance.
[0009] Thus, a system that incorporates multiple models to predict coating composition performance characteristics is desirable.SUMMARY
[0010] The present disclosure describes devices, systems, and techniques for estimating values of properties of coating compositions, and coating composition formulations and manufacturing processes defined thereby. A manufacturing process may be controlled based on whether a candidate coating composition formulation has desired values of one or more properties, whether they be mechanical properties, chemical properties, or application properties. In this disclosure, a coating composition formula is said to have a value of a property if a coating produced according to the coating composition formula has the value of the property. Additionally, the techniques provide for various advantageous user interfaces. As described herein, a computing system may obtain an input dataset comprising entries for a plurality of coating composition formulations. For each coating composition formulation of the plurality of coating composition formulations, the entry for the coating composition formulation includes values of one or more formulation variables1079-127US01 / US231094of the coating composition formulation and measured values of one or more properties of the coating composition formulation. The computing system may train a clustering model to partition the coating composition formulations of the input dataset into a plurality of clusters based on measured values of one or more target properties. The one or more properties of the coating composition formulations include the one or more target properties. Additionally, the computing system may train a membership prediction model to predict, based on values of one or more condition-basis variables, whether a coating composition formulation is a member of a selected cluster of the plurality of clusters. Each of the one or more condition-basis variables is one of the formulation variables of the coating composition formulations. The computing system may apply the prediction model to predict whether a candidate coating composition formulation has values of the one or more properties associated with the selected cluster. The candidate coating composition formulation may have one or more values of the one or more formulation variables different from the values of the formulation variables of the entries for the coating composition formulations in the input dataset. Based on the prediction model predicting that the candidate coating composition formulation has the properties associated with the selected cluster, the computing system may configure a manufacturing system to prepare a coating composition according to the candidate coating composition formulation.
[0011] In one aspect, this disclosure describes a computer-implemented method comprising: obtaining, by one or more processors of a computing system, an input dataset comprising entries for a plurality of coating composition formulations, wherein for each coating composition formulation of the plurality of coating composition formulations, the entry for the coating composition formulation includes values of one or more formulation variables of the coating composition formulation and measured values of one or more properties of the coating composition formulation; training, by the one or more processors, a clustering model to partition the coating composition formulations of the input dataset into a plurality of clusters based on measured values of one or more target properties, wherein the clustering model is a first ML model and the one or more properties of the coating composition formulations include the one or more target properties; and training, by the one or more processors, a prediction model to predict, based on values of one or more condition-basis variables, whether a coating composition formulation is a member of a selected cluster of the plurality of clusters, wherein the prediction model is a second1079-127US01 / US231095ML model and each of the one or more condition-basis variables is one of the formulation variables of the coating composition formulations.
[0012] In another aspect, this disclosure describes a system comprising: one or more data storage units configured to store an input dataset comprising entries for a plurality of coating composition formulations, wherein for each coating composition formulation of the plurality of coating composition formulations, the entry for the coating composition formulation includes values of one or more formulation variables of the coating composition formulation and measured values of one or more properties of the coating composition formulation; one or more processors implemented in circuitry and communicatively coupled to the one or more data storage units, the one or more processors configured to: train a clustering model to partition the coating composition formulations of the input dataset into a plurality of clusters based on measured values of one or more target properties, wherein the clustering model is a first ML model and the one or more properties of the coating composition formulations include the one or more target properties; and train a prediction model to predict, based on values of one or more condition-basis variables, whether a coating composition formulation is a member of a selected cluster of the plurality of clusters, wherein the prediction model is a second ML model and each of the one or more condition-basis variables is one of the formulation variables of the coating composition formulations.
[0013] In another aspect, this disclosure describes one or more non-transitory computer-readable data storage media having processor-executable instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: store, in one or more data storage devices an input dataset comprising entries for a plurality of coating composition formulations, wherein for each coating composition formulation of the plurality of coating composition formulations, the entry for the coating composition formulation includes values of one or more formulation variables of the coating composition formulation and measured values of one or more properties of the coating composition formulation; train a clustering model to partition the coating composition formulations of the input dataset into a plurality of clusters based on measured values of one or more target properties, wherein the clustering model is a first ML model and the one or more properties of the coating composition formulations include the one or more target properties; and train a prediction model to predict, based on values of one or more condition-basis variables, whether a coating composition formulation is a member of a selected cluster of the plurality of1079-127US01 / US231096clusters, wherein the prediction model is a second ML model and each of the one or more condition-basis variables is one of the formulation variables of the coating composition formulations.
[0014] The details of one or more aspects of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the techniques described in this disclosure will be apparent from the description, drawings, and claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0015] FIG. 1 is a block diagram illustrating example components of a system that includes a computing system, in accordance with one or more aspects of this disclosure.
[0016] FIG. 2 is a flowchart illustrating an example operation in accordance with one or more aspects of this disclosure.
[0017] FIG. 3 is a flowchart illustrating an example operation for predicting whether a candidate coating composition formulation has values of one or more properties associated with the selected cluster, in accordance with one or more techniques of this disclosure.
[0018] FIG. 4 is a conceptual diagram illustrating an example user interface in accordance with one or more techniques of this disclosure.DETAILED DESCRIPTION
[0019] This technology relates to the use of ML models to identify promising coating composition formulations. Specifically, the technology involves a multi-model system that identifies variables in coating composition formulations that affect a target property of the resulting coating compositions. As described herein, a computing system may obtain an input dataset comprising entries for a plurality of coating composition formulations. For each coating composition formulation of the plurality of coating composition formulations, the entry for the coating composition formulation includes values of one or more formulation variables of the coating composition formulation and measured values of one or more properties of the coating composition formulation. Additionally, the computing system may train a clustering model to partition the coating composition formulations of the input dataset into a plurality of clusters based on measured values of one or more target properties. The one or more properties of the coating composition formulations include the one or more target properties. Additionally, the computing1079-127US01 / US231097system may train a prediction model to predict, based on values of one or more condition-basis variables, whether a coating composition formulation is a member of a selected cluster of the plurality of clusters. Each of the one or more condition-basis variables is one of the formulation variables of the coating composition formulations. The computing system may apply the prediction model to predict whether a candidate coating composition formulation has values of the one or more properties associated with the selected cluster. The candidate coating composition formulation may have one or more values of the one or more formulation variables different from the values of the formulation variables of the entries for the coating composition formulations in the input dataset. Training the clustering model and prediction model in this way may support improved manufacturing processes, improved user interfaces, improved understanding of how coating composition formulations are recommended, and so on. Additionally, training the clustering model and prediction model in this way may help users understand what material properties influence whether a coating has a good performance since the prediction model was created with material properties as the features.
[0020] FIG. 1 is a block diagram illustrating example components of a system 100 that includes a computing system 102, in accordance with one or more aspects of this disclosure. In the example of FIG. 1, system 100 also includes a manufacturing system 104. FIG. 1 illustrates only one particular example of computing system 102, and many other example configurations of computing system 102 exist.
[0021] Computing system 102 includes one or more computing devices, each of which may include one or more processors. For instance, computing system 102 may include one or more mobile devices (e.g., smartphones, tablet computers, etc.), server devices, personal computer devices, handheld devices, or other types of devices. Actions described in this disclosure as being performed by computing system 102 may be performed by one or more of the computing devices of computing system 102.
[0022] As shown in the example of FIG. 1, computing system 102 includes one or more processors 112, one or more communication units 114, one or more storage devices 116, one or more input devices 118, one or more output devices 120, a power source 128, and one or more communication channels 126. Output devices 120 may include a display device 122. Computing system 102 may include other components. For example, computing system 102 may include physical buttons, microphones, speakers, communication ports, and so on.1079-127US01 / US231098
[0023] Communication channels 126 may interconnect each of processors 112, communication units 114, input devices 118, output devices 120, display device 122, and storage devices 116 (physically, communicatively, and / or operatively). In some examples, communication channels 126 may include a system bus, a network connection, an inter-process communication data structure, or any other method for communicating data. Power source 128 may provide electrical energy to one or more of processors 112, communication units 114, input devices 118, output devices 120, display device 122, and storage devices 116.
[0024] Storage devices 116 may store information required for use during operation of computing system 102. In some examples, storage devices 116 have the primary purpose of being a shortterm and not a long-term computer-readable storage medium. Storage devices 116 may include volatile memory and may therefore not retain stored contents if powered off. In some examples, storage devices 116 includes non-volatile memory that is configured for long-term storage of information and for retaining information after power on / off cycles. In some examples, processors 112 of computing system 102 may read and execute instructions stored by storage devices 116.
[0025] Computing system 102 may include one or more input devices 118 that computing system 102 uses to receive user input. Examples of user input include tactile, audio, and video user input. Input devices 118 may include presence-sensitive screens, touch-sensitive screens, mice, keyboards, voice responsive systems, microphones, motion sensors capable of detecting gestures, or other types of devices for detecting input from a human or machine.
[0026] Communication units 114 may enable computing system 102 to send data to and receive data from one or more other computing devices (e g., via a communication network, such as a local area network or the Internet). Examples of communication units 114 may include network interface cards, Ethernet cards, optical transceivers, radio frequency transceivers, or other types of devices that are able to send and receive information. Other examples of such communication units may include BLUETOOTH™, 3G, 4G, 5G, and WI-FI™ radios, Universal Serial Bus (USB) interfaces, etc.
[0027] Output devices 120 may generate output. Examples of output include tactile, audio, and video output. Output devices 120 may include presence-sensitive screens, sound cards, video graphics adapter cards, speakers, liquid crystal displays (LCD), light emitting diode (LED) displays, or other types of devices for generating output. Output devices 120 may include a display1079-127US01 / US231099screen. In some examples, output devices 120 may include virtual reality, augmented reality, or mixed reality display devices.
[0028] Manufacturing system 104 may include devices for manufacturing coating compositions. For example, manufacturing system 104 may include raw material storage tanks, mixing tanks, dispersers, emulsifiers, filters, filling machines, and control systems. The control systems may control mixing time and speed, amounts of materials added, timing and order of materials added, temperatures of materials, and so on. In some examples, manufacturing system 104 may include different types of mixing vessel and / or different types of mixing blades. The control systems may control which materials are added to which types of mixing vessels and which types of mixing blades are used to mix the materials.
[0029] Processors 112 may read processor-executable instructions from storage devices 116 and may execute the processor-executable instructions stored by storage devices 116. Execution of the instructions by processors 112 may configure or cause computing system 102 to provide at least some of the functionality ascribed in this disclosure to computing system 102 or components thereof (e.g., processors 112). As shown in the example of FIG. 1, storage devices 116 include computer-readable instructions associated with a training system 130 and a prediction system 132. In some examples, training system 130 and prediction system 132 may be included in different computing devices and / or different computing systems. In addition, storage devices 116 may store a tabular input dataset 134, a clustering model 136, and a prediction model 138.
[0030] In a first step, computing system 102 obtains input dataset 134. Input dataset 134 may be generated from previous experiments. Input dataset 134 may represent tabular data. Columns of input dataset 134 may correspond to formulation variables of coating composition formulations and measured properties. Rows of input dataset 134 may correspond to experiments run on the different coating composition formulations. Thus, each entry (e.g., row) in input dataset 134 is associated with a respective coating composition formulation and the entry for a respective coating composition formulation includes values of formulation variables of the respective coating composition formulation and values of measured properties of the coating composition formulation. In other examples, rows and columns may be switched.
[0031] In a second step, training system 130 may perform a data cleaning process on input dataset 134 to generate a cleaned input dataset. For example, training system 130 may remove anomalies from the tabular input data. In some examples, training system 130 applies an isolation forest1079-127US01 / US2310910algorithm to remove anomalies from the tabular input data. The isolation forest algorithm is an unsupervised machine learning algorithm used for anomaly detection. In other examples, training system 130 uses one or more other types of algorithms to remove anomalies from the tabular input data.
[0032] In a third step, training system 130 may train clustering model 136. Clustering model 136 is a model that has been trained to identify whether a coating composition formulation belongs to one of a plurality of clusters. For example, clustering model 136 may be a -means clustering model, a £-means++ model, a support vector clustering model, a density-based clustering model (e.g., a density -based spatial clustering of applications with noise (DBSCAN) model), or another type of clustering model. For ease of explanation, this disclosure primarily describes clustering model 136 in terms of a £-means clustering model. In such examples, training system 130 may partition the coating composition formulations of the cleaned input dataset 134 into a plurality of k clusters based on a set of one or more target properties such as mechanical properties (e.g., abrasion or burnish resistance, scratch resistance, adhesion), appearance properties (e.g., hide, tint strength, gloss / sheen), chemical properties (e.g., washability, resistance to stains, chemical resistance, food contact qualification, VOC content), and application properties (e.g., brushability, sprayability, and surface quality of applied coating). The target properties may include each or a subset of the measured properties of the coating composition formulations of input dataset 134. For example, each coating composition formulation may be characterized as a datapoint in a Euclidean coordinate space defined by measurements of target properties. The number of clusters (i.e., the value of k) is a value greater than 1. In some examples, the number of clusters may be set by a user or may be a predetermined value.
[0033] By training the clustering model 136, training system 130 identifies clusters of coating composition formulations having generally similar values of the target properties. This allows a user (or the computing system 102) to select a cluster of coating composition formulations having desirable properties. For example, the target properties of the coating composition formulations may include a first target property related to washability for tea stains and a second target property related to washability for wine stains. In this example, coating composition formulations in a first cluster may have high measured values of the first target property (e.g., high washability for tea stains) and low measured values of the second target property (e.g., poor washability for wine stains). Conversely, in this example, coating composition formulations in a second cluster may1079-127US01 / US2310911have low measured values of the first target property (i.e., poor washability for tea stains) and high measured values of the second target property (i.e., high washability for wine stains). In this example, the user may select the first cluster if washability for tea stains is more important for the user’s goals, or may select the second cluster if washability for wine stains is more important for the user’s goals. In another example, the target properties of the coating composition formulations may include multiple target properties related to washability. In this example, coating composition formulations in a first cluster may have high measured values of the target properties related to washability. In this example, coating composition formulations in a second cluster may have relatively lower measured values of the target property for washability. Note that it is often not possible to have a coating composition that has all possible desirable properties, so trade-offs and selections may need to be made. However, after cluster identification, it may not yet be apparent which of the formulation variables are responsible for coating composition formulations being included in or being excluded from the selected cluster.
[0034] Hence, in a fourth step, training system 130 may train prediction model 138 to predict, based on values of one or more condition-basis variables, whether a coating composition formulation is a member of a selected cluster of the plurality of clusters. Each of the one or more condition-basis variables is one of the formulation variables of the coating composition formulation. The set of condition-basis variables may be a subset of the formulation variables. In other words, not all of the formulation variables of the coating composition formulations in input dataset 134 are included in the condition-basis variables. Prediction model 138 may be implemented in one of several ways. For example, prediction model 138 may comprise a decision tree model. In some examples, prediction model 138 may comprise one or more neural network models, random forest models, gradient tree boosting models, LightGBM models, XGBoost models, Support Vector Machine (SVM) models, and so on.
[0035] Thus, by using clustering model 136, a set of existing coating composition formulations having a set of desirable target properties may be identified, and by using prediction model 138, computing system 102 can identify which ones of the formulation variables (i.e., the conditionbasis variables) of the coating composition formulations are most significant for controlling whether a coating composition formulation ends up having the set of target properties. The identified formulation variables are the variables used as conditions for nodes in the decision tree model.1079-127US01 / US2310912
[0036] A user or prediction system 132 may use the resulting prediction model 138 to guide selection of a candidate coating composition formulation, such as for further testing, with a higher likelihood of having the target properties. Therefore, the user or prediction system 132 may design experiments that vary the values of the identified formulation variables (i.e., values of the condition-basis variables) and use prediction model 138 to determine whether a candidate coating composition formulation having the varied values of the identified formulation variables would fall within the selected cluster of coating composition formulations (i.e., the set of coating composition formulations having the target set of properties). For example, prediction system 132 or the user may determine a candidate coating composition formulation having an altered value of the one or more condition -basis variables. In other words, the values of one or more of the formulation variables may be different from values of the formulation variables of any of the coating composition formulations in the input dataset. Prediction system 132 may apply prediction model 138 to predict whether the candidate coating composition formulation has properties associated with the selected cluster.
[0037] Determining whether a candidate coating composition formulation has values of the properties associated with the selected cluster in this way may be more computationally efficient than other machine learned models for similar purposes. For instance, training clustering model 136, followed by training the prediction model 138 based on the clusters identified using clustering model 136 may allow the computing system to identify the condition-basis variables. Subsequently, when evaluating whether a coating composition formulation is likely to have desired values of one or more properties, computing system 102 need not process any information with respect to the potentially numerous formulation variables other than the condition-basis variables. Restricting evaluation to the condition-basis variables may therefore conserve computational resources. Additionally, the representation of prediction model 138 may be considerably smaller than other machine-learned models, such as neural networks trained for the same purpose. Furthermore, the techniques of this disclosure may be integrated into a process that controls manufacturing system 104. For instance, based on prediction model 138 predicting that the candidate coating composition formulation has the properties associated with the selected cluster, prediction system 132 may configure manufacturing system 104 to prepare a coating composition according to the candidate coating composition formulation. Thus, based on prediction model 138 predicting that the candidate coating composition formulation has the properties associated with1079-127US01 / US2310913the selected cluster, manufacturing system 104 may manufacture a coating composition according to the candidate coating composition formulation. This may be extremely useful in producing coating composition samples on which to conduct experiments to confirm the properties of the coating composition formulations. It is much less likely that coating composition samples that do not have the desired properties will be produced.
[0038] In some examples, prediction system 132 may apply prediction model 138 multiple times to identify a set of multiple coating composition formulations that have the values of the one or more properties associated with a selected cluster. Prediction system 132 may configure manufacturing system 104 to produce coating composition samples based on the identified set of coating composition formulations. In this way, manufacturing system 104 may automatically produce a set of coating composition samples for experiments based on the values of the properties of a selected cluster. Manufacturing system 104 may include one or more robotic devices and perform automated laboratory testing through robotics. Producing coating composition samples in this way may address technical problems with existing processes for manufacturing coating composition samples. For example, existing processes for manufacturing coating composition samples may produce coating composition samples that are unlikely to have desired values of one or more properties, leading to a waste of time and resources. Controlling manufacturing system 104 to produce a coating composition based on a coating composition formula determined according to techniques of this disclosure may reduce such waste, especially when large numbers of samples are to be produced. The techniques may thus filter inapplicable coating composition formulas from a large batch of coating composition formulas prior to manufacture.
[0039] In some examples, prediction system 132 may cause display device 122 to display a user interface 140. User interface 140 may include selectable features corresponding to the clusters identified by training clustering model 136. User interface 140 may also specify the values of properties associated with the clusters. Thus, user interface 140 may easily enable the user to select a set of desirable values of properties from among a set of automatically generated sets of values of properties.
[0040] Furthermore, user interface 140 may include input features (e.g., text boxes, drop-boxes, radio buttons, etc.) for input of values of formulation variables of a candidate coating composition formulation. In some examples, user interface 140 only displays input features associated with the condition-basis variables of prediction model 138 instead of all of the formulation variables. In1079-127US01 / US2310914some examples, input features associated with input of values of formulation variables other than the condition-basis variables may be deprioritized, hidden until requested to be shown by the user, grayed-out, indicated as not being relevant to the properties, or otherwise made less prominent in user interface 140 than input features associated with the condition-basis variables. In some examples, input elements associated with the condition-basis variables are separated (spatially or temporally) in user interface 140 from input elements associated with formulation variables other than the condition-basis variables. In this way, user interface 140 may be made more streamlined and efficient than conventional user interfaces for inputting coating composition formulations for purposes of the user experimenting with values of formulation parameters likely to affect values of properties of coating compositions. In such conventional user interfaces, input features may be displayed for each formulation variable, regardless of whether the formulation variables are relevant to achieving desired values of properties. Thus, in such conventional user interfaces, the user is given no indications of which formulation variables are relevant.
[0041] In some examples, prediction system 132 may prepopulate input elements associated with the condition-basis variables with values that would lead prediction system 132 to predict that a coating composition formulation having the values would have the values of the properties associated with the selected cluster. For instance, prediction system 132 may determine a set of values of the condition-basis variables that lead down a path in the decision tree model to a node associated with the selected cluster. In some such examples, the set of values may be midpoints or extreme points of ranges of the values of the condition-basis variables that lead down a path in the decision tree model to the node associated with the selected cluster. The user may be free to experiment with coating compositions by making adjustments relative to these values of the condition-basis variables. This may increase the ease of use of user interface 140 relative to conventional user interfaces because the user may be able to start from values of formulation variables that are expected to lead to the desired values of properties.
[0042] In some examples where user interface 140 includes input features associated with both condition-basis variables and other formulation variables, prepopulating values in the input elements associated with the condition-basis variables allows the user to experiment with coating composition formulations by adjusting values of the other formulation variables while still maintaining the expectation that coating compositions manufactured using such values of the other formulation variables have the values of the properties of the selected cluster. This may1079-127US01 / US2310915significantly increase the usability of user interface 140 for purposes of enabling the user to experiment with values of the other formulation variables. This, in turn, improves the functionality of computing system 102 for purposes of enabling a user to efficiently create coating composition formulations for testing. In some such examples, the values of the condition-basis variables may be locked while allowing the user to use input elements of user interface 140 to adjust values of the other formulation variables. This may further increase ease of use.
[0043] The techniques of this disclosure may be used in projects for reformulating coatings. The capabilities of prediction system 132 may be especially useful in such projects because most reformulation projects do not start from scratch and consider replacing every raw material. Usually, these projects involve finding a replacement for one or two raw materials because they have been discontinued or need to be removed due to regulatory reasons. Other projects may be trying to improve one or two properties while maintaining the performance in all other respects. So, typically formulators will know to change certain types of raw materials to address these one or two properties, without needing to change all types of raw materials.
[0044] In some examples, prediction system 132 may be used to computationally review other formulas from a large dataset of formulas to determine which formulas would be most susceptible to property changes based on a particular change in an input variable (e.g., a raw material change or quantity change). This would be a more efficient way to determining the potential breadth of formulas to which a particular cluster would apply. In some examples, prediction system 132 may be used to identify formulas that already embody that particular values of a condition-based variable so as to identify formulas that already have that benefit.
[0045] FIG. 2 is a flowchart describing a training process in accordance with one or more techniques of this disclosure. The flowcharts of this disclosure are provided as examples. Other examples in accordance with the techniques of this disclosure may include more, fewer, or different steps, or the steps may be performed in different orders.
[0046] In the example of FIG. 2, computing system 102 obtains input dataset 134 (200). For example, computing system 102 may obtain input dataset 134 from another computing system, from a remote data store, from a computer-readable media, from manual user input, and / or from one or more other sources. Input dataset 134 may be formatted in a Comma Separated Value (CSV) format or another format.1079-127US01 / US2310916
[0047] Computing system 102 may store input dataset 134 in one or more of storage devices 116. Input dataset 134 may be represented in a tabular format. Columns of the tabular input data include columns that correspond to formulation variables of coating composition formulations and measured properties. Rows of the tabular input data include rows that correspond to experiments run on the different coating composition formulations. Thus, each entry (e.g., row) in the input dataset is associated with a respective coating composition formulation and the entry for a respective coating composition formulation includes values of formulation variables of the respective coating composition formulation and values of measured properties of the coating composition formulation. Example formulation variables may include each of the specific raw materials selected and their respective amounts (e.g. binders, thickeners, surfactants, dispersants, pigments, extenders, coalescents, plasticizers, carriers / solvents, and other additives), properties of those raw materials (e.g., density, HLB values, solubility parameters, particle size, etc.), and chemical composition of those raw materials (e.g., monomer ratios of polymeric binders and thickeners, surfactant compositions, dispersant compositions, etc.).
[0048] Furthermore, training system 130 may perform a data cleaning process on input dataset 134 (202). For example, training system 130 may apply an isolation forest model to input dataset 134 to remove anomalous coating composition formulations from the input dataset. A core idea of the isolation forest algorithm is the isolation principle. The isolation principle is an assumption that anomalies are few and different, making the anomalies easier to isolate. The isolation forest process isolates observations by randomly selecting a feature (e.g., a formulation variable or a mechanical property) and then randomly selecting a split value between the maximum and minimum values of the selected feature. The isolation forest algorithm builds multiple isolation trees (iTrees). Each isolation tree is constructed by recursively partitioning the data. The process continues until all data points are isolated or a predefined depth is reached. The number of splits required to isolate a data point is known as the path length. Anomalies, being distinct, tend to have shorter path lengths compared to normal data points. An anomaly score is calculated based on the average path length of a data point across all trees. A lower average path length indicates a higher likelihood of the point being an anomaly. Data points with anomaly scores above a certain threshold are flagged as anomalies.
[0049] Training system 130 may train clustering model 136 to partition the coating composition formulations of input dataset 134 into a plurality of clusters based on measured values of one or1079-127US01 / US2310917more target properties (204). The one or more properties of the coating composition formulations include the one or more target properties. In some examples, the one or more target properties include one or more properties related to washability of coating compositions produced using the coating composition formulations.
[0050] In some examples, training clustering model 136 involves initializing two or more (k) centroid points. Each point is then assigned to the closest of the centroid points. The position of each of the centroid points is then updated, e.g., by calculating a mean of the positions of the points assigned to the centroid point. The process of assigning points and updating the positions of the centroids is then repeated multiple times until a stopping criterion is reached. In other words, as part of training clustering model 136, training system 130 may represent the plurality of coating composition formulations as datapoints having coordinates in a coordinate space defined by the measured values of the one or more target properties. Additionally, training system 130 may initialize two or more centroid points having coordinates in the coordinate space. Training system 130 may iteratively update the coordinates of the centroid points based on distances between the coordinates of the centroid points and the coordinates of the datapoints that represent the coating composition formulations.
[0051] In some examples, training system 130 outputs, for presentation by one or more of output devices 120 (e.g., display device 122), information based on values of the one or more properties that correspond to the coordinates of the centroid points. For example, prediction system 132 may cause display device 122 to display a user interface that specifies the values of the properties that correspond to the coordinates of the centroid points. In this way, a user may learn expected values of the properties of the clusters.
[0052] Furthermore, in the example of FIG. 2, training system 130 may train prediction model 138 to predict, based on values of one or more condition-basis variables, whether a coating composition formulation is a member of a selected cluster of the plurality of clusters (206). Each of the one or more condition-basis variables is one of the formulation variables of the coating composition formulations. In some examples where prediction model 138 comprises a decision tree model, the process of training prediction model 138 includes the following steps:
[0053] 1) Add a root node to the decision tree model. All coating composition formulations of input dataset 134 are associated with the root node.1079-127US01 / US2310918
[0054] 2) For each node added to the decision tree model, evaluate conditions for dividing coating composition formulations associated with the node into two subsets based on the measured values of the target properties. For example, training system 130 may test different threshold levels of different formulation variables (i.e., condition-basis variables) to determine how to divide the coating composition formulations. For instance, in an example where PVC content is one of the target properties, training system 130 may determine that PVC > 47.76 divides the coating composition formulations associated with the node among those that are included in the selected cluster and those that are excluded from the selected cluster.
[0055] 3 ) Create two child nodes of the node, assigning each to a different condition (e g., first child node assigned to PVC > 47.76; second child node assigned to PVC < 47.76).
[0056] 4) Repeat from step 2 until no evaluation condition divides the remaining coating composition formulations associated with the node. For instance, training system 130 does not create any additional nodes if all of the coating composition formulations associated with the node are in the selected cluster or none of the coating composition formulations associated with the node are in the selected cluster.
[0057] Thus, in a first round of a decision tree training process, training system 130 may add a root node to a decision tree model. The root node is associated with each coating composition formulation in the plurality of coating composition formulations. Training system 130 may then perform one or more subsequent rounds of the decision tree training process after performing the first round of the decision tree training process. For each subsequent round of the one or more subsequent rounds and for at least one node added to the decision tree model in a previous round of the decision tree training process: training system 130 may determine whether there are threshold levels of the one or more condition-basis variables for dividing coating composition formulations associated with the node into one or more coating composition formulations associated with a first subset and one or more coating composition formulations associated with a second subset. The coating composition formulations associated with the first subset have measured values of the one or more target properties that correspond to the selected cluster, and the coating composition formulations associated with the second subset have measured properties of the one or more target properties that are excluded from the selected cluster. Based on there being threshold levels of the one or more condition-basis variables for dividing the coating composition formulations associated with the node into the coating composition formulations1079-127US01 / US2310919associated with the first subset and the coating composition formulations associated with the second subset, training system 130 may add a first child node and a second child node to the decision tree model. The first child node is associated with the coating composition formulations associated with the first subset and the second child node is associated with the coating composition formulations associated with the second subset.
[0058] After training prediction model 132, prediction system 132 may apply prediction model 132 to predict whether a candidate coating composition formulation has values of the one or more properties associated with the selected cluster. The candidate coating composition formulation may have one or more values of the one or more formulation variables different from the values of the formulation variables of the entries for the coating composition formulations in the input dataset.
[0059] FIG. 3 is a flowchart illustrating an example operation for predicting whether a candidate coating composition formulation has values of one or more properties associated with the selected cluster, in accordance with one or more techniques of this disclosure. In the example of FIG. 3, prediction system 132 may obtain data describing a candidate coating composition formulation (300). The data describing the candidate coating composition formulation may comprise values of one or more of the formulation variables. For example, the data describing the candidate coating composition formulation may include values of one or more of the following types of variables:• Cured coating mechanical properties, e g. scrubs, Martindale, washability• Other wet and dry paint characteristics (gloss, pH, contrast ratio, color)• Analytical paint characteristics (e.g. DMA, tensile, surface tension)• Calculated Paint Characteristics - e.g., PVC / N-vinylpyrrolidone (NW), VOC• Raw material properties (solubility, particle size)• Empirically determined raw material properties (data sheets / analytical / etc.)• Raw material composition (monomer composition of a polymer, thickener NW)• Quantitative Structure / Analytical (QSAR) Parameters• QSAR for surfactant hydrophobes• QSAR for thickener hydrophobes• Hansen solubility parameters for polymers1079-127US01 / US2310920• Hansen solubility parameters for coalescents• Hansen solubility parameters for surfactants hydrophobes• Hansen solubility parameters for thickener hydrophobes• QS AR for monomers• Hansen solubility parameters for monomers
[0060] The candidate coating composition formulation may have one or more values of one or more formulation variables different from the values of the formulation variables of the entries for the coating composition formulations in the input dataset. In some examples, prediction system 132 may output, for display, a graphical user interface that includes user input features that allow the user to input the candidate coating composition formulation. Examples of such graphical user interfaces are provided with greater detail elsewhere in this disclosure.
[0061] Additionally, prediction system 132 may determine a selected cluster of the plurality of clusters of clustering model 136 (302). For example, prediction system 132 may receive an indication of user input to specify the selected cluster. In some examples, prediction system 132 may receive an indication of user input specifying values of one or more properties. Prediction system 132 may determine the selected cluster as the cluster that is closest to the specified values of the one or more properties. Thus, prediction system 132 may select, based on the coordinates of the centroid points and predetermined values of the one or more target properties, the selected cluster from among the plurality of clusters of clustering model 136. Prediction system 132 may receive indications of the predetermined values from a user. In some examples, prediction system 132 may output, for display by display device 122, a user interface that allows a user to provide user input to select a cluster or specify desired values of properties.
[0062] Prediction system 132 may apply prediction model 138 to predict whether the candidate coating composition formulation has values of the one or more properties associated with the selected cluster (304). For instance, in an example where prediction model 138 comprises a decision tree model, prediction system 132 may traverse a path through the decision tree model from a root node to a leaf node by evaluating the values of the formulation variables of the candidate coating composition formulation with respect to conditions of nodes on the path. The leaf nodes are associated with the clusters identified by clustering model 136, and hence different sets of values of properties. In the example of FIG. 3, prediction system 132 may output1079-127US01 / US2310921information based on the values of properties (308). For instance, prediction system 132 may output values of the properties of the cluster associated with the leaf node of the traversed path.
[0063] In some examples, prediction system 132 may use prediction model 138 to perform backwards prediction. That is, prediction system 132 may use the prediction model 138 to determine a coating composition formulation based on a selected cluster. In an example where prediction model 138 is a decision tree model, prediction system 132 may perform backward prediction by identifying a leaf node of the decision tree model that is associated with values of properties of a coating composition (i.e., properties associated with a selected cluster). Prediction system 132 may then trace a path back through the decision tree model from the identified leaf node to the root node. Prediction system 132 may determine, based on the conditions of the nodes along the path, which values of formulation variables could be used to lead down the path from the root node to the identified leaf node. In this way, prediction system 132 may determine potential values of formulation variables that could lead to desired values of the properties.
[0064] Furthermore, by performing this backward prediction for each of the clusters, prediction system 132 may determine which of the formulation variables (and / or which values of the formulation variables) are most significant for determining which clusters a coating composition formulation will belong to.
[0065] Based on prediction model 138 predicting that the candidate coating composition formulation has the properties associated with the selected cluster, prediction system 132 may configure manufacturing system 104 to manufacture a coating composition according to the candidate coating composition formulation (306). For example, prediction system 132 may generate and output a control script (e.g., a recipe) that configures machines of manufacturing system 104 to manufacture the paint according to the candidate coating composition formulation. Thus, based on the prediction model 138 predicting that the candidate coating composition formulation has the properties associated with the selected cluster, manufacturing system 104 may manufacture a coating composition according to the candidate coating composition formulation (308).
[0066] FIG. 4 is a conceptual diagram illustrating an example user interface 400 in accordance with one or more techniques of this disclosure. User interface 400 may be an example of user interface 140 as discussed above. In the example of FIG. 4, user interface 400 may include an input1079-127US01 / US2310922feature 402 (e.g., a drop box) for selecting a cluster. Additionally, user interface 400 may include an area 404 for displaying values of properties of the selected cluster.
[0067] Furthermore, user interface 400 may include features 406A, 406B, 406C (collectively, “features 406”) for input and / or display of values of condition-basis variables. In a real user interface, user interface 400 may indicate the names of the condition-basis variables next to features 406. In some examples, features 406 may be prepopulated with values that would lead prediction system 132 to predict that a coating composition formulation is in the selected cluster.
[0068] Additionally, user interface 400 may include features 408A, 408B (collectively, “features 408” for input and / or display of values of other formulation parameters. In a real user interface, user interface 400 may indicate the names of the condition-basis variables next to features 406. In some examples, the values in features 406 may be fixed while the values in features 408 may be changed, or vice versa. As discussed above, separating or otherwise distinguishing condition-basis variables from other formulation variables may increase ease of use.
[0069] In other examples, more or fewer of features 406 and 408 may be included in user interface 400. In some examples, the number of features 406 is not preprogrammed but rather is dependent on how many condition-basis variables are identified when training prediction model 138.
[0070] User interface 400 includes a button 410. Prediction system 132 may apply prediction model 138 to predict whether a candidate coating composition formulation defined by values in features 406, 408 has values of the one or more properties associated with the selected cluster. User interface 400 may output an indication of whether the candidate coating composition formation has values of the one or more properties associated with the selected cluster. In some examples, if prediction system 132 determines that the candidate coating composition formulation does not have values of the one or more properties associated with the selected cluster, prediction system 132 may determine which of the condition-basis variables has a value in the candidate coating composition formulation that caused the candidate coating composition formulation to fall off a path in a decision tree model of prediction model 138 to a node associated with the selected cluster. User interface 400 may indicate the determined condition-basis variable. For instance, user interface 400 may highlight one or more of features 406 associated with the determined conditionbasis variable. This may increase the ease of use of user interface 400.
[0071] Furthermore, user interface 400 may include a button 412. After values have been entered in features 406 and 408, the user may select button 412. Based on receiving an indication of user1079-127US01 / US2310923input to select button 412, prediction system 132 may configure manufacturing system 104 to prepare a coating composition according to the coating composition formulation defined by the values of the formulation parameters entered in features 406 and 408. In some examples, button 412 may only be enabled after the coating composition formulation is still predicted to be in the selected cluster. In some examples, user interface 400 may display a warning if the coating composition formulation has not been checked and confirmed to still predicted to be in the selected cluster.
[0072] In some examples, prediction system 132 may receive an indication of user input via input feature 402 of user interface 400 to select a cluster. Area 404 may then show values of the properties of the selected cluster. Additionally, features 406 and 408 may show example values of formulation variables that may result in a coating composition that is in the selected cluster. Prediction system 132 may use backward prediction on prediction model 138 to determine the example values of the formulation variables. Prediction system 132 may subsequently receive indications of user input to change one or more of the example values of the formulation variables and to check whether the changes values of the formulation variables still lead to a coating composition in the selected cluster.
[0073] For processes, apparatuses, and other examples or illustrations described herein, including in any flowcharts or flow diagrams, certain operations, acts, steps, or events included in any of the techniques described herein can be performed in a different sequence, may be added, merged, or left out altogether (e.g., not all described acts or events are necessary for the practice of the techniques). Moreover, in certain examples, operations, acts, steps, or events may be performed concurrently, e.g., through multi -threaded processing, interrupt processing, or multiple processors, rather than sequentially. Further certain operations, acts, steps, or events may be performed automatically even if not specifically identified as being performed automatically. Also, certain operations, acts, steps, or events described as being performed automatically may be alternatively not performed automatically, but rather, such operations, acts, steps, or events may be, in some examples, performed in response to input or another event.
[0074] Further, certain operations, techniques, features, and / or functions may be described herein as being performed by specific components, devices, and / or modules. In other examples, such operations, techniques, features, and / or functions may be performed by different components, devices, or modules. Accordingly, some operations, techniques, features, and / or functions that may1079-127US01 / US2310924be described herein as being attributed to one or more components, devices, or modules may, in other examples, be attributed to other components, devices, and / or modules, even if not specifically described herein in such a manner.
[0075] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer-readable media generally may correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0076] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. It should be understood, however, that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but are instead directed to non-transitory, tangible storage media. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.1079-127US01 / US2310925
[0077] Instructions may be executed by one or more processors, such as one or more DSPs, general purpose microprocessors, ASICs, microcontrollers, FPGAs, or other equivalent integrated or discrete logic circuitry, as well as any combination of such components. Accordingly, the term “processor,” as used herein, may refer to any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules. Also, the techniques could be fully implemented in one or more circuits or logic elements.
[0078] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless communication device or wireless handset, a microprocessor, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, various units may be combined in a hardware unit or provided by a collection of interoperative hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware. Various examples have been described. These and other examples are within the scope of the following claims. It is, of course, not possible to describe every conceivable combination of components or methodologies for purposes of describing the present specification, but one of ordinary skill in the art may recognize that many further combinations and permutations of the present specification are possible. Each of the systems, components, and / or methodologies described above may be combined or added together in any permutation. Accordingly, the present specification is intended to embrace all such alterations, modifications and variations that fall within the spirit and scope of the appended claims. Furthermore, to the extent that the term “includes” is used in either the detailed description or the claims, such term is intended to be inclusive in a manner similar to the term “comprising” as “comprising” is interpreted when employed as a transitional word in a claim.
[0079] Illustrative examples have been described, hereinabove and below in a List of Exemplary Embodiments. It will be apparent to those skilled in the art that the above methods, systems, compositions, and methods of manufacturing may incorporate changes and modifications without departing from the general scope of this disclosure. The general scope of the disclosure is intended to encompass all such modifications and alterations.1079-127US01 / US2310926LIST OF EXEMPLARY EMBODIMENTS
[0080] Embodiment 1. A computer-implemented method comprising: obtaining, by one or more processors of a computing system, an input dataset comprising entries for a plurality of coating composition formulations, wherein for each coating composition formulation of the plurality of coating composition formulations, the entry for the coating composition formulation includes values of one or more formulation variables of the coating composition formulation and measured values of one or more properties of the coating composition formulation; training, by the one or more processors, a clustering model to partition the coating composition formulations of the input dataset into a plurality of clusters based on measured values of one or more target properties, wherein the one or more properties of the coating composition formulations include the one or more target properties; and training, by the one or more processors, a prediction model to predict, based on values of one or more condition-basis variables, whether a coating composition formulation is a member of a selected cluster of the plurality of clusters, wherein each of the one or more condition-basis variables is one of the formulation variables of the coating composition formulations.
[0081] Embodiment 2. The computer-implemented method of Embodiment 1, wherein the prediction model comprises a decision tree model.
[0082] Embodiment 3. The computer-implemented method of any one of Embodiments 1-2, wherein the coating composition formula is a first coating composition formula and the selected cluster is a first selected cluster, and the method further comprises applying, by the one or more processors, the prediction model to predict a second coating composition formula based on a second selected cluster.
[0083] Embodiment 4. The computer-implemented method of any one of Embodiments 1-3, further comprising applying, by the one or more processors, the prediction model to predict whether a candidate coating composition formulation has values of the one or more properties associated with the selected cluster, the candidate coating composition formulation having one or more values of the one or more formulation variables different from the values of the formulation variables of the entries for the coating composition formulations in the input dataset.1079-127US01 / US2310927
[0084] Embodiment 5. The computer-implemented method of Embodiment 4, further comprising, based on the prediction model predicting that the candidate coating composition formulation has the values of the one or more properties associated with the selected cluster, configuring, by the one or more processors, a manufacturing system to prepare a coating composition according to the candidate coating composition formulation.
[0085] Embodiment 6. The computer-implemented method of any one of Embodiments 4 or 5, further comprising, based on the prediction model predicting that the candidate coating composition formulation has the values of the one or more properties associated with the selected cluster, manufacturing, by a manufacturing system, a coating composition according to the candidate coating composition formulation.
[0086] Embodiment 7. The computer-implemented method of any one of Embodiments 1-6, further comprising, prior to training the clustering model, removing, by the one or more processors, anomalous coating composition formulations from the input dataset.
[0087] Embodiment 8. The computer-implemented method of any one of Embodiments 1-7, wherein training the clustering model comprises: representing the plurality of coating composition formulations as datapoints having coordinates in a coordinate space defined by the measured values of the one or more target properties; initializing two or more centroid points having coordinates in the coordinate space; and iteratively updating the coordinates of the centroid points based on distances between the coordinates of the centroid points and the coordinates of the datapoints.
[0088] Embodiment 9. The computer-implemented method of Embodiment 8, further comprising selecting, by the one or more processors, based on the coordinates of the centroid points and predetermined values of the one or more target properties, the selected cluster from among the plurality of clusters.
[0089] Embodiment 10. The computer-implemented method of Embodiment 9, further comprising receiving, by the one or more processors, indications of the predetermined values from a user.
[0090] Embodiment 11. The computer-implemented method of any one of Embodiments 8-10, further comprising outputting, by the one or more processors, for presentation by an output device, information based on values of the one or more properties that correspond to the coordinates of the centroid points.1079-127US01 / US2310928
[0091] Embodiment 12. The computer-implemented method of Embodiment 11, further comprising, after outputting the information, receiving, by the one or more processors, an indication of user input to specify the selected cluster.
[0092] Embodiment 13. The computer-implemented method of any one of Embodiments 1-12, wherein the prediction model comprises a decision tree model and training the prediction model comprises: in a first round of a decision tree training process, adding a root node to the decision tree model, wherein the root node is associated with each coating composition formulation in the plurality of coating composition formulations; performing one or more subsequent rounds of the decision tree training process after performing the first round of the decision tree training process, wherein for each subsequent round of the one or more subsequent rounds, performing the subsequent round comprises, for at least one node added to the decision tree model in a previous round of the decision tree training process: determining whether there are threshold levels of the one or more condition-basis variables for dividing coating composition formulations associated with the node into one or more coating composition formulations associated with a first subset and one or more coating composition formulations associated with a second subset, wherein the coating composition formulations associated with the first subset have measured values of the one or more target properties that correspond to the selected cluster, and the coating composition formulations associated with the second subset have measured properties of the one or more target properties that are excluded from the selected cluster; and based on there being threshold levels of the one or more condition-basis variables for dividing the coating composition formulations associated with the node into the coating composition formulations associated with the first subset and the coating composition formulations associated with the second subset, adding a first child node and a second child node to the decision tree model, wherein the first child node is associated with the coating composition formulations associated with the first subset and the second child node is associated with the coating composition formulations associated with the second subset.
[0093] Embodiment 14. The computer-implemented method of any one of Embodiments 1-13, wherein the one or more target properties include one or more properties related to washability of coating compositions produced using the coating composition formulations.
[0094] Embodiment 15. A system comprising: one or more data storage devices configured to store an input dataset comprising entries for a plurality of coating composition formulations, wherein for each coating composition formulation of the plurality of coating composition1079-127US01 / US2310929formulations, the entry for the coating composition formulation includes values of one or more formulation variables of the coating composition formulation and measured values of one or more properties of the coating composition formulation; and one or more processors implemented in circuitry and communicatively coupled to the one or more data storage devices, the one or more processors configured to: train a clustering model to partition the coating composition formulations of the input dataset into a plurality of clusters based on measured values of one or more target properties, wherein the one or more properties of the coating composition formulations include the one or more target properties; and train a prediction model to predict, based on values of one or more condition-basis variables, whether a coating composition formulation is a member of a selected cluster of the plurality of clusters, wherein each of the one or more condition-basis variables is one of the formulation variables of the coating composition formulations.
[0095] Embodiment 16. The system of Embodiment 15, wherein the prediction model comprises a decision tree model.
[0096] Embodiment 17. The system of any one of Embodiments 15-16, wherein the coating composition formula is a first coating composition formula and the selected cluster is a first selected cluster, and the one or more processors are further configured to apply the prediction model to predict a second coating composition formula based on a second selected cluster.
[0097] Embodiment 18. The system of any one of Embodiments 15-17, wherein the one or more processors are further configured to apply the prediction model to predict whether a candidate coating composition formulation has values of the one or more properties associated with the selected cluster, the candidate coating composition formulation having one or more values of the one or more formulation variables different from the values of the formulation variables of the entries for the coating composition formulations in the input dataset.
[0098] Embodiment 19. The system of Embodiment 18, wherein the one or more processors are configured to, based on the prediction model predicting that the candidate coating composition formulation has the values of the one or more properties associated with the selected cluster, configure a manufacturing system to prepare a coating composition according to the candidate coating composition formulation.
[0099] Embodiment 20. The system of any one of Embodiments 18-19, further comprising a manufacturing system configured to, based on the prediction model predicting that the candidate coating composition formulation has the values of the one or more properties associated with the1079-127US01 / US2310930selected cluster, manufacture a coating composition according to the candidate coating composition formulation.
[0100] Embodiment 21. The system of any one of Embodiments 15-20, wherein the one or more processors are configured to, prior to training the clustering model, remove anomalous coating composition formulations from the input dataset.
[0101] Embodiment 22. The system of any one of Embodiments 15-21, wherein the one or more processors are configured to, as part of training the clustering model: represent the plurality of coating composition formulations as datapoints having coordinates in a coordinate space defined by the measured values of the one or more target properties; initialize two or more centroid points having coordinates in the coordinate space; and iteratively update the coordinates of the centroid points based on distances between the coordinates of the centroid points and the coordinates of the datapoints.
[0102] Embodiment 23. The system of Embodiment 22, wherein the one or more processors are further configured to select, based on the coordinates of the centroid points and predetermined values of the one or more target properties, the selected cluster from among the plurality of clusters.
[0103] Embodiment 24. The system of any one of Embodiments 15-23, wherein the prediction model comprises a decision tree model and the one or more processors are configured to, as part of training the prediction model: in a first round of a decision tree training process, add a root node to the decision tree model, wherein the root node is associated with each coating composition formulation in the plurality of coating composition formulations; and perform one or more subsequent rounds of the decision tree training process after performing the first round of the decision tree training process, wherein for each subsequent round of the one or more subsequent rounds, the one or more processors are configured to, as part of performing the subsequent round, for at least one node added to the decision tree model in a previous round of the decision tree training process: determine whether there are threshold levels of the one or more condition-basis variables for dividing coating composition formulations associated with the node into one or more coating composition formulations associated with a first subset and one or more coating composition formulations associated with a second subset, wherein the coating composition formulations associated with the first subset have measured values of the one or more target properties that correspond to the selected cluster, and the coating composition formulations associated with the second subset have measured properties of the one or more target properties1079-127US01 / US2310931that are excluded from the selected cluster; and based on there being threshold levels of the one or more condition-basis variables for dividing the coating composition formulations associated with the node into the coating composition formulations associated with the first subset and the coating composition formulations associated with the second subset, add a first child node and a second child node to the decision tree model, wherein the first child node is associated with the coating composition formulations associated with the first subset and the second child node is associated with the coating composition formulations associated with the second subset.
[0104] Embodiment 25. A coating composition wherein a coating composition formulation for the coating composition is defined by the computer-implemented method of any one of Embodiments 1-14.
[0105] Embodiment 26. A process of manufacturing a coating composition wherein the coating composition is prepared according to a coating composition formulation defined by the computer-implemented method of any one of Embodiments 1-14.
Claims
1079-127US01 / US2310932WHAT IS CLAIMED IS:
1. A computer-implemented method comprising:obtaining, by one or more processors of a computing system, an input dataset comprising entries for a plurality of coating composition formulations, wherein for each coating composition formulation of the plurality of coating composition formulations, the entry for the coating composition formulation includes values of one or more formulation variables of the coating composition formulation and measured values of one or more properties of the coating composition formulation;training, by the one or more processors, a clustering model to partition the coating composition formulations of the input dataset into a plurality of clusters based on measured values of one or more target properties, wherein the clustering model is a first machine learning (ML) model and the one or more properties of the coating composition formulations include the one or more target properties; andtraining, by the one or more processors, a prediction model to predict, based on values of one or more condition-basis variables, whether a coating composition formulation is a member of a selected cluster of the plurality of clusters, wherein the prediction model is a second ML model and each of the one or more condition-basis variables is one of the formulation variables of the coating composition formulations.
2. The computer-implemented method of claim 1, wherein the prediction model comprises a decision tree model.
3. The computer-implemented method of any one of claims 1-2, wherein the coating composition formula is a first coating composition formula and the selected cluster is a first selected cluster, and the method further comprises applying, by the one or more processors, the prediction model to predict a second coating composition formula based on a second selected cluster.
4. The computer-implemented method of any of claims 1-3, further comprising applying, by the one or more processors, the prediction model to predict whether a candidate coating1079-127US01 / US2310933composition formulation has values of the one or more properties associated with the selected cluster, the candidate coating composition formulation having one or more values of the one or more formulation variables different from the values of the formulation variables of the entries for the coating composition formulations in the input dataset.
5. The computer-implemented method of claim 4, further comprising, based on the prediction model predicting that the candidate coating composition formulation has the values of the one or more properties associated with the selected cluster, configuring, by the one or more processors, a manufacturing system to prepare a coating composition according to the candidate coating composition formulation.
6. The computer-implemented method of any one of claims 4-5, further comprising, based on the prediction model predicting that the candidate coating composition formulation has the values of the one or more properties associated with the selected cluster, manufacturing, by a manufacturing system, a coating composition according to the candidate coating composition formulation.
7. The computer-implemented method of any one of claims 1-6, further comprising, prior to training the clustering model, removing, by the one or more processors, anomalous coating composition formulations from the input dataset.
8. The computer-implemented method of any one of claims 1-7, wherein training the clustering model comprises:representing the plurality of coating composition formulations as datapoints having coordinates in a coordinate space defined by the measured values of the one or more target properties;initializing two or more centroid points having coordinates in the coordinate space; and iteratively updating the coordinates of the centroid points based on distances between the coordinates of the centroid points and the coordinates of the datapoints.1079-127US01 / US23109349. The computer-implemented method of claim 8, further comprising selecting, by the one or more processors, based on the coordinates of the centroid points and predetermined values of the one or more target properties, the selected cluster from among the plurality of clusters.
10. The computer-implemented method of claim 9, further comprising receiving, by the one or more processors, indications of the predetermined values from a user.
11. The computer-implemented method of any of claims 8-10, further comprising outputting, by the one or more processors, for presentation by an output device, information based on values of the one or more properties that correspond to the coordinates of the centroid points.
12. The computer-implemented method of claim 11, further comprising, after outputting the information, receiving, by the one or more processors, an indication of user input to specify the selected cluster.
13. The computer-implemented method of any one of claims 1-12, wherein the prediction model comprises a decision tree model and training the prediction model comprises:in a first round of a decision tree training process, adding a root node to the decision tree model, wherein the root node is associated with each coating composition formulation in the plurality of coating composition formulations;performing one or more subsequent rounds of the decision tree training process after performing the first round of the decision tree training process, wherein for each subsequent round of the one or more subsequent rounds, performing the subsequent round comprises, for at least one node added to the decision tree model in a previous round of the decision tree training process:determining whether there are threshold levels of the one or more condition-basis variables for dividing coating composition formulations associated with the node into one or more coating composition formulations associated with a first subset and one or more coating composition formulations associated with a second subset, wherein the coating composition formulations associated with the first subset have measured values of the one or more target properties that correspond to the selected cluster, and the coating1079-127US01 / US2310935composition formulations associated with the second subset have measured properties of the one or more target properties that are excluded from the selected cluster; and based on there being threshold levels of the one or more condition-basis variables for dividing the coating composition formulations associated with the node into the coating composition formulations associated with the first subset and the coating composition formulations associated with the second subset, adding a first child node and a second child node to the decision tree model, wherein the first child node is associated with the coating composition formulations associated with the first subset and the second child node is associated with the coating composition formulations associated with the second subset.
14. The computer-implemented method of any one of claims 1-13, wherein the one or more target properties include one or more properties related to washability of coating compositions produced using the coating composition formulations.
15. A system comprising:one or more data storage devices configured to store an input dataset comprising entries for a plurality of coating composition formulations, wherein for each coating composition formulation of the plurality of coating composition formulations, the entry for the coating composition formulation includes values of one or more formulation variables of the coating composition formulation and measured values of one or more properties of the coating composition formulation; andone or more processors implemented in circuitry and communicatively coupled to the one or more data storage devices, the one or more processors configured to:train a clustering model to partition the coating composition formulations of the input dataset into a plurality of clusters based on measured values of one or more target properties, wherein the clustering model is a first machine learning (ML) model and the one or more properties of the coating composition formulations include the one or more target properties; andtrain a prediction model to predict, based on values of one or more condition-basis variables, whether a coating composition formulation is a member of a selected cluster of1079-127US01 / US2310936the plurality of clusters, wherein the prediction model is a second ML model and each of the one or more condition-basis variables is one of the formulation variables of the coating composition formulations.
16. The system of claim 15, wherein the prediction model comprises a decision tree model.
17. The system of any one of claims 15-16, wherein the coating composition formula is a first coating composition formula and the selected cluster is a first selected cluster, and the one or more processors are further configured to apply the prediction model to predict a second coating composition formula based on a second selected cluster.
18. The system of any one of claims 15-17, wherein the one or more processors are further configured to apply the prediction model to predict whether a candidate coating composition formulation has values of the one or more properties associated with the selected cluster, the candidate coating composition formulation having one or more values of the one or more formulation variables different from the values of the formulation variables of the entries for the coating composition formulations in the input dataset.
19. The system of claim 18, wherein the one or more processors are configured to, based on the prediction model predicting that the candidate coating composition formulation has the values of the one or more properties associated with the selected cluster, configure a manufacturing system to prepare a coating composition according to the candidate coating composition formulation.
20. The system of any one of claims 18-19, further comprising a manufacturing system configured to, based on the prediction model predicting that the candidate coating composition formulation has the values of the one or more properties associated with the selected cluster, manufacture a coating composition according to the candidate coating composition formulation.1079-127US01 / US231093721. The system of any one of claims 15-20, wherein the one or more processors are configured to, prior to training the clustering model, remove anomalous coating composition formulations from the input dataset.
22. The system of any one of claims 15-21, wherein the one or more processors are configured to, as part of training the clustering model:represent the plurality of coating composition formulations as datapoints having coordinates in a coordinate space defined by the measured values of the one or more target properties;initialize two or more centroid points having coordinates in the coordinate space; and iteratively update the coordinates of the centroid points based on distances between the coordinates of the centroid points and the coordinates of the datapoints.
23. The system of claim 22, wherein the one or more processors are further configured to select, based on the coordinates of the centroid points and predetermined values of the one or more target properties, the selected cluster from among the plurality of clusters.
24. The system of any one of claims 15-23, wherein the prediction model comprises a decision tree model and the one or more processors are configured to, as part of training the prediction model:in a first round of a decision tree training process, add a root node to the decision tree model, wherein the root node is associated with each coating composition formulation in the plurality of coating composition formulations; andperform one or more subsequent rounds of the decision tree training process after performing the first round of the decision tree training process, wherein for each subsequent round of the one or more subsequent rounds, the one or more processors are configured to, as part of performing the subsequent round, for at least one node added to the decision tree model in a previous round of the decision tree training process:determine whether there are threshold levels of the one or more condition-basis variables for dividing coating composition formulations associated with the node into one or more coating composition formulations associated with a first subset and one or more1079-127US01 / US2310938coating composition formulations associated with a second subset, wherein the coating composition formulations associated with the first subset have measured values of the one or more target properties that correspond to the selected cluster, and the coating composition formulations associated with the second subset have measured properties of the one or more target properties that are excluded from the selected cluster; and based on there being threshold levels of the one or more condition-basis variables for dividing the coating composition formulations associated with the node into the coating composition formulations associated with the first subset and the coating composition formulations associated with the second subset, add a first child node and a second child node to the decision tree model, wherein the first child node is associated with the coating composition formulations associated with the first subset and the second child node is associated with the coating composition formulations associated with the second subset.
25. A coating composition wherein a coating composition formulation for the coating composition is defined by the computer-implemented method of any one of claims 1-14.
26. A process of manufacturing a coating composition wherein the coating composition is prepared according to a coating composition formulation defined by the computer-implemented method of any one of claims 1-14.