Property prediction of chemical mixtures

By assigning the components of a chemical mixture formulation to predefined material clusters and utilizing a rule-based machine learning model, the problem of difficulty in predicting the properties of chemical mixtures in existing technologies is solved, achieving rapid and accurate property prediction and reducing the need for experimental verification.

CN115668237BActive Publication Date: 2026-05-19BASF COATINGS GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BASF COATINGS GMBH
Filing Date
2021-05-20
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and accurately predict the properties of chemical mixtures, especially coating formulations, requiring extensive experimental verification to ensure that the properties remain within acceptable ranges.

Method used

By assigning components in a chemical mixture formulation to predefined clusters of substances and training a data-driven model using a rule-based machine learning model, the complexity of the training dataset is reduced, the correlation between substance clusters and properties is identified, and the properties of new chemical mixtures are predicted.

Benefits of technology

It significantly reduces the need for experimental verification, improves the accuracy and efficiency of predicting the properties of chemical mixtures, and lowers development costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115668237B_ABST
    Figure CN115668237B_ABST
Patent Text Reader

Abstract

The invention relates to a property prediction of chemical mixtures. A computer-implemented method for training a data-driven model for predicting properties of chemical mixtures is proposed. The method comprises the steps of obtaining data comprising historical and / or calibration data of a plurality of chemical mixture recipes and a property of each chemical mixture recipe, each chemical mixture recipe comprising two or more ingredients; assigning at least one ingredient in each chemical mixture recipe to one of pre-defined substance clusters, each pre-defined substance cluster representing one ingredient or a group of ingredients having similar chemical properties; modifying each chemical mixture recipe by substituting the at least one ingredient with the assigned pre-defined substance cluster; and providing the modified chemical mixture recipes and the properties of the chemical mixture recipes to a machine learning process in order to train a data-driven model which can be used to predict a feature of a property of a new chemical mixture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to predicting the properties of chemical mixtures, and more particularly to computer-implemented methods and associated apparatus for training data-driven models for predicting chemical mixtures, computer-implemented methods and associated apparatus for predicting the properties of chemical mixtures, computer program products, and computer-readable media. Background Technology

[0002] Chemical mixtures (such as automotive coatings, multi-component nutritional blends, etc.) are typically formulated to achieve desired properties as expressed by property measurements. However, laboratory personnel must expend considerable effort to develop these formulations to provide the correct balance of properties.

[0003] For example, automotive coatings or coating formulations comprise complex mixtures of colorants (hues), binders, additives, and solvents formulated to provide a balance of color, appearance, durability, application, and film properties. Models can be used to quantitatively predict the color of the mixture, but not other properties. Therefore, labor-intensive validation experiments are required to measure the properties of the coating formulation to ensure that the values ​​are within acceptable limits.

[0004] Such experiments are necessary because the relationship between mixture components and the measured properties is often complex and unknown. In these cases, it would be advantageous to develop predictive models that can correlate mixture components with properties so that the properties of new mixtures can be estimated. Summary of the Invention

[0005] It may be necessary to predict the properties of chemical mixtures.

[0006] The objectives of the invention are achieved through the subject matter of the independent claims, wherein further embodiments are included in the dependent claims. It should be noted that the following described aspects of the invention also apply to computer-implemented methods and associated apparatus for training data-driven models for predicting the properties of chemical mixtures, computer-implemented methods and associated apparatus for predicting the properties of chemical mixtures, computer program products, and computer-readable media.

[0007] According to a first aspect of the present invention, a computer-implemented method is provided for training a data-driven model for predicting the properties of chemical mixtures. The method includes the following steps:

[0008] - Obtain historical and / or calibration data including multiple chemical mixture formulations and data on the characteristics of each chemical mixture formulation, wherein each chemical mixture formulation includes two or more components;

[0009] - Assign at least one component in each chemical mixture formulation to one of a predefined cluster of substances, wherein each predefined cluster of substances represents a single component or a group of components having similar chemical properties;

[0010] - The formulation of each chemical mixture is modified by replacing at least one component with a predefined cluster of substances; and

[0011] - Provide the machine learning process with modified chemical mixture formulations and their properties to train a data-driven model that can be used to predict the properties of new chemical mixtures.

[0012] In other words, empirical data can be obtained from libraries or databases (such as commercial databases or proprietary databases of companies). Empirical data includes historical and / or calibration data from historical and / or calibration experiments. Instead of directly using empirical data to train a data-driven model for predicting the properties of chemical mixtures, the proposed method modifies the empirical data by assigning at least one component in each chemical mixture formulation to a predefined cluster of chemical substances, replacing at least one component with the predefined cluster of substances assigned in each chemical mixture formulation. For example, a chemical mixture formulation may include component A, component B, and ethanol. Ethanol may be assigned to a solvent cluster named "alcohol". Therefore, a modified chemical mixture formulation including component A, component B, and the substance cluster "alcohol" is used as training data. In this way, the data-driven model cannot further distinguish components of the same substance with similar chemical properties. For example, 70 resins may cluster into 15 resin clusters. Therefore, the proposed training method does not consider 70 resins (some of which may have similar chemical properties), but only 15 resin clusters. This significantly reduces the complexity of the training dataset and thus further reduces the complexity of the data-driven model.

[0013] According to an embodiment of the present invention, the computer-implemented method further includes the step of identifying the correlation between at least one predefined material cluster and one or more properties based on training.

[0014] In other words, it is recommended to observe the percentage of substance clusters, such as resin clusters, additive clusters, or solvent clusters, within the formulation to understand the impact of substance clusters in different formulations. For example, it can be determined whether the amount of a substance cluster is positively or negatively correlated with its properties. If the amount of a substance cluster is positively correlated with its properties, increasing the amount of that substance cluster will achieve better, i.e., more desirable property characteristics. On the other hand, if the amount of a substance cluster is negatively correlated with its properties, increasing the amount of substance clusters in the formulation will achieve worse (i.e., less desirable) property characteristics.

[0015] According to embodiments of the present invention, the characteristics of each chemical mixture formulation further include, for each measured characteristic, a corresponding performance score indicating the performance evaluation of the corresponding chemical mixture formulation.

[0016] For example, performance scores can be ordinal measurements, such as a decimal categorical ordinal scale from 1 (very good, i.e., desirable) to 5 (very poor, i.e., undesirable). Including performance scores in characteristics allows for the evaluation of the performance of chemical mixture formulations.

[0017] The performance score for each chemical mixture formulation can be given based on customer feedback, customer expectations or specifications, or comparison with competitor materials.

[0018] According to embodiments of the present invention, at least one component selected from resins and / or additives is represented by a substance cluster.

[0019] According to embodiments of the present invention, the chemical mixture includes a coating formulation.

[0020] For example, a coating formulation may be an automotive coating formulation.

[0021] According to embodiments of the present invention, the characteristics of the coating formulation include the characteristics of the undried coating and / or the characteristics of the coating thus formed.

[0022] According to embodiments of the present invention, the chemical mixture includes at least one of the following: agricultural multicomponent mixtures, pharmaceutical multicomponent mixtures, nutritional multicomponent mixtures, ink multicomponent mixtures, chemical mixtures for construction purposes, and chemical mixtures used in petroleum production.

[0023] According to embodiments of the present invention, the data-driven model includes a rule-based machine learning model.

[0024] Rule-based machine learning models include any machine learning method that identifies, learns, or evolves “rules” for storage, manipulation, or application.

[0025] According to embodiments of the present invention, rule-based machine learning models include at least one of a learning classifier system, association rule learning, and an artificial immune system.

[0026] According to a second aspect of the present invention, a computer-implemented method for predicting the properties of chemical mixtures is provided. The method includes the following steps:

[0027] - To obtain a chemical mixture formulation that includes two or more ingredients;

[0028] - Assign at least one component to one of a predefined cluster of substances, wherein each predefined cluster of substances represents a single component or a group of components having similar chemical properties;

[0029] - Modify the formulation of a chemical mixture by replacing at least one component with a predefined cluster of substances;

[0030] - A data-driven model is used to process modified chemical mixture formulations to predict the characteristic measurements of the chemical mixture formulations, wherein the data-driven model has been trained according to the method according to any one of the preceding claims; and

[0031] - Outputs predicted properties of chemical mixture formulations by measurement.

[0032] In other words, data-driven models trained to predict the properties of chemical mixtures will not distinguish between components of the same cluster, because these components have similar chemical properties.

[0033] According to an embodiment of the present invention, the computer-implemented method further includes the steps of comparing the predicted characteristic measurement with the characteristic performance target and adjusting the chemical mixture formulation to meet the characteristic performance target.

[0034] According to embodiments of the present invention, the characteristics of each chemical mixture formulation further include, for each measured characteristic, a corresponding performance score indicating the performance evaluation of the corresponding chemical mixture formulation.

[0035] According to embodiments of the present invention, at least one component selected from resins and / or additives is represented by a substance cluster.

[0036] According to embodiments of the present invention, the chemical mixture includes a coating formulation.

[0037] According to embodiments of the present invention, the characteristics of the coating formulation include the characteristics of the undried coating and / or the characteristics of the coating thus formed.

[0038] According to embodiments of the present invention, the chemical mixture includes at least one of the following: agricultural multicomponent mixtures, pharmaceutical multicomponent mixtures, nutritional multicomponent mixtures, ink multicomponent mixtures, chemical mixtures for construction purposes, and chemical mixtures used in petroleum production.

[0039] According to embodiments of the present invention, the data-driven model includes a rule-based machine learning model.

[0040] According to embodiments of the present invention, rule-based machine learning models include at least one of a learning classifier system, association rule learning, and an artificial immune system.

[0041] According to a third aspect of the invention, an apparatus including a training module is provided, the training module being configured to perform the method described according to the first aspect and any associated examples.

[0042] According to a fourth aspect of the invention, an apparatus including a prediction module is provided, the prediction module being configured to perform the method described according to the second aspect and any associated examples.

[0043] According to another aspect of the present invention, a computer program product is provided, comprising a computer program having program code for performing the methods described above and below.

[0044] According to a further aspect of the invention, a computer-readable medium storing program elements is provided.

[0045] Advantageously, the benefits provided by any of the above aspects also apply to all other aspects, and vice versa.

[0046] As used herein, the term "learning" in the context of machine learning refers to identifying and training suitable algorithms to accomplish a task of interest. The term "learning" includes, but is not limited to, association learning, classification learning, clustering, and numerical prediction.

[0047] As used herein, the term "machine learning" refers to the field of computer science that studies the design of computer programs that can generalize patterns, regularities, or rules from past experience in order to respond appropriately to future data or to describe data in a meaningful way.

[0048] As used herein, the term "data-driven model" in the context of machine learning refers to a suitable algorithm that learns based on appropriate training data.

[0049] It should be understood that all combinations of the foregoing concepts and the additional concepts discussed in more detail below (provided that these concepts are not contradictory) are considered part of the inventive subject matter disclosed herein. In particular, all combinations of the claimed subject matter appearing at the end of this disclosure are considered part of the inventive subject matter disclosed herein.

[0050] These and other aspects of the invention will be apparent from and will be illustrated with respect to the embodiments described below. Attached Figure Description

[0051] In the accompanying drawings, similar reference characters generally refer to the same portion throughout different views. Furthermore, the drawings are not necessarily drawn to scale, but rather the emphasis is generally placed on illustrating the principles of the invention.

[0052] Figure 1 This is a flowchart illustrating a computer-implemented method according to some embodiments of the present invention.

[0053] Figure 2 This is a flowchart illustrating a computer-implemented method according to some embodiments of the present invention.

[0054] Figure 3 The chemical structure of melamine-formaldehyde resin is shown.

[0055] Figure 4 The central part of the chemical structure of the pyrrolopyrrole dione pigment is shown.

[0056] Figure 5 One criterion for classifying pigments into the same cluster can be the same central part of different pigments.

[0057] Figure 6 Training and prediction modules according to some embodiments of the present disclosure are shown. Detailed Implementation

[0058] According to a first aspect of this disclosure, a computer-implemented method 100 is provided for training a data-driven model for predicting the properties of chemical mixtures. The computer-implemented method includes the following steps:

[0059] - Obtain historical and / or calibration data for 110 chemical mixture formulations, including data on the characteristics of each chemical mixture formulation, wherein each chemical mixture formulation comprises two or more components;

[0060] - Assign at least one component in each chemical mixture formulation to one of the predefined material clusters in a predefined material cluster, wherein each predefined material cluster represents a single component or a group of components having similar chemical properties;

[0061] - By modifying the formulation of each of the 130 chemical mixtures by replacing at least one component with a predefined cluster of substances; and

[0062] - Provide the machine learning process with 140 modified chemical mixture formulations and their properties to train a data-driven model that can be used to predict the properties of new chemical mixtures.

[0063] Figure 1 This is a flowchart illustrating a computer-implemented method 100 according to a first aspect of this disclosure.

[0064] Empirical data collection

[0065] In step 110, empirical data may be obtained from a library or database (such as a commercial database or a company's proprietary database). Empirical data includes historical and / or calibration data from multiple chemical mixture formulations and the characteristics of each chemical mixture formulation.

[0066] Each chemical mixture formulation comprises two or more components. In some examples, a single chemical mixture formulation may include up to 50 different raw materials, or components. Two or more components are expressed as fractional concentrations of the total chemical mixture. Typically, the properties of a chemical mixture depend on the fractional concentration of its components, not the total amount of the mixture. Mixture formulations can be expressed in units of weight, volume, or other quantities. Fractional concentration is the amount of each component in the chemical mixture divided by the total amount of the mixture. The sum of fractional concentrations is one. Fractional concentration is a continuous variable ranging from 0 to 1.

[0067] The properties of a chemical mixture can be any measurable characteristic. This characteristic can be a continuous, ordinal, or nominal measurement. For example, a formulated coating can have a measurement of the viscosity of the liquid mixture on a continuous scale. For example, a measurement of orange peel wrinkles on an applied coating film can use a decimal ordinal scale from 1 (very rough) to 10 (very smooth). In another example, the performance of each chemical mixture formulation further includes, for each measured characteristic, a corresponding performance score indicating a performance evaluation of the respective chemical mixture formulation, for example, from 1 (very good) to 5 (very poor). An example of a nominal measurement could be a coded category for the acceptance or rejection of certain defects observed.

[0068] In some examples, the chemical mixture may be an automotive paint formulation. The characteristics of an automotive paint formulation may include, for example, physical properties (viscosity, sag) and appearance (concealment, gloss, image sharpness), depending on the formulation of the chemical mixture, such as the amount of paint components.

[0069] Table 1 shows exemplary characteristics and corresponding properties of water-based primers.

[0070] Table 1

[0071] Exemplary characteristics and corresponding properties of water-based primers.

[0072]

[0073]

[0074] Those skilled in the art will understand that the methods disclosed herein can also help predict the properties of other types of chemical mixtures, whether solid or liquid, including but not limited to other types of coatings and coatings, inks including inkjet inks, alcohols, diesel fuels, oils, plastics, polymer blends, films, etc.

[0075] The following will describe some further examples of chemical mixtures and the properties measured:

[0076] 1. Agricultural multi-component mixtures

[0077] For example, there are mixtures used for agricultural purposes, such as formulations for sprays used to treat crops with insecticides, fungicides, etc. Therefore, on the one hand, the sprayability of the active ingredient is guaranteed by the residual components within the formulation. That is, different other components in the formulation besides the active ingredient are used to obtain a formulation suitable for a given spraying process. In other words, sprayability (e.g., droplet size formation, ease of forming such droplets, etc.) may be a characteristic influenced by the different components of this formulation and the properties of the active ingredient.

[0078] Furthermore, the adsorption of the spray formulation on plants and the adsorption (absorption in this context) of the active ingredient or the whole spray formulation also depend on the active ingredient and residual components in the formulation. Moreover, the targeting mode of the active ingredient within the plant / organism—or more precisely, the movement of the active ingredient to the target site within the cell—is also influenced by these residual components within the formulation. That is, the rate of effect formation and the effect formation itself depend on these proportions of the formulation.

[0079] 2. Pharmaceutical multi-component mixtures

[0080] In addition to the active ingredient, the components present in the pharmaceutical formulation also affect the entire life cycle of the pharmaceutical product—from preparation to excretion or "digestion".

[0081] For example, these formulation portions define whether the active ingredient is provided as a pill, suppository, or liquid, which is primarily a dispersion of the active ingredient.

[0082] In addition, these formulation proportions define the locations in which the active ingredients are released and adsorbed in the organism.

[0083] Finally, these formulation proportions define which parts of the body's cells the active ingredient is transported to and digested there to produce the desired effect; or if it is not "digested" at all within the organism and is excreted without being "digested".

[0084] Each of these characteristics may be important in finding the right formulation, i.e., the composition of a pharmaceutical multicomponent mixture.

[0085] 3. Multi-component nutrient mixture

[0086] Many foods can be viewed as multi-component mixtures comprising different subgroups of chemicals essential for the normal functioning of our organisms. Nutritional additives (such as vitamins and minerals) are also part of food, so it is important to integrate them into the food formulation in a way that allows them to be obtained from the correct parts of the organism. Furthermore, both of these parameters can be affected by the residual proportions in the food formulation. For example, the correct way of providing mineral nutrients to an organism ensures good absorption, while a poor way of providing them can reduce absorption and lead to health problems.

[0087] 4. Inks as multi-component mixtures

[0088] Similarly, inks for coatings are also multi-component mixtures; that is, they can also be defined as ink formulations. Furthermore, the residual components, besides the color-providing components (primarily dyes in this case), ensure the ink's stability, processability, and fixation on the surface to be coated. Therefore, the different tasks and characteristics are very similar to those described in the section on coating properties, with a more detailed description of water-based primers in Table 1.

[0089] Here, particularly important characteristics include adhesion to the surface to which the ink will be applied, the sag resistance or viscosity stability of the formulation after application, and the lightfastness of the resulting print, i.e., the non-fading property of the resulting print.

[0090] 5. Chemical mixtures for construction purposes

[0091] Furthermore, many materials used in construction applications can be considered chemical mixtures. For example, concrete is formed from a mixture of cement, stones of different sizes, and water. In addition, modern concrete mixes also contain concrete additives and concrete admixtures, which are additives used to trigger and customize specific properties of the concrete mix. These properties include, for example, the construction behavior of concrete in wet or dry states, settlement behavior, hardening, tensile strength, flexural properties, and durability. All of these properties can be affected by concrete additives and concrete admixtures. While most substances used as concrete additives are inorganic, such as stone powder, fly ash, or silica fume, substances used as concrete admixtures can also have organic properties, such as acrylic acid or other oligomers or polymers.

[0092] Related applications can also involve chemical mixtures used as plastering materials. Therefore, formulations similar to concrete mixes are used. However, these plaster mortars are typically limited by the size of the aggregate. That is, the aggregate size is limited to 4mm, and larger sizes are not permitted for these mortars. The main properties (which also need to be achieved through the use of the correct additives, which are very similar to those mentioned above) are primarily in terms of workability, and correspondingly, processability. Pumpability, smoothness, and adhesion properties are typically evaluated during the development of such plastering mixes.

[0093] 6. Chemical mixtures used in petroleum production

[0094] Furthermore, chemical mixtures are used in oil production to optimize oil recovery efficiency. In hydraulic fracturing and conventional oil recovery methods, especially later in the wellbore's lifespan, efficiency levels are improved by pumping these formulations into the wellbore. Therefore, water containing organic polymers is primarily used. Overall, oil recovery efficiency is a parameter of the effectiveness of the additives used. In detail, characteristics such as the ability to release oil from the rock or the ability to generate pressure and viscosity under these conditions are likely important properties.

[0095] matter clusters

[0096] The collected empirical data is not directly used as training data. Instead, in step 120, at least one component in each chemical mixture formulation is assigned to a predefined cluster of substances. Each predefined cluster of substances can represent a group of components with similar chemical properties. However, it should be noted that a cluster of substances may also be a single component, meaning that a cluster may exist containing only one chemical substance. This can happen if the chemical properties of a component are not similar to those of one of the other components, thus preventing it from being classified into a cluster. For example, the "oligomeric or polyethylene glycol" cluster may only contain triethylene glycol, as this is the only component used outside of the oligomeric or polyethylene glycol chemical group.

[0097] For example, both monomethyl melamine-formaldehyde resin and monobutyl melamine-formaldehyde resin, which do not contain free OH groups, can be considered highly etherified melamine-formaldehyde resins. However, due to their different degrees of etherification, their physical and / or chemical behaviors will differ from each other. Therefore, two distinct clusters are generated.

[0098] For example, substances named "alcohol" can include ethanol and methanol because they have similar chemical structures and are completely soluble in and miscible with water. More cluster examples will be described later.

[0099] Components can be clustered in various ways. For example, automated clustering algorithms can be used to cluster components. For instance, experts can define multiple clusters of substances, each with representative components sufficient to represent the entire cluster. Automated clustering algorithms can compare the components of a chemical mixture formulation with the representative components of each cluster, for example, based on binary fingerprints, graph properties, or maximum common substructure, and assign components to clusters if the similarity is within acceptable limits.

[0100] In the example, a binary substructure fingerprint can be generated for a chemical structure. A substructure is a fragment of a chemical structure. A fingerprint is an ordered list of binary (1 / 0) bits. Each bit represents a Boolean determination or test for the presence of element counts, ring system types, atom pairings, atomic environments (nearest neighbors), etc., in the chemical structure.

[0101] In another example, the greatest common substructure (GCS) can be used for clustering. GCS is a graph-based similarity concept defined as the largest shared substructure (sub-graph) between two compounds, which can be used to calculate the same similarity coefficient.

[0102] Alternatively or additionally, a person (such as a formulation person) may assign clusters of substances to the ingredients of each collected chemical mixture formulation.

[0103] To facilitate understanding of the material clusters described herein, typical material clusters of automotive coating components will be described thereafter. Automotive coatings or coating formulations comprise complex mixtures of resins, pigments, solvents, and additives formulated to provide a balance of color matching, appearance, durability, application, and film properties.

[0104] First subgroup: Resins (i.e., adhesives)

[0105] Figure 3 The chemical structure of melamine-formaldehyde resin (mfr) used as a crosslinking agent in water-based primer systems is shown.

[0106] Therefore, R1, R2, R3, R4, R5, and R6 are selected from groups consisting of methyl, butyl, or hydrogen. By analyzing and comparing different raw materials, i.e., different MFRs from different suppliers, these MFRs can be grouped, for example, whether they are simply monomethylated MFRs without free OH- groups, or monobutylated MFRs without free OH- groups, etc. Based on this, such clusters can be classified. For the sake of completeness, it must be noted that amine groups, oligooxymethylene groups, and structures obtained through self-crosslinking reactions can be included in typical melamine-formaldehyde resins. All of these groups can be used to define clusters.

[0107] Further examples include polyurethane resins, which are used as "resin = binder" in waterborne primer systems. These resins are characterized by the content of so-called polyurethane bonds, achieved through the reaction of an alcohol (characterized by free OH- groups) with an isocyanate (characterized by NCO- groups). However, the chemical structures of the isocyanates and alcohols used to prepare such polyurethanes can be quite different. Therefore, key parameters of the precipitates in this synthesis process (OH- number, NCO- number, molecular weight, glass transition number, precipitate molecular ratio, etc.) define the structure and properties of the polyurethane. If these key parameters are known, a "similarity" arising from very similar key parameters can be declared, and a cluster can be categorized.

[0108] For example, three different polyurethanes are used in waterborne primer systems, which are described in WO 92 / 15405A1, page 14, lines 13-15, line 20. This polymer is then described similarly.

[0109] In this way, 70 different resins can be clustered into, for example, 15 resin clusters.

[0110] Second subgroup: Pigments

[0111] Here, a cluster can also be categorized based on the chemical structure of the pigment itself. For example, white pigments are typically based on titanium dioxide; however, each specific type differs in terms of surface treatment and / or particle size distribution. Nevertheless, these different types are ultimately all titanium dioxide, and therefore these pigments exhibit very similar chemical behaviors, making it reasonable to categorize them into a cluster based on this similarity.

[0112] The second example could be a pyrrolopyrrole dione pigment, which is made from... Figure 4 The basic chemical structure definition is shown in the figure.

[0113] Figure 4 The central portion of a pyrrolopyrrole dione pigment is shown. While the color of a particular pigment depends on the chemical structures of R1 and R2, both of which can vary and thus form different pigments, the chemical behavior of these pigments largely depends on this central portion. Therefore, classification into a particular cluster is based on this central unit. Examples of this are shown below. Figure 5 As shown in the image.

[0114] Third subgroup: Solvents

[0115] Solvents are grouped on one side by their chemical properties (e.g., whether they are alcohols or esters) and on the other side by their physical properties. For example, depending on the chemical structure of a particular alcohol, it can be considered a water-soluble or insoluble substance.

[0116] For example, ethanol and methanol are completely soluble in water and miscible with water, and therefore they can be classified into the same cluster.

[0117] Another example is the classification of alcohols that are immiscible or only poorly miscible with water, such as 2-ethylhexanol, 1-octanol, or isothietol.

[0118] Fourth subgroup: Additives

[0119] The most significant variations are used in the substances described. However, classification into a particular cluster is also accomplished here by observing the chemical structure of the substances. An example might be the polypropylene glycols used, namely Pluracol 1010, Uniol 1000, and Pluriol P900. All of these are polypropylene glycols, so the number after the brand name indicates the average molecular weight of the polypropylene glycol present in that raw material. Thus, the classification into the cluster of polypropylene glycols is completed.

[0120] Another example is the cluster of modified siloxanes, such as Byk-345, Byk-346, and Byk-347, which are all ethylene oxide-modified siloxanes. This information is used to determine "similarity" and thus identify which cluster they belong to.

[0121] In this way, more than 70 additives can be clustered into approximately 20 additive clusters.

[0122] Training data preparation

[0123] In step 130, each chemical mixture formulation is modified by replacing at least one component with an assigned predefined cluster of substances. In other words, the training dataset is constructed using modified empirical data. In the training dataset, a single example of a set of chemical mixture formulation inputs and characteristic outputs is called a paradigm. In each chemical mixture formulation input, if a component in the chemical mixture formulation is assigned to a predefined cluster of substances, the assigned cluster of substances will replace the corresponding component in the chemical mixture formulation.

[0124] For example, if a chemical mixture formulation includes components A, B, C, and a methylated melamine-formaldehyde resin without free OH- groups, the corresponding chemical mixture formulation input could be components A, B, C, and a substance group named "mfr-pure methylation". Similarly, if a chemical mixture formulation includes components B, D, and a monobutyl compound without free OH- groups, the corresponding chemical mixture formulation input could be components B, D, and a substance group named "mfr-monobutyl".

[0125] In other words, the training dataset does not distinguish between components of the same material cluster because these components have similar chemical properties. Accordingly, the complexity of the training dataset can be reduced, thereby also reducing the complexity of the training process for the data-driven model.

[0126] Data-driven model training

[0127] In step 140, the modified chemical mixture formulation and its properties are fed together to the machine learning process to train a data-driven model that can be used to predict the properties of the new chemical mixture.

[0128] In the context of machine learning, the term "data-driven model" refers to a suitable algorithm that learns based on appropriate training data. In this case, such a learned data-driven model aims to predict the properties of a chemical mixture based on the composition and clusters of substances in the corresponding chemical mixture formulation.

[0129] For example, a data-driven model might be a rule-based machine learning model. A rule-based machine learning model can include any machine learning approach that identifies, learns, or evolves “rules” to store, manipulate, or apply. The defining characteristic of a rule-based machine learning machine is that it identifies and utilizes a set of relational rules that collectively represent the knowledge captured by the system. This contrasts with other machine learning machines, which typically identify a single model that can be universally applied to any instance to make predictions.

[0130] Rule-based machine learning methods can include learning classifier systems, association rule learning, artificial immune systems, and any other methods that rely on a set of rules, each of which covers contextual knowledge.

[0131] For example, association rule learning algorithms can be used to predict one or more machine learning algorithms, which are selected from: feature evaluation algorithms, feature subset selection algorithms, Bayesian networks (see Cheng and Greiner (1999), Comparing Bayesian network classifiers. Proceedings UAI, pp. 101-107), instance-based algorithms, support vector machines (see, for example, Shevade et al., (1999), Improvements to SMO Algorithm for SVM Regression. Technical Report CD-99-16, Control Division Dept of Mechanical and Production Engineering, National University of Singapore; Smola et al., (1998). A Tutorial on Support Vector Regression. NeuroCOLT2 Technical Report Series—NC2-TR-1998-030; Scholkopf, (1998). SVMs—a practical consequence of learningtheory. IEEE Intelligent Systems. IEEE Intelligent Systems 13.4:18-21; Boser et al., (1992), A Training Algorithm for Optimal Margin Classifiers V 144-52; and Burges (1998), A tutorial on support vector machines for pattern recognition. Data Mining and Knowledge Discovery 2 (1998): 121-67), voting algorithm, cost-sensitive classifier, stacking algorithm, classification rules, and decision tree algorithm (see Witten and Frank (2005), DataMining Practical machine learning Tools and Techniques. Morgan Kaufmann, San Francisco, Second Edition).

[0132] Correlation between predefined material clusters and properties

[0133] Optionally, the computer-implemented method may include a step of identifying the correlation between at least one predefined cluster of matter and one or more properties based on training.

[0134] In the example, correlation allows for the identification of which raw materials are included in a formulation assigned a specific property value.

[0135] In another example, correlation can allow determining which combinations of raw materials occur frequently for a given characteristic value.

[0136] In a further example, this correlation can allow for the identification of which raw materials are likely to produce good and bad property values. In other words, it can be determined which raw materials are positively correlated with the property value and which are negatively correlated.

[0137] For example, the "impact resistance" property, which is very important for the quality of automotive coatings, has been found to be related to the amount and nature of the crosslinking agent in the waterborne primer. However, a higher crosslinking rate caused by a higher amount of crosslinking agent (as defined in cluster mfr pure methylation) results in a coating with poorer impact resistance, while a lower crosslinking rate caused by a lower amount of crosslinking agent (as defined in cluster mfr pure methylation) results in a coating with improved impact resistance.

[0138] predict

[0139] Once trained, the data-driven model can provide a model of the relationship between chemical mixture formulation inputs and measured property outputs. Note that in each chemical formulation input, at least one component is replaced by an assigned predefined cluster of substances. In other words, the proposed trained data-driven model does not distinguish between components of the same substance cluster. This reduces the complexity of both the input data and the data-driven model.

[0140] Therefore, according to a second aspect of this disclosure, a computer-implemented method 200 for predicting the properties of chemical mixtures is provided. The method includes the following steps:

[0141] - Obtain a chemical mixture formulation comprising two or more ingredients;

[0142] - Assign at least one component 220 to one of the predefined material clusters, each predefined material cluster representing a single component or a group of components having similar chemical properties;

[0143] - Modify the formulation of 230 chemical mixtures by replacing at least one component with a predefined cluster of substances;

[0144] - A data-driven model is used to process 240 modified chemical mixture formulations to predict the property measurements of the chemical mixture formulations, wherein the data-driven model has been trained according to the method described in accordance with the first aspect and any associated examples; and

[0145] - Output 250 chemical mixture formulations with predicted property measurements.

[0146] Figure 2 This is a flowchart illustrating a computer-implemented method 200 according to a second aspect of this disclosure.

[0147] In step 210, a chemical mixture formulation is obtained. The chemical mixture formulation includes two or more components. Examples of chemical mixtures may include, but are not limited to: coating formulations, agricultural multicomponent mixtures, pharmaceutical multicomponent mixtures, nutritional multicomponent mixtures, ink multicomponent mixtures, chemical mixtures for construction purposes, and chemical mixtures used in petroleum production.

[0148] In step 220, at least one component of the chemical mixture formulation is assigned to one of a predefined cluster of substances. Each predefined cluster represents a component or group of components with similar chemical properties. For example, ethanol and methanol are completely soluble in water and miscible with water, so they can be classified into the same cluster.

[0149] In step 230, the chemical mixture formulation is modified by replacing at least one component with an assigned predefined cluster of substances. In other words, if a component in the chemical mixture formulation is assigned to a cluster of substances, the assigned cluster of substances, rather than the component, is provided as input to the trained data-driven model.

[0150] In step 240, a data-driven model is used to process the modified chemical mixture formulation to predict the property measurements of the chemical mixture formulation. The data-driven model has been trained according to the method described in the first aspect and any associated examples. For example, if the data-driven model is a rule-based machine learning model, a set of relational rules will be derived from the training dataset. Based on these rules, the properties of the new formulation composition can be predicted with a high probability.

[0151] In step 250, a predicted property measurement of the chemical mixture formulation is provided.

[0152] Optionally, the computer-implemented method may include the steps of comparing predicted characteristic measurements with characteristic performance targets and adjusting the chemical mixture formulation to meet the characteristic performance targets.

[0153] The computer-implemented methods 100 and 200 can be implemented as a device, module, or associated component within a set of logical instructions stored in a non-transitory machine or computer-readable storage medium (such as random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), firmware, flash memory, etc.), in configurable logic (such as, for example, a programmable logic array (PLA), a field-programmable gate array (FPGA), a complex programmable logic device (CPLD)), in fixed-function hardware logic using circuitry techniques (such as, for example, application-specific integrated circuits (ASIC), complementary metal-oxide-semiconductor (CMOS), or transistor-transistor logic (TTL) technology), or any combination thereof. For example, the computer program code for performing the operations shown in methods 100 and 200 can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as JAVA, SMALLTALK, C++, Python, or similar languages, and traditional procedural programming languages ​​such as the "C" programming language or similar programming languages.

[0154] It should also be understood that, unless expressly stated to the contrary, in any method claimed herein that includes more than one step or action, the order of the steps or actions of the method is not necessarily limited to the order in which the steps or actions of the method are detailed.

[0155] Training module

[0156] According to a third aspect of the invention, an apparatus 10 is provided for training a data-driven model for predicting the properties of chemical mixtures. The apparatus includes a training module 12 configured to perform the methods described according to the first aspect and any associated examples.

[0157] Figure 6 The apparatus 10 according to a third aspect of this disclosure is shown. For example, an association rule learning model can be used as a data-driven model.

[0158] Therefore, training module 12 may refer to or include application-specific integrated circuits (ASICs), electronic circuits, processors (shared, dedicated, or grouped) and / or memories (shared, dedicated, or grouped), which execute one or more software or firmware programs, combinational logic circuits, and / or other suitable components that provide the said functionality. Furthermore, such training module 12 may be connected to volatile or non-volatile storage devices, display interfaces, communication interfaces, and similar devices known to those skilled in the art. Those skilled in the art will understand that the implementation of training module 12 depends on computational intensity and latency requirements, implied by the selection of signals used to represent positional information in a particular implementation.

[0159] Prediction module

[0160] According to a fourth aspect of the invention, an apparatus 20 is provided including a prediction module 22 configured to perform the methods described according to the second aspect of the present disclosure and any associated examples.

[0161] According to the fourth aspect of this disclosure, device 20 is also... Figure 3 As shown in the image.

[0162] Therefore, prediction module 22 may refer to or include application-specific integrated circuits (ASICs), electronic circuits, processors (shared, dedicated, or grouped) and / or memories (shared, dedicated, or grouped), which execute one or more software or firmware programs, combinational logic circuits, and / or other suitable components that provide the said functionality. Furthermore, such prediction module 22 may be connected to volatile or non-volatile memory devices, display interfaces, communication interfaces, and similar devices known to those skilled in the art. Those skilled in the art will understand that the implementation of prediction module 22 depends on computational intensity and latency requirements, implied by the selection of signals used to represent positional information in a particular implementation.

[0163] All definitions defined and used herein should be understood as controlling dictionary definitions, definitions referenced in included files, and / or the general meaning of the definition terms.

[0164] The indefinite articles “a” and “an” used herein in this specification and claims shall be understood as “at least one” unless expressly stated to the contrary.

[0165] The phrase “and / or” as used herein in this specification and claims shall be understood as “one or both” of the elements so combined, that is, elements that appear together in some cases and separately in others. Multiple elements listed using “and / or” shall be interpreted in the same manner, that is, “one or more” elements so combined. In addition to the elements explicitly identified by the “and / or” clause, other elements may optionally appear, whether related to or unrelated to those explicitly identified elements.

[0166] As used herein in this specification and claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when items are separated in a list, “or” or “and / or” should be interpreted as including, i.e., including multiple elements or at least one element from a list of elements, but also including more than one element, as well as optional additional items not listed. Only terms that explicitly indicate the opposite, such as “only one of them” or “exactly one of them,” or when used in the claims as “consisting of…”, will refer to including multiple elements or exactly one element from a list of elements. Generally, the term “or” as used herein, only when preceded by an exclusive term (such as “any,” “one of them,” “only one of them,” or “exactly one of them”), should be interpreted as indicating an exclusive alternative (i.e., “one or the other, but not both”).

[0167] As used in this specification and claims, when referring to a list of one or more elements, the phrase "at least one" should be understood as at least one element selected from any one or more elements in the list, but does not necessarily include at least one element from every and all elements specifically listed in the list, nor exclude any combination of elements in the list. This definition also allows for the optional appearance of other elements, whether related to or unrelated to those specifically identified elements, in addition to those specifically identified elements in the list referred to by the phrase "at least one".

[0168] In the claims and the foregoing description, all transitional phrases such as "comprising," "including," "carrying," "having," "containing," "involving," "holding," "forming," etc., should be understood as open-ended, meaning including but not limited to. Only the transitional phrases "consisting of..." and "consisting substantially of..." should be closed or semi-closed transitional phrases, respectively.

[0169] Furthermore, in this detailed description, those skilled in the art should note that quantitative qualifying terms such as “general,” “substantially,” “most,” and other terms are generally used to refer to the majority of the subject matter being mentioned, referring to the objects, features, or qualities that constitute the majority of the subject matter. The meaning of any of these terms depends on the context in which they are used and may be explicitly modified.

[0170] In another exemplary embodiment of the present invention, a computer program or computer program element is provided, characterized by being adapted to perform method steps of a method according to one of the above embodiments on a suitable system. Therefore, the computer program element can be stored on a computer unit, which may also be part of an embodiment of the present invention. The computing unit can be adapted to perform or direct the execution of the steps of the above-described method. Furthermore, it can also be adapted to operate components of the above-described apparatus. The computing unit can be adapted to automatically operate and / or execute user commands. The computer program can be loaded into the working memory of a data processor. Therefore, the data processor can be equipped to execute the method of the present invention.

[0171] This exemplary embodiment of the invention includes two aspects: a computer program that uses the invention from the outset and a computer program that uses the invention by means of an update to transform an existing program into a program that uses the invention.

[0172] Furthermore, computer program elements may be able to provide all the necessary steps for completing the exemplary embodiments of the methods described above.

[0173] According to a further exemplary embodiment of the present invention, a computer-readable medium, such as a CD-ROM, is proposed, wherein the computer-readable medium has computer program elements stored thereon, the computer program elements being as described in the previous section.

[0174] Computer programs may be stored and / or distributed on suitable media, such as optical or solid-state storage media provided with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems.

[0175] However, computer programs can also be presented via networks like the World Wide Web and can be downloaded from such networks to the working memory of a data processor. According to a further exemplary embodiment of the invention, a medium is provided for making computer program elements available for download, these computer program elements being arranged to perform a method according to one of the foregoing embodiments of the invention.

[0176] All features can be combined to provide a synergistic effect that is not merely a simple addition of features.

[0177] While several embodiments of the invention have been described and illustrated herein, those skilled in the art will readily conceive of various other components and / or structures for performing the functions and / or obtaining the results and / or one or more advantages described herein, and each of such variations and / or modifications is considered to be within the scope of the embodiments of the invention described herein. More generally, it will be readily understood by those skilled in the art that all parameters, dimensions, materials, and configurations described herein are intended as exemplary and that actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications for which the teachings of the invention are applied. Those skilled in the art will recognize, or can determine, many equivalents to the specific embodiments of the invention described herein, simply through conventional experimentation. Therefore, it should be understood that the above embodiments are presented by way of example only and within the scope of the appended claims and their equivalents, that the embodiments of the invention may be practiced in places other than those specifically described and claimed. The embodiments of the invention disclosed herein are directed to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the invention disclosed herein, provided that such features, systems, articles, materials, kits, and / or methods are not contradictory.

Claims

1. A computer-implemented method (100) for training a data-driven model for predicting the properties of chemical mixtures, comprising: - Obtain (110) historical and / or calibration data of multiple chemical mixture formulations and data on the characteristics of each chemical mixture formulation, wherein each chemical mixture formulation comprises two or more components; - Assign (120) at least one component in each chemical mixture formulation to one of a predefined cluster of substances, wherein each predefined cluster of substances represents a group of components having similar chemical properties; - Modify the formulation of each chemical mixture by replacing at least one component with a predefined cluster of substances; and - Provide the machine learning process with (140) a modified chemical mixture formulation and the properties of the chemical mixture formulation in order to train the data-driven model for predicting the properties of new chemical mixtures.

2. The computer-implemented method according to claim 1, further comprising: - Based on the training, identify the correlation between at least one predefined material cluster and one or more features of the property.

3. A computer-implemented method for predicting the properties of chemical mixtures, comprising: - Obtain (210) a chemical mixture formulation comprising two or more ingredients; - Assign at least one component (220) to one of a predefined cluster of substances, wherein each predefined cluster of substances represents a group of components having similar chemical properties; - Modify the chemical mixture formulation (230) by replacing at least one of the components with a predefined cluster of substances; - A data-driven model is employed to process (240) the modified chemical mixture formulation to predict the characteristic measurements of the chemical mixture formulation, wherein the data-driven model has been trained using the method according to any one of claims 1 to 2; and - Output (250) the predicted property measurements of the chemical mixture formulation.

4. The computer-implemented method according to claim 3, further comprising: - Compare the predicted characteristic measurement with the characteristic performance target; as well as - Adjust the formulation of the chemical mixture to meet the stated performance objectives.

5. The computer-implemented method according to any one of claims 1 to 4, in, The aforementioned characteristics of each chemical mixture formulation further include, for each measured characteristic, a corresponding performance score indicating the performance evaluation of the respective chemical mixture formulation.

6. The computer-implemented method according to any one of claims 1 to 4, in, At least one component selected from resins and / or additives is represented by a substance cluster.

7. The computer-implemented method according to any one of claims 1 to 4, in, The chemical mixture includes coating formulations.

8. The computer-implemented method according to claim 7, in, The properties of the coating formulation include the properties of the undried coating and / or the properties of the coating formed therefrom.

9. The computer-implemented method according to any one of claims 1 to 4, in, The chemical mixture includes at least one of the following: - Agricultural multi-component mixtures; - Pharmaceutical multi-component mixtures; - A multi-component nutritional mixture; - Multi-component ink mixture; - Chemical mixtures for building applications; and - A chemical mixture used in petroleum production.

10. The computer-implemented method according to any one of claims 1 to 4, in, The data-driven model includes rule-based machine learning models.

11. The computer-implemented method according to claim 10, in, The rule-based machine learning model includes at least one of the following: - Learning classifier systems; - Association rule learning; and - Artificial immune system.

12. An apparatus including a training module configured to perform the method according to any one of claims 1, 2, and 5 to 11.

13. An apparatus including a prediction module configured to perform the method according to any one of claims 3 to 11.

14. A computer program product comprising a computer program having program code for performing the method according to any one of claims 1 to 11.

15. A computer-readable medium storing the computer program of claim 14.