Method for predicting odor of waterborne polymer compositions

A decision tree ensemble-based method accurately predicts the odor of aqueous polymer compositions and coatings by analyzing VOC concentrations, addressing the limitations of traditional human assessment methods.

JP2025528027APending Publication Date: 2025-08-26DOW GLOBAL TECHNOLOGIES LLC +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025504073
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-08-02
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing methods for predicting the odor of aqueous polymer compositions and coatings are tedious, subjective, and prone to human error, lacking accuracy and consistency, especially due to the complex interactions of volatile organic compounds (VOCs) in these systems.

Method used

A computer-assisted method using a decision tree ensemble trained with concentration data from analytical characterization of VOCs in aqueous polymer compositions or coatings, predicting odor intensity through a supervised machine learning approach.

Benefits of technology

Provides standardized, automated odor prediction with improved accuracy and consistency, reducing labor costs and exposure to harmful substances, while enabling efficient evaluation of multiple samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025528027000001_ABST
    Figure 2025528027000001_ABST
Patent Text Reader

Abstract

A method and system (400) for predicting the odor of an aqueous polymer composition, such as a polymerized coating, comprising: analytically characterizing the aqueous polymer composition with a detector (4013), thereby generating concentration data of volatile organic compounds in the aqueous polymer composition from the analytical characterization; inputting the concentration data into a decision tree ensemble configured to predict the odor intensity of the aqueous polymer composition after polymerization based on the concentration data; and outputting a predicted odor intensity of the aqueous polymer composition from the decision tree ensemble.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and system for odor prediction of waterborne polymer compositions and coatings made therefrom by machine learning, particularly suited for coating applications.

[0002] Introduction Aqueous or water-borne binder or coating compositions are becoming increasingly important as environmentally friendly alternatives to solvent-based compositions. In the architectural coatings industry, particularly for interior applications, some manufacturers and end users are also sensitive to the odor of aqueous compositions. Traditional odor evaluation of binder and coating compositions primarily relies on human sensory panels. Such odor panel tests tend to be tedious, time-consuming, and subjective, and usually require repeated testing to obtain consistent test results. In addition, prolonged exposure to odors can pose potential hazards to odor panelists. Existing computer-implemented odor prediction methods are typically developed for relatively simple systems, such as tobacco, and are not applicable to aqueous polymer compositions or coatings in the coatings industry, which have more complex chemical compositions. These aqueous polymer compositions typically contain a wide variety of volatile organic compounds (e.g., more than 30 types) at a wide range of concentrations (e.g., parts per billion to parts per million of the aqueous polymer composition), and interactions between these compounds usually have a significant impact on the odor of aqueous polymer compositions. Therefore, developing a method or system for odor prediction of aqueous polymer systems is more challenging.

[0003] It is therefore desirable to provide a method for predicting the odor of an aqueous polymer composition or a coating made therefrom. Summary of the Invention

[0004] The present invention provides a novel computer-assisted method and system that overcomes the aforementioned problems. The method involves inputting novel combinations and concentrations of volatile organic compounds (VOCs) in an aqueous polymer composition or a coating made therefrom into a specific supervised machine learning module to predict the odor intensity of the composition or coating. The machine learning module used in the present invention is a decision tree ensemble that is trained to predict the odor intensity of an aqueous polymer composition or coating using a training dataset that employs multiple training samples. The training dataset includes concentration data and corresponding actual odor intensities assessed by human panelists for aqueous polymer compositions or coatings containing VOCs. Such a method or system enables a standardized, automated approach to accurately predict odor by assessing odor intensity. Compared to odor assessment by human panelists, the method of the present invention significantly improves the consistency and accuracy of testing without the influence of panelist variability, experimental conditions, and environmental conditions. At the same time, the method can also significantly improve productivity and efficiency for quickly evaluating dozens of samples in one sequence, thus reducing labor costs and time for training and evaluation, and preventing panelists from being exposed to harmful substances through inhalation.

[0005] In a first aspect, the present invention is a method for predicting the odor of an aqueous polymer composition. The method comprises: analytically characterizing the aqueous polymer composition with a detector, thereby generating concentration data for volatile organic compounds in the aqueous polymer composition from the analytical characterization; inputting the concentration data into a decision tree ensemble configured to predict an odor intensity of the waterborne polymer composition based on the concentration data; and outputting a predicted odor intensity of the water-based polymer composition from the decision tree ensemble.

[0006] In a second aspect, the present invention is a method for predicting the odor of a coating. The method comprises: analytically characterizing the coating with a detector, thereby generating concentration data of volatile organic compounds in the coating, the coating being obtained by drying an aqueous polymer composition; inputting the concentration data into a decision tree ensemble configured to predict an odor intensity of the coating based on the concentration data; and outputting a predicted odor intensity of the coating from the decision tree ensemble.

[0007] In a third aspect, the present invention is a system for predicting the odor of an aqueous polymer composition or a coating made therefrom. The system comprises: The system includes a detector configured to analytically characterize the aqueous polymer composition of the coating, thereby generating concentration data of volatile organic compounds from the analytical characterization, and a computing device having a deployed decision tree ensemble configured to input the concentration data and output a predicted odor intensity of the aqueous polymer composition or coating. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a schematic diagram of a simplified example of a decision tree for prediction. [Figure 2] 1 is a flowchart of an odor prediction method according to an embodiment of the present invention. [Figure 3] FIG. 1 is a schematic diagram of a cloud-based server cluster according to one embodiment of the present invention. [Figure 4] 1 is a schematic block diagram of an odor prediction system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0009] If a test method is not indicated with a date along with the test method number, it refers to the test method most recent as of the priority date of this document.

[0010] Products identified by trade names refer to compositions available under those trade names as of the priority date of this document.

[0011] "And / or" means "and, or alternatively." All ranges are inclusive of the endpoints unless otherwise indicated.

[0012] "Odor" refers to the sensation perceived through the nose by the olfactory nerves.

[0013] "Odor intensity" is a measure of how strong an odor is based on initial perception. Odor intensity can be assessed based on the criteria set out in the VDA 270 standard. The VDA 270 standard (Determination of the Odor Properties of Trim Materials in Motor Vehicles) is developed by the German Association of the Automotive Industry (VDA).

[0014] "Actual odor intensity" is the odor intensity as assessed by human panelists according to the odor panel test described in the Examples section below. "Volatile organic compound" ("VOC") refers to any organic compound having a normal boiling point of less than or equal to 250 degrees Celsius (°C) at a pressure of 101.3 kilopascals (kPa).

[0015] "VOC profile" means a data set containing the identification of VOCs and their concentrations.

[0016] As used herein, an "aqueous" polymer composition refers to a composition comprising a polymer present in an aqueous medium, e.g., polymer particles dispersed in an aqueous medium. As used herein, "aqueous medium" refers to water and 0 to 30% by weight, based on the weight of the medium, of a water-miscible compound, such as, for example, an alcohol, a glycol, a glycol ether, a glycol ester, or a mixture thereof.

[0017] As used herein, "acrylic (co)polymer" refers to a homopolymer of an acrylic monomer or a copolymer containing structural units of an acrylic monomer and one or more additional monomers. "Acrylic" in the present invention includes (meth)acrylic acid, (meth)alkyl acrylate, (meth)acrylamide, (meth)acrylonitrile, and modified forms thereof, such as (meth)hydroxyalkyl acrylate. Throughout this document, the fragment "(meth)acrylic" refers to both "methacrylic" and "acrylic." For example, (meth)acrylic acid refers to both methacrylic acid and acrylic acid, and methyl (meth)acrylate refers to both methyl methacrylate and methyl acrylate. Specific examples of acrylic (co)polymers include acrylic homopolymers, styrene-acrylic copolymers, or mixtures thereof.

[0018] A "structural unit," also known as a "polymerized unit," of a specified monomer refers to the remainder of the monomer after polymerization, i.e., the polymerized monomer or the polymerized form of the monomer. For example, the structural unit of methyl methacrylate is:

[0019] [ka] where the dotted lines represent the points of attachment of the structural units to the polymer backbone.

[0020] "Machine learning" refers to a set of methods that "learn" from data to improve performance for a specific task. Machine learning algorithms build models based on historical data, also known as training data, to make predictions as the model output.

[0021] A "decision tree" is a type of machine learning model in which a model has a tree-like structure to represent "tests" against a set of attributes. Each internal node represents a "test," and each leaf node represents the decision that follows the outcome of the attribute.

[0022] The method of the present invention is useful for predicting the odor of an aqueous polymer composition or a coating made therefrom (collectively referred to as a "test sample"). In particular, the method is useful for predicting the odor of an aqueous polymer composition.

[0023] The method of the present invention involves analytically characterizing the aqueous polymer composition or coating with a detector, thereby generating concentration data for volatile organic compounds in the aqueous polymer composition or coating.

[0024] Aqueous polymer compositions, especially in coating applications, typically contain various types of VOCs, e.g., more than 20 types. These VOCs can be present in a wide range of concentrations, from parts per billion to parts per million of the composition. The VOC types and the interactions between these VOCs usually significantly affect the perceived odor of an aqueous polymer composition. Therefore, the odor intensity from a mixture of individual chemicals cannot reflect the actual odor of the aqueous polymer composition containing these individual chemicals. By using concentration data for the mixture of VOCs in an aqueous composition obtained from analytical characterization to train a specific machine learning module, the present invention can achieve surprisingly higher prediction accuracy than modules that generate predicted odor intensities by inputting individual chemical structure data (chemical molecule descriptors) for training or by measuring individual chemicals for training.

[0025] Waterborne polymer composition The aqueous polymer compositions useful in the present invention may contain one or more acrylic (co)polymers. The acrylic (co)polymers may contain any one or any combination of two or more alkyl esters of (meth)acrylic acid. "Alkyl" means a straight-chain, branched-chain, or cyclic alkyl group. Examples of alkyl esters of (meth)acrylic acid include C1 to C6 alkyl esters of (meth)acrylic acid, such as methyl acrylate, methyl methacrylate, ethyl acrylate, butyl acrylate, butyl methacrylate, 2-ethylhexyl acrylate, isobutyl (meth)acrylate, hexyl (meth)acrylate, lauryl (meth)acrylate, stearyl (meth)acrylate, cyclohexyl (meth)acrylate, benzyl (meth)acrylate, oleyl (meth)acrylate, palmityl (meth)acrylate, nonyl (meth)acrylate, decyl (meth)acrylate, dodecyl (meth)acrylate, pentadecyl (meth)acrylate, hexadecyl (meth)acrylate, octadecyl (meth)acrylate, and mixtures thereof. 20 , C1~C 10 , or C1 to C8 alkyl. Desirably, the alkyl ester of (meth)acrylic acid comprises methyl methacrylate, methacrylate, ethyl acrylate, butyl methacrylate, butyl acrylate, 2-ethylhexyl acrylate, or a mixture thereof. The acrylic (co)polymer may comprise structural units of the alkyl ester of (meth)acrylic acid in a concentration of 5 wt% to 100 wt%, based on the weight of the acrylic (co)polymer, and may be 15 wt% to 99 wt%, 30 wt% to 98 wt%, 50 wt% to 95 wt%, 55 wt% to 90 wt%, or 60 wt% to 90 wt%.

[0026] The acrylic (co)polymers useful in the present invention may or may not contain structural units of one or more ethylenically unsaturated functional monomers having at least one functional group selected from amide, ureido, carbonyl, carboxyl, carboxylic anhydride, hydroxyl, sulfonic acid, sulfonate, sulfate, or sulfate groups, or salts thereof (hereinafter "functional monomers"). Suitable functional monomers include α,β-ethylenically unsaturated carboxylic acids, including monomers having an acid such as methacrylic acid, acrylic acid, itaconic acid, maleic acid, or fumaric acid; or monomers having an acid-forming group that generates or can subsequently be converted into an acid group (such as anhydride, (meth)acrylic anhydride, or maleic anhydride); vinylphosphonic acid, allylphosphonic acid, phosphoalkyl(meth)acrylates, e.g., phosphoethyl(meth)acrylate, phosphopropyl(meth)acrylate, phosphobutyl(meth)acrylate, salts thereof, or mixtures thereof. The functional monomer may include phosphorus-containing monomers; sulfonic acid monomers and their salts, such as 2-acrylamido-2-methyl-1-propanesulfonic acid, the sodium salt of 2-acrylamido-2-methyl-1-propanesulfonic acid, and the ammonium salt of 2-acrylamido-2-methyl-1-propanesulfonic acid; sodium vinyl sulfonate, the sodium salt of allyl ether sulfonic acid; monomers having carbonyl groups, such as acetoacetoxyethyl methacrylate (AAEM) and diacetone acrylamide (DAAM); acrylamide; methacrylamide; or mixtures thereof. Desirably, the functional monomer includes acrylamide, acrylic acid, methacrylic acid, phosphoethyl (meth)acrylate, the sodium salt of 2-acrylamido-2-methyl-1-propanesulfonic acid, or mixtures thereof. The acrylic (co)polymer may contain structural units of functional monomers at a concentration of 0 to 10% by weight, based on the weight of the acrylic (co)polymer, and can be 0.1% to 8% by weight, 0.3% to 5% by weight, or 0.5% to 3% by weight, or 0.7% to 2% by weight.

[0027] The acrylic (co)polymers useful in the present invention may or may not contain structural units of one or more additional ethylenically unsaturated nonionic monomers other than alkyl esters of (meth)acrylic acid and the functional monomers described above. The term "nonionic monomer" refers to a monomer that does not have an ionic charge at pH = 1 to 14. Suitable additional ethylenically unsaturated nonionic monomers may include vinyl aromatic monomers such as styrene and substituted styrenes (including, for example, α-methylstyrene, p-methylstyrene, t-butylstyrene, and vinyltoluene), glycidyl (meth)acrylate, vinyl trialkoxysilanes such as vinyltrimethoxysilane, and ethylenically unsaturated monomers having at least one alkoxysilane functionality, including (meth)acryloxyethyltrimethoxysilane and (meth)acryloxypropyltrimethoxysilane; α-olefins such as ethylene, propylene, and 1-decene; vinyl acetate, vinyl acetate, vinyl butyrate, vinyl versatate, and other vinyl esters; acrylonitrile; or mixtures thereof. Desirably, the additional ethylenically unsaturated nonionic monomer comprises styrene, vinyl acetate, or a mixture thereof. The acrylic (co)polymer may or may not comprise structural units of polyfunctional nonionic monomers such as butadiene, divinylbenzene, and allyl (meth)acrylate, typically at a concentration of 0 to 5 wt %, based on the weight of the acrylic (co)polymer, and may be 0 to 2 wt %, 0.1 wt % to 1 wt %, or 0.1 wt % to 0.5 wt %.

[0028] The acrylic (co)polymers useful in the present invention may contain structural units of additional ethylenically unsaturated nonionic monomers at a concentration of 0 to 95% by weight, based on the weight of the acrylic (co)polymer, and may be 1 to 85%, 2 to 70%, 5 to 50%, or 10 to 45% by weight. Desirably, the acrylic (co)polymer may contain styrene structural units at a concentration of 0 to 60% by weight, based on the weight of the acrylic (co)polymer, and may be 5 to 55%, 10 to 50%, or 20 to 40% by weight. Alternatively, the acrylic (co)polymer may contain vinyl acetate structural units at a concentration of 0 to 95% by weight, based on the weight of the acrylic (co)polymer, and may be 5 to 90%, 10 to 85%, 20 to 80%, or 50 to 70% by weight.

[0029] The acrylic (co)polymer useful in the present invention can be a vinyl acrylic (co)polymer containing 5% to 30% structural units of an alkyl ester of (meth)acrylic acid, 70% to 95% structural units of vinyl acetate, and 0 to 5% structural units of a functional monomer. Alternatively, the acrylic (co)polymer can be a styrene acrylic copolymer containing 30% to 60% structural units of styrene, 40% to 70% structural units of an alkyl ester of (meth)acrylic acid, and 0 to 5% structural units of a functional monomer. Alternatively, the acrylic (co)polymer can contain 95% to 100% structural units of an alkyl ester of (meth)acrylic acid and 0 to 5% structural units of a functional monomer.

[0030] The aqueous polymer compositions useful in the present invention may contain one or more VOCs. Depending on the synthesis process, such as the polymerization process, for preparing the aqueous polymer composition, the VOCs in the aqueous polymer composition may be selected from aldehydes, ketones, acrylic monomers, alcohols, saturated esters, ethers, aromatic hydrocarbons, or mixtures thereof.

[0031] The aqueous polymer composition can be a binder emulsion or an aqueous coating composition. Commercially available aqueous polymer compositions include, for example, PRIMAL DC-420 emulsion, PRIMAL DC-430V emulsion, PRIMAL SF-155 emulsion, PRIMAL DC-460 emulsion, PRIMAL DC-480 emulsion, PRIMAL LE-328V emulsion, PRIMAL AS-2010 emulsion, PRIMAL AS-356 emulsion, PRIMAL SF-508 emulsion, PRIMAL SF-105 emulsion, PRIMAL SF-308 emulsion, PRIMAL LE-318V emulsion, or mixtures thereof (PRIMAL is a trademark of The Dow Chemical Company).

[0032] The aqueous polymer compositions useful in the present invention may or may not contain pigments, extenders, or mixtures thereof. Pigments may include particulate inorganic materials that can substantially contribute to the opacity or hiding power of the coating. Such materials typically have a refractive index greater than 1.8. Examples of suitable pigments include titanium dioxide (TiO), zinc oxide, zinc sulfide, iron oxide, barium sulfate, barium carbonate, or mixtures thereof. The aqueous polymer compositions may or may not contain one or more extenders. Extenders may include particulate inorganic materials that typically have a refractive index less than or equal to 1.8 and greater than 1.5. Examples of suitable extenders include calcium carbonate, aluminum oxide (Al2O3), clay, calcium sulfate, aluminosilicates, silicates, zeolites, mica, diatomaceous earth, solid or hollow glass, ceramic beads, and opaque polymers, such as ROPAQUE™ Ultra E opaque polymer available from The Dow Chemical Company (ROPAQUE is a trademark of The Dow Chemical Company), or mixtures thereof. The pigment and / or extender may be present in a concentration of 0 to 40 wt%, 5 to 30 wt%, 10 to 25 wt%, or 15 to 20 wt%, based on the weight of the aqueous polymer composition. The aqueous polymer composition may or may not further include any one or combination of the following additives commonly used in coating applications: antifoaming agents, thickeners, dispersants, biocides, and coalescents.

[0033] Analytical Characterization The method of the present invention includes analytically characterizing an aqueous polymer composition using a detector, thereby generating concentration data for VOCs in the aqueous polymer composition. "VOC concentration data," also referred to as "VOC concentration data," refers to a plurality of concentration values ​​for VOCs in the aqueous polymer composition. Alternatively, the concentration data can include concentration values ​​for VOCs each having a concentration of 0.1 parts per million (ppm) or greater, 0.2 ppm or greater, 0.3 ppm or greater, or even 0.5 ppm or greater by weight of the aqueous polymer composition. The concentration of VOCs refers to the weight concentration of the VOC in ppm based on the weight of the aqueous polymer composition.

[0034] Surprisingly, it has also been found that the concentrations of some specific VOCs, such as compounds with 1 to 11 carbon atoms and normal boiling points below 220°C at a pressure of 101.3 kPa, play a crucial role in the accuracy of the prediction of odor intensity for such aqueous polymer compositions. Preferably, the concentration data input into the decision tree ensemble includes concentrations of VOCs selected from acetone, 2-methylpropanol, 1-butanol, methyl methacrylate, butyl acetate, 4-heptanone, 2-heptanone, butyl ether, styrene, butyl acrylate, anisole, propanoic acid, butyl ester, methylethylbenzene, 3-methyl-4-heptanone, propenylbenzene, propylbenzene, benzaldehyde, acetophenone, butyl methacrylate, isobutylvinyl acetate, butanoic acid, butyl ester, 2-butenoic acid, butyl ester, diethylbenzene isomers, cyclohexyl methacrylate (2-propenoic acid, 2-methyl-, cyclohexyl ester), 2-ethylhexyl acrylate, xylene, ethylbenzene, or mixtures thereof. More preferably, the concentration data includes or consists of concentrations of VOCs selected from 2-methylpropanol, methyl methacrylate, butyl acetate, 2-heptanone, 3-methyl-4-heptanone, styrene, butyl acrylate, propanoic acid, butyl esters, 2-ethylhexyl acrylate, ethylbenzene, xylene, benzaldehyde, or mixtures thereof.

[0035] Analytical characterization of the aqueous polymer composition may include analytical characterization techniques with a lower limit of detection (LLOD) of less than 0.1 ppm. LLOD (also known as analytical sensitivity) is the minimum amount of an analyte that can be reliably detected. A typical analytical characterization technique may be gas chromatography (GC) coupled with different detectors, such as GC-MSD (mass selective detector) (hereinafter referred to interchangeably as "GC-MS"), GC-FID (flame ionization detector), and GC-ECD (electron capture detector). GC-MS typically includes a GC instrument and an MSD. The detector in GC-MS can be used to identify VOCs by their MS spectrum and can also be used to obtain the peak area of ​​the VOCs for further quantification of their concentration. The concentration of each VOC can be obtained by comparing the peak area of ​​the VOCs in the aqueous polymer composition with that of an external standard. Conventional GC-MS, such as an Agilent 6890 gas chromatograph coupled with an Agilent 5975C MSD, can be used. Various sample preparation techniques for GC-MS can be used, including solid phase microextraction (SPME) coupled with GC-MS (also referred to as "SPME GC-MS"), needle trap microextraction (NTM) coupled with GC-MS, and Tenax absorbent cartridge (TC) coupled with GC-MS. Conventional NTM GC-MS and TC GC-MS are available from Shinwa Ltd. (Japan) and Gerstel Co., Ltd. (Germany), respectively.

[0036] Preferably, VOCs are extracted, identified, and quantified from aqueous polymer compositions using SPME coupled with a GC-MS, including SPME coupled to a GC equipped with an MSD instrument. SPME is a sample preparation technique for integrating operations including sample collection, extraction, and analyte concentration from the sample headspace. The SPME technique of the present invention can be coupled with a GC and used to extract analytes (such as VOCs) from liquid samples, such as aqueous polymer compositions. A typical SPME procedure involves two steps: (1) partitioning the analytes between an extraction phase and a sample matrix, and (2) subsequently desorbing the concentrated extract into an analytical instrument, such as a GC-MS. The SPME procedure can be performed manually or automatically. For example, the SPME procedure can be automated using a multi-purpose sampler (MPS).

[0037] Decision Tree Ensemble The VOC concentration data generated from the analytical characterization is used as input to a decision tree ensemble configured to predict the odor intensity of the aqueous polymer composition based on the VOC concentration data, e.g., the decision tree ensemble is trained to predict odor intensity.

[0038] A decision tree is a model that uses a tree-like decision structure and its possible outcomes, such as the likelihood of an event outcome or a predicted value. A decision tree model, an example of a supervised learning model, is trained using input data paired with output data by optimizing model parameters to minimize the difference between the model predicted value and the actual value.

[0039] The trained model can then predict new outputs given a set of input data from a test sample (also "new sample"). A "decision tree ensemble" (also known as a "decision tree ensemble model") refers to a model that combines multiple decision trees. "Multiple" means two or more, and can be ten or more, but generally no more than 1000 or no more than 100.

[0040] Figure 1 shows the input X={x1,x2...x p 1 shows a schematic diagram of an example of a simplified decision tree with a decision-making process of predicting y based on p features as {. For example, in a decision tree using two features x1 and x2, each node is split into different branches based on the condition of x1 or x2. Based on the condition of x1, the decision tree is first split into three branches with different ranges of x1. If x1>0.5, the decision tree output is y=5. If x1<0.5, there is another branch that splits based on the condition of x2. In these new branches, the predicted value of y can be 2, 3, or 4 depending on different conditions of the value of x2 under the condition of x1<0.5. For example, if x1<0.5 and x2>0.6, the predicted y is 4.

[0041] Decision trees for regression can be constructed by using standard deviation reduction to create leaf nodes. In decision tree regression tasks (where the predicted y is a continuous numeric value), the decision tree is constructed top-down from the root node to the leaf nodes, and involves dividing data containing similar values ​​by the standard deviation. A decision tree has two or more branches, each branch representing the value of the attribute being tested. Each leaf node represents a decision regarding the target. The count m represents the total number of points in the node,

[0042]

number

[0043]

number

[0044] The standard deviation (S) is for the tree branches.

[0045]

number

[0046] The coefficient of variation (CV) is used to stop branching.

[0047]

number

[0048] The sum of standard deviations of multiple attributes (S(target,C)) can be calculated from the combination of the standard deviations of the individual attributes (S), and the probability (P) of condition (c) occurring under all conditions (C) of the decision.

[0049]

number

[0050] Standard deviation reduction (SDR) can be achieved after the dataset is split on features. A decision tree is constructed to find the highest SDR, which is calculated by the difference in the standard deviation of the nodes between before splitting (S(target)) and after splitting (S(target, C)). SDR(Target,C)=S(Target)-S(Target,C)

[0051] Branching is stopped after the CV reaches a specified threshold. Sometimes branching is also limited by other criteria, including a maximum tree depth and a minimum number of samples per node.

[0052] The decision tree ensemble method combines several decision trees to produce better predictive performance than using a single tree. The decision tree ensemble used in the present invention is trained to predict the odor intensity of a sample by building a correlation model from VOC concentrations to odor intensities evaluated by human sensory panelists (hereinafter also referred to as "actual odor intensity") during the training phase. During the training phase, the concentrations of VOCs in the training samples serve as inputs to the decision tree ensemble model, and the actual odor intensities of these training samples evaluated by the panelists serve to guide the learning of the output to minimize the difference between the actual and predicted values. The trained decision tree ensemble model can be used to predict the odor intensity of a new sample based on its VOC concentration.

[0053] First, the concentrations of each of the selected VOCs for a total of p features are collected: {P1, P2...P p}, where each feature set P contains n samples: P={p1, p2...p n} These are rescaled so that all values ​​are in the interval [0,1].

[0054]

number

[0055] The decision tree ensemble model takes rescaled inputs from the training data and predicts odor intensity as output. The model parameters are the actual values ​​y i and predicted value

[0056]

number

[0057]

number

[0058]

number

[0059] MSE is a metric used to measure the average of the squares of the difference between predicted and actual values. RMSE stands for root mean square error, a metric for measuring the difference between predicted and actual values. Percentage RMSE is the square root of MSE normalized by the ensemble mean for a dimensionless number expressed as a percentage. Percentage RMSE can be calculated as follows:

[0060]

number

[0061]

number

[0062]

number

[0063] R 2 R stands for discrimination coefficient, which is a commonly used performance metric for regression that calculates the proportion of variance explained by the regression model. 2 is typically in the range of 0 to 1 and can be calculated according to equation (III) below:

[0064]

number

[0065]

number

[0066]

number

[0067] For example, multiple decision tree-based ensemble learning models can be used in the present invention, including a Random Forest (RF) model, a Gradient Boosting (GB) model, and an eXtreme Gradient Boosting (XGB) model.

[0068] The RF model uses bagging, which creates several subsets of data from the training samples while randomly selecting training features to develop multiple decision trees using the selected samples and features. The decision trees are then averaged to obtain the final RF model. The RF model can effectively reduce the variance of a single decision tree.

[0069] The GB model uses boosting, another ensemble technique that combines multiple weak learners to create a strong learner. In the GB model, each individual decision tree is trained sequentially, with early learners learning a simple model for the data and later learners analyzing the prediction error by learning the error from the earlier learners. It uses a gradient descent algorithm, where a differentiable loss function allows the weights of the successive learners to be optimized to recover the difference between the actual and predicted values.

[0070] The XGB model offers several adjustments to the GB algorithm, including changing the gradient descent algorithm in the GB algorithm to a Newton-Raphson optimization algorithm. The XGB model also uses additional randomization parameters, proportional shrinkage at leaf nodes, tree penalization, and feature selection in addition to the original GB algorithm.

[0071] The RF, GB, and XGB models can be constructed in a computing environment such as Python, R, or MATLAB. These models are available in open source libraries such as the Scikit-learn library. For example, the RF and GB models can be constructed using the Scikit-learn package, as described in "Scikit-learn: Machine Learning in Python," Pedregosa, F. et al., Journal of Machine Learning Research, 2011, 12, 2825-2830. The XGB model can be constructed using the XGBoost package, as described in "XGBoost: A Scalable Tree Boosting System," Chen, T., and Guestrin, C., Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016, 785-794, Association for Computing Machinery (ACM), New York, USA.

[0072] Hyperparameters for a decision tree ensemble model may be default settings in a machine learning model library, such as the Scikit-learn library. A "hyperparameter" is a defined value used to control the learning process that is set before the model training process begins. Hyperparameters are distinct from other parameters of the model because the other parameters are learned during the training process. Examples of hyperparameters include the number of estimators, which represents how many trees are used for the ensemble learning model, and the minimum number of samples per split, which represents the criterion for making splits in a decision tree model.

[0073] A decision tree ensemble useful in the present invention is trained to predict the odor intensity of an aqueous polymer composition using a training dataset that uses multiple training samples, such as concentration data for aqueous compositions containing VOCs (i.e., a trained decision tree ensemble). The training dataset includes VOC concentration data (as input) for each training sample paired with corresponding actual odor intensity data (as output) evaluated by human panelists for such training samples, thereby providing a trained decision tree ensemble. The resulting training results can be validated with a validation dataset using validation samples to monitor model prediction accuracy. The decision tree ensemble model can be further tested with a test dataset using a separate set of test samples that are prepared separately at a different time than the samples used to generate the training dataset. Models that meet certain thresholds for training performance, validation accuracy, and test accuracy are selected.

[0074] In the present invention, the concentrations of selected VOCs in each training sample, such as the aqueous polymer composition described above or the coating described below, are input into a decision tree ensemble model to train the model. Desirably, the training data set includes concentration data for volatile organic compounds each having a concentration of ≥ 0.1 ppm by weight based on the weight of the training sample (i.e., the training aqueous polymer composition). More preferably, the training dataset includes concentration data for volatile organic compounds, including acetone, 2-methylpropanol, 1-butanol, methyl methacrylate, butyl acetate, 4-heptanone, 2-heptanone, butyl ether, styrene, butyl acrylate, anisole, propanoic acid, butyl ester, methylethylbenzene, 3-methyl-4-heptanone, propenylbenzene, propylbenzene, benzaldehyde, acetophenone, butyl methacrylate, isobutylvinyl acetate, butanoic acid, butyl ester, 2-butenoic acid, butyl ester, diethylbenzene or isomers, cyclohexyl methacrylate, 2-ethylhexyl acrylate, xylene, ethylbenzene, or mixtures thereof. The training samples are analytically characterized to determine the concentrations of VOCs using methods such as GC-MS (as input), and odor intensity from human sensory results (i.e., human panelist ratings) is also evaluated as "actual odor intensity." The actual odor intensity of the training samples can be evaluated by an odor panel test according to the VDA 270 standard on a scale of 1 to 6, where 1 = not perceptible, 2 = perceptible but not bothersome, 3 = clearly perceptible but not bothersome, 4 = bothersome, 5 = very bothersome, and 6 = unacceptable (further details are provided in the Examples section below). A decision tree ensemble model is trained to model the correlation between VOC concentration data and the odor intensity of the aqueous polymer composition as evaluated by the panelists. The decision tree ensemble model in the present invention is calculated using the training R as calculated by equation (III) above. 2The trained decision tree ensemble model can be trained to provide an accuracy of >0.85. The trained decision tree ensemble model can be further validated using a validation dataset that uses multiple validation samples. The validation dataset includes VOC concentration data (as input) for volatile organic compounds in each validation sample and corresponding actual odor intensity data (as output) for such validation samples as assessed by human panelists.

[0075] The trained decision tree ensemble model can then predict the odor of an aqueous polymer composition tested as a test sample (and as a "new sample") using a test dataset containing concentration data of VOCs in the test sample obtained from analytical characterization. A user can input the concentration data of VOCs for the test sample into the trained decision tree ensemble via a web-based user interface (described in more detail below).

[0076] In one embodiment of the present invention, a decision tree ensemble model is trained using a total of 39 data pairs containing actual odor intensities evaluated by human panelists, which are divided into a training data set with 31 data pairs and a validation data set with 8 data pairs. A "data pair" herein refers to the VOC concentration data for a particular sample and its corresponding odor intensity evaluated by the panelists. The model is then evaluated or tested using a new data set collected in a new batch of experiments to evaluate the predictive performance of the model.

[0077] 2 shows a flowchart of an odor prediction method according to one embodiment of the present invention. The method includes raw data extraction by identifying the type of VOC to be analyzed, obtaining a VOC profile including concentration data of the VOC, and collecting odor intensity data from an odor panel test (further details are provided in the Examples section below). The odor intensity data is used for model development along with the concentration data as data pairs. The method includes training and validating multiple models using the obtained concentration data paired with the odor intensity data, and then performing a training R 2 and selecting an appropriate model based on desired model criteria, including the training RMSE and validation RMSE. Underfitting is a scenario in which a model fails to capture the relationship between input and output variables. An underfitted model will have an undesirably low training R 2 Training R 2 The threshold, i.e., training R 2 >0.85 is used to filter out underfitted models. Nevertheless, model overfitting is when the model is matched too closely to the training data, and the learned representation cannot accurately predict the validation data. A validation percentage RMSE threshold, i.e., validation percentage RMSE<30%, is used to filter out overfitted models. After meeting the threshold, the trained model is used to predict the odor intensity of new samples.

[0078] The method of the present invention may further include adjusting the synthesis process of the aqueous polymer composition, for example, adjusting the polymerization process, particularly the emulsion polymerization process, for preparing the polymer in the aqueous polymer composition based on the predicted odor intensity. Parameters that can be adjusted include, for example, the type and amount of surfactant, the type and amount of initiator, the type and source of monomer, the reaction temperature, steam stripping parameters, or a combination thereof.

[0079] coating The present invention also relates to a method for predicting the odor of a coating (also referred to as a "coating film"). The method includes analytically characterizing the coating using a detector to generate concentration data for VOCs in the coating, inputting the concentration data into a decision tree ensemble configured to predict the odor intensity of the coating based on the VOC concentration data, and outputting a predicted odor intensity of the coating from the decision tree ensemble. The coating can be obtained by drying or pre-drying the aqueous polymer composition described above (i.e., the dried aqueous polymer composition). The aqueous polymer composition can be applied to a substrate and dried or pre-drying the applied polymer composition to form a coating. The coating can have a dry film thickness of 50 to 60 μm. The aqueous polymer composition can be used alone or in combination with other coatings to form multi-layer coatings. The aqueous polymer composition can be applied to a substrate by conventional means, including brushing, dipping, rolling, and spraying. Drying can be performed at room temperature (20°C to 25°C) or at an elevated temperature, e.g., 35°C to 60°C. Because some of the VOCs in the aqueous polymer composition may evaporate from the composition after drying, the resulting coating may have a different VOC profile than the aqueous polymer composition. The coating is suitable for various coating applications, such as architectural coatings, wood coatings, and protective coatings. The analytical characterization useful for generating VOC concentration data for the coating is the same as that described above for analytically characterizing the aqueous polymer composition. The decision tree ensemble used in the method for predicting the odor of a coating is the same as that described above in the method for predicting the odor of an aqueous polymer composition, except that the training sample, validation sample, and test sample used in the decision tree ensemble are coating samples obtained by drying the aqueous polymer composition.

[0080] In particular, the novel combination of VOC concentration data as input with a specific machine learning model, i.e., a decision tree ensemble, enables the method of the present invention to predict the odor intensity of aqueous polymer compositions with higher accuracy than other machine learning models, such as an artificial neural network (ANN) model, a partial least squares regression (PLS) model, or a ridge regression (RR) model. The ANN model is based on a set of connected units or nodes called artificial neurons for predictive modeling. The PLS model finds a linear regression model by projecting predictor and observable variables into a new space. The RR model builds a multiple regression model in scenarios where independent variables are highly correlated by creating a ridge regression estimator.

[0081] In one embodiment, the method includes a training R 2 The training R is greater than 0.85, and the validation percentage RMSE is less than 30% (<30%) as calculated according to equation (II) above. 2 " is an R query on the training dataset. 2 "Validation percentage RMSE" refers to the percentage RMSE for the validation data set, and "Test percentage RMSE" refers to the percentage RMSE for the test data set. In particular, when using an RF model, the method can provide even higher prediction accuracy, exhibiting a test percentage RMSE of less than 20%.

[0082] A system for odor prediction. The present invention also relates to a system for predicting the odor of an aqueous polymer composition or coating. The system of the present invention includes a detector configured to analytically characterize the aqueous polymer composition or coating, thereby generating concentration data of volatile organic compounds in the aqueous polymer composition or coating from the analytical characterization, and a computing device configured to input the concentration data and output a predicted odor intensity, and having a decision tree ensemble developed thereon. The decision tree ensemble and the training of the decision tree ensemble are as described above. For example, the decision tree ensemble developed on the computing device is trained using a training dataset that uses a plurality of training samples, where the training dataset, training samples, and resulting trained decision tree ensemble are as described above. The analytical characterization and detector are as described above, such as a GC-MS, preferably an SPME GC-MS.

[0083] A computing device useful in the present invention may include a processor and data storage that stores computer-executable instructions that, when executed by the processor, cause the computing device to perform functions including inputting VOC concentration data into a decision tree ensemble and outputting a predicted odor intensity.

[0084] Computing devices useful in the present invention may be client devices (e.g., devices actively operated by a user), server devices (e.g., devices that provide computing services to client devices), or some other type of computing platform. Some server devices can sometimes act as client devices to perform certain operations, and some client devices can incorporate server functionality.

[0085] Processors useful in the present invention may be one or more of any type of computer processing element, such as a central processing unit (CPU), a coprocessor (e.g., a mathematics, graphics, neural network, or cryptography coprocessor), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a network processor, and / or any form of integrated circuit or controller that performs processor operations.

[0086] The data storage may include one or more data storage arrays including one or more drive array controllers configured to manage read and write access to a group of hard disk drives and / or solid state drives.

[0087] In some embodiments, computing devices may be deployed to support a clustered architecture. The exact physical location, connectivity, and configuration of these computing devices may be unknown and / or unimportant to the client device. Thus, the computing devices may be referred to as "cloud-based" devices that may be housed in various remote data center locations, such as cloud-based server clusters. Preferably, the computing devices are cloud-based server clusters, and inputting cardinality data into the decision tree ensemble is performed via a web-based user interface accessible by a user.

[0088] FIG. 3 shows a schematic diagram of a cloud-based server cluster 300 according to one embodiment of the present invention. Desirably, computing device operations may be distributed among server devices 302, data storage 304, and routers 306, all of which may be connected by a local cluster network 308. The number of server devices 302, data storage 304, and routers 306 in the server cluster 300 may depend on the computing task(s) and / or applications assigned to the server cluster 300. For example, the server devices 302 may be configured to perform various computing tasks of the computing devices. Thus, computing tasks may be distributed among one or more of the server devices 302. By way of example, the data storage 304 may store any type of database, such as a structured query language (SQL) database or trained model checkpoints. Furthermore, any database in the data storage 304 may be monolithic or distributed across multiple physical devices. The router 306 may include network equipment configured to provide internal and external communications for the server cluster 300. For example, router 306 may include one or more packet switching and / or routing devices (including switches and / or gateways) configured to provide (i) network communications between server device 302 and data storage 304 via cluster network 308, and / or (ii) network communications between server cluster 300 and other devices via communication link 310 to network 312. Server device 302 may be configured to send and receive data to and from cluster data storage 304. Furthermore, server device 302 may organize the received data into web page representations. Such representations may take the form of a markup language, such as Hypertext Markup Language (HTML), Extensible Markup Language (XML), or some other standardized or proprietary format.Additionally, server device 302 may be capable of executing various types of computerized scripting languages, such as Perl, Python, PHP Hypertext Preprocessor (PHP), Active Server Pages (ASP), or JavaScript. Computer program code written in these languages ​​may facilitate the delivery of web pages to client devices, as well as client device interaction with the web pages.

[0089] 4 shows a schematic block diagram of an odor prediction system 400 according to one embodiment of the present invention. Desirably, the odor prediction system 400 includes a GC-MS 401 for determining the concentration of volatile organic compounds (VOCs) in a sample and a cloud-based computing platform 402, such as a cloud-based server cluster. The trained decision tree ensemble model is deployed on the cloud-based computing platform 402 to provide web-based access to end users. The GC-MS 401 includes a sample injector 4011, a GC column 4012, and an MSD 4013. The sample is first analytically characterized using the GC-MS 401 to generate VOC concentration data from the sample, and the resulting concentration data is then input into the cloud-based computing platform 402. The odor prediction system 400 may further include a sampling unit 403 including a headspace vial 4032 and an SPME fiber 4033. Desirably, sample 4031 (e.g., an aqueous polymer composition or coating) is added to headspace vial 4032, and then SPME fiber 4033 is inserted into headspace vial 4032 and exposed to the headspace of the vial to extract VOCs in the sample. The VOCs extracted from sample 4031 are then injected into sample injector 4011 and analyzed in GC-MS 401 to generate a VOC profile. The resulting VOC profile is used as input and uploaded via a web-based user interface to a cloud-based computing platform 402 deployed with a model. The model then predicts the odor intensity of each sample based on the VOC profile of such sample. [Example]

[0090] Some embodiments of the present invention will now be described in the following examples. PRIMAL, FORMASHIELD, and RHOPLEX are all trademarks of The Dow Chemical Company.

[0091] Methyl methacrylate, butyl acetate, xylene (including three isomers), butyl ether, styrene, butyl acrylate, propanoic acid, butyl ester, 3-methyl-4-heptone, and benzaldehyde are available from Sinopharm Chemical Reagent Co., Ltd.

[0092] Thirty-nine commercially available water-based acrylic binders from various suppliers (including PRIMA™, FORMASHIELD™, and RHOPLEX™ emulsions available from The Dow Chemical Company, ACRONAL emulsions available from BASF, ARCHSOL™ emulsions available from Wanhua, and RS series emulsions available from BATF) were used as training and validation samples, with 31 samples randomly selected as training samples and 8 samples randomly selected as validation samples. These samples are randomly labeled with sample codes "1" through "39" in Tables 2-4.

[0093] Six other commercially available acrylic binders from various sources as described above for the training and validation samples are used as new samples for testing the model and are randomly labeled with the numbers "41" to "46" as sample codes in Table 5.

[0094] PRIMA™ SF-180 emulsion, PRIMAL™ DC-430V styrene acrylic copolymer emulsion, and PRIMAL™ SF-508M styrene acrylic emulsion, which contain 100% acrylic polymer, are all commercially available from The Dow Chemical Company.

[0095] The following standard analytical equipment and methods are used in the examples and in determining the properties and characteristics described herein.

[0096] SPME coupled with GC-MS analysis 1) Preparation of external standards Standard mixtures were prepared by mixing methyl methacrylate, butyl acetate, xylene (including three isomers), butyl ether, styrene, butyl acrylate, propanoic acid, butyl ester, 3-methyl-4-heptone, and benzaldehyde with PRIMAL™ SF-180 emulsion, with each of these compounds present at a concentration of 10,000 ppm based on the wet weight of PRIMAL™ SF-180 emulsion. The prepared standard mixtures were further diluted to different concentrations for each compound, such as 0.005 ppm, 0.01 ppm, 0.05 ppm, 0.1 ppm, 0.5 ppm, 1 ppm, 5 ppm, 10 ppm, 50 ppm, 100 ppm, and 500 ppm based on the wet weight of PRIMAL™ SF-180 emulsion, for use as external standards for quantification of different VOCs.

[0097] 2) Automated SPME GC-MS Test Parameters Approximately 0.05 g of aqueous binder sample was weighed into a 20 mL headspace vial.

[0098] A multipurpose sampler (MPS) (Gerstel) with an SPME unit is connected to the GC. The SPME fiber available from Anple is polydimethylsiloxane / divinylbenzene (PDMS / DVB, 65 millimeters (mm), catalog 57345-U, Supelco Co., Ltd.). The parameters for the SPME process are as follows: incubation temperature: 60°C, extraction time: 30 minutes, and desorption time: 90 seconds.

[0099] GC-MS analysis was carried out using an Agilent 6890 gas chromatograph coupled with a mass spectrometry detector (Agilent 5975C MSD) under the conditions listed in Table 1 .

[0100] [Table 1]

[0101] Odor Panel Test The tests were carried out according to the VDA270 standard, using the experimental procedures and sample preparation described below. In one sensory test, three test samples were prepared for evaluation by each panelist. Two benchmark samples, PRIMAL™ DC-430V and PRIMAL™ SF-508M, were used in each sensory test for reference. The odor intensity values ​​of PRIMAL™ SF-508M and PRIMAL™ DC-430V emulsions are rated as "1.5" and "5.0", respectively, according to the criteria set forth in the VDA 270 standard:

[0102] Aliquots of 0.5 g of binder sample were placed in 100 mL glass vials with odorless caps. The vials containing the samples were equilibrated at room temperature for 2 hours before evaluation. Each panelist independently received one set of three samples. Randomly selected labels were assigned to each sample for blind sample identification. The order in which the samples were presented to the panelists was also randomized.

[0103] For each test, 8–10 well-trained panelists (certified by SGS Co., Ltd. for odor intensity training) were recruited to perform sensory evaluations of these samples. Analysis of variance for the sensory evaluations was automatically calculated by CSAS (Conventional Sensory Analysis System) sensory software (ISENSO Co., Ltd., Shanghai, China).

[0104] Odor intensity is assessed and rated according to the criteria set out in the VDA 270 standard (the smallest unit is 0.5): 1. Not perceptible, 2. Perceptible but not bothersome, 3. Clearly perceptible but not bothersome, 4. Bothersome, 5. Very bothersome, 6. Unacceptable.

[0105] The average odor intensity as assessed by the panelists is reported for each binder sample and is designated as "Actual Odor Intensity."

[0106] Model The random forest (RF) model, gradient boosting (GB) model, artificial neural network (ANN) model, partial least squares regression (PLS) model, and ridge regression (RP) model used in the following examples (IE1-3 and CE1-3) were each independently constructed based on Python Scikit-learn (further details can be found in "Scikit-learn: Machine Learning in Python," Pedregosa, F. et al., Journal of Machine Learning Research, 2011, 12, 2825-2830). Each of these models was independently constructed using Scikit-learn 0.23.2 in a Python 3.8 environment.

[0107] The eXtreme Gradient Boosting (XGB) model was built using the XGBoost Package version 1.4.2 with a Python 3.8 environment (further details can be found in "XGBoost: A Scalable Tree Boosting System," Chen, T., and Guestrin, C., Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016, 785-794, ACM, New York, USA).

[0108] Each model was trained on a total of 39 data pairs from the training samples, divided into a training dataset with 31 data pairs (dataset type: "train") and a validation dataset with 8 data pairs (dataset type: "validation"). The input to the model was the concentration of 27 selected VOCs measured by GC-MS, and the output of the model was the odor intensity of the training samples as assessed by human panelists according to the odor panel test described above. The results are shown in Tables 2-4. To verify the effectiveness and accuracy of the method using each model, the VOC concentrations of six new samples (i.e., test samples) were measured and input into each trained model. The predicted odor intensity values ​​were obtained and compared with the manual evaluation results of odor intensity assessed by human panelists in Table 5.

[0109] IE-1 An automated odor intensity prediction tool was established by combining SPME GC-MS, odor panel testing, and an RF model. The hyperparameters of the RF model are listed below. Number of estimators = 100, minimum number of samples per split = 2, minimum samples per leaf node = 1, minimum weight fraction of leaf nodes = 0, maximum feature value of split

[0110]

number

[0111] IE-2 IE-2 was performed following the same procedure as IE-1, except that the RF model was replaced with a GB model with the hyperparameters shown below. Learning rate = 0.1, number of estimators = 100, fraction of samples used for each base learner = 1, minimum number of samples per split = 2, minimum samples per leaf node = 1, minimum weight fraction for leaf nodes = 0, maximum individual tree depth = 3, and minimum impurity reduction = 0.

[0112] IE-3 IE-3 was performed following the same procedure as IE-1, except that the RF model was replaced with an XGB model with the hyperparameters shown below. Step size reduction = 0.3, minimum impurity reduction = 0, maximum individual tree depth = 6, minimum sum of instance weights = 1, maximum delta step = 0, minimum samples per leaf node = 1, cosamples by tree, cosamples by leaf, cosamples by node = 1, L2 regularization term on weights = 1, and L1 regularization term on weights = 0.

[0113] CE-1 CE-1 was performed following the same procedure as IE-1, except that the RF model was replaced with an ANN model with the hyperparameters shown below. Hidden layer size = 100, activation = ReLU (rectified linear units), solver = Adam, alpha = 0.0001, batch size = number of samples, learning rate = 0.001, max number of iterations = 200, shuffle = true, early stopping = false, beta 1 = 0.9, beta 2 = 0.999, epsilon = 10 -8 .

[0114] CE-2 CE-2 was performed following the same procedure as IE-1, except that the RF model was replaced with a PLS model with the following hyperparameters: number of components = 2, scale = true, maximum number of iterations = 500, tolerance = 10. -6 , copy = true ;

[0115] CE-3 CE-3 was performed following the same procedure as IE-1, except that the RF model was replaced with an RR model with the hyperparameters shown below. Alpha=0.2, Fit Intercept=True, Normalize=False, Copy X=True, Max Number of Replicates=15000, Tolerance=10 -3 , solver=auto, true=false.

[0116] The raw data (including VOC concentrations and actual odor intensities) and model predictions for the training and validation samples according to the methods described in the above Examples (IE-1 to IE-3 and CE-1 to CE-3) are shown in Tables 2 to 4. Table 5 shows the odor evaluation results for the test samples. In these tables, actual odor intensity refers to the odor intensity assessed by panelists according to the odor panel test described above, and "\" means that the concentration is below the detection limit.

[0117] [Table 2-1]

[0118] [Table 2-2]

[0119] [Table 3-1]

[0120] [Table 3-2]

[0121] [Table 4]

[0122] [Table 5]

[0123] Based on the data shown in Tables 2 to 5, training R 2The validation percentage RMSE and test percentage RMSE were calculated based on the above equations (II) and (III), and the results are shown in Table 6. As shown in Table 6, the IE-1 (RF model), IE-2 (GB model), and IE-3 (XGB model) methods achieved higher training Rs than the CE-1 (ANN model), CE-2 (PLS model), and CE-3 (RR model). 2 The validation and test results for the IE-1 to IE-3 methods all met the requirement of percentage RMSE < 30%, while the CE-1 to CE-3 methods all provided test percentage RMSEs higher than 30%. This indicates that the IE-1 to IE-3 methods, all based on decision tree ensembles, exhibited higher prediction accuracy than the CE-1 to CE-3 methods, which were based on ANN, PLS, and RR models, respectively. In particular, IE-1 with the RF model exhibited even higher prediction accuracy (test percentage RMSE = 18.7%) than IE-2 (test percentage RMSE = 21.3%) and IE-3 (test percentage RMSE = 20.6%).

[0124] [Table 6]

Claims

1. 1. A method for predicting odor of an aqueous polymer composition, comprising: analytically characterizing the aqueous polymer composition with a detector, thereby generating concentration data for volatile organic compounds in the aqueous polymer composition from the analytical characterization; inputting the concentration data into a decision tree ensemble configured to predict an odor intensity of the waterborne polymer composition based on the concentration data; and outputting a predicted odor intensity of the waterborne polymer composition from the decision tree ensemble.

2. 2. The method of claim 1, wherein the decision tree ensemble is trained to predict the odor intensity of the aqueous polymer composition using a training dataset using a plurality of training samples, the training dataset including the concentration data of volatile organic compounds in each training sample paired with actual odor intensity data evaluated by human panelists for such training samples, thereby providing a trained decision tree ensemble.

3. 3. The method of claim 2, wherein the trained decision tree ensemble is validated on a validation dataset using a plurality of validation samples, the validation dataset including the concentration data of volatile organic compounds in each validation sample paired with the actual odor intensity data evaluated by a human panelist for such validation sample.

4. The method of claim 1 , wherein the decision tree ensemble is selected from a random forest model, a gradient boosting model, or an eXtreme gradient boosting model.

5. 3. The method of claim 2, wherein the training data set includes concentration data for volatile organic compounds each having a concentration of ≧0.1 ppm by weight, based on the weight of the aqueous polymeric composition.

6. The method of claim 1 , wherein the concentration data input to the decision tree ensemble is performed via a web-based user interface.

7. 2. The method of claim 1, wherein the concentration data input to the decision tree ensemble is a concentration of a volatile organic compound comprising acetone, 2-methylpropanol, 1-butanol, methyl methacrylate, butyl acetate, 4-heptanone, 2-heptanone, butyl ether, styrene, butyl acrylate, anisole, propanoic acid, butyl ester, methyl ethyl benzene, 3-methyl-4-heptanone, propenyl benzene, propyl benzene, benzaldehyde, acetophenone, butyl methacrylate, isobutyl vinyl acetate, butanoic acid, butyl ester, 2-butenoic acid, butyl ester, diethylbenzene or an isomer, cyclohexyl methacrylate, 2-ethylhexyl acrylate, xylene, ethylbenzene, or a mixture thereof.

8. 10. The method of claim 1, wherein analytically characterizing the aqueous polymer composition comprises an analytical characterization selected from solid phase microextraction coupled with gas chromatography-mass spectrometry, needle trap microextraction coupled with gas chromatography-mass spectrometry, or Tenax sorbent cartridge coupled with gas chromatography-mass spectrometry.

9. The decision tree ensemble has a test percentage RMSE < 30% and a training R 2 3. The method of claim 2, wherein the method exhibits a predictive accuracy indicated by >0.

85.

10. 10. The method of claim 1, further comprising adjusting the polymerization process for preparing the aqueous polymer composition based on the predicted odor intensity.

11. The method of claim 1 , wherein the aqueous polymer composition comprises an acrylic (co)polymer.

12. 1. A method for predicting the odor of a coating, comprising: analytically characterizing the coating using a detector, thereby generating concentration data of volatile organic compounds in the coating from the analytical characterization, the coating being obtained by drying an aqueous polymer composition; inputting the concentration data into a decision tree ensemble configured to predict an odor intensity of the coating based on the concentration data; and outputting a predicted odor intensity of the coating from the decision tree ensemble.

13. 1. A system for predicting the odor of an aqueous polymer composition or a coating made therefrom, comprising: a detector configured to analytically characterize the aqueous polymer composition or the coating, thereby generating volatile organic compound concentration data from the analytical characterization; a computing device having a decision tree ensemble deployed thereon, the computing device being configured to input the concentration data and output a predicted odor intensity of the aqueous polymer composition or the coating.

14. 14. The system of claim 13, wherein the decision tree ensemble deployed on the computing device is trained to predict the odor intensity of the aqueous polymer composition or the coating using a training dataset using a plurality of training samples, the training dataset including the concentration data of volatile organic compounds in each training sample paired with actual odor intensity data evaluated by human panelists for such training samples, thereby providing a trained decision tree ensemble.

15. 14. The system of claim 13, wherein the computing device is a cloud-based server cluster and inputting the concentration data into the decision tree ensemble is performed via a web-based user interface.

Citation Information

Patent Citations

  • Water-dispersed acrylic pressure-sensitive adhesive composition, pressure-sensitive adhesive sheet, and method for producing the same

    JP2011093956A

  • Superabsorbent polymer having improved odor prevention performance and manufacturing method therefor

    JP2015168824A

  • Methods for identifying, compounds identified and compositions thereof

    US20200399558A1

  • Method and system for providing machine learning service

    US20200401886A1