Method for learning prediction model

By using ensemble learning methods and data preprocessing, the method addresses imbalanced data issues in cosmetic risk prediction, improving accuracy in predicting skin irritation and other risks from cosmetics.

JP2025078464APending Publication Date: 2025-05-20POLA CHEMICAL INDUSTRIES INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023191058
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

Conventional prediction models struggle to accurately predict the level of risk to the human body from cosmetics due to imbalanced ingredient data, leading to overlearning or underlearning, especially when distinguishing between low-risk and high-risk cosmetics.

Method used

A method for learning a predictive model using raw material blend data as features and risk level as an objective variable, employing ensemble learning methods like Balanced Random Forest, One-Class Support Vector Machine, Local Outlier Factor, and Isolation Forest, with data preprocessing to balance high-risk and low-risk data.

Benefits of technology

The method improves prediction accuracy by suppressing overlearning and underlearning, effectively distinguishing between normal and abnormal data, even with imbalanced learning data, enhancing the model's ability to predict skin irritation and other risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025078464000001_ABST
    Figure 2025078464000001_ABST
Patent Text Reader

Abstract

To provide a method for learning a prediction model capable of improving the prediction accuracy in predicting a risk degree of cosmetics to a human body by a prediction model.SOLUTION: An information processing device 1 acquires learning data by executing high correlation check processing and spareness check processing to unpressed data with the total compounding amount of each raw material category in formulations of a plurality of cosmetics as a feature amount and with the risk degree in the formulations of the plurality of cosmetics to the human bod as an objective variable, and executes learning a prediction model being a division model type by a Balanced Random Forest by using the learning data.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method for learning a prediction model that performs learning of a prediction model that predicts the level of risk of cosmetics to the human body. [Background technology]

[0002] A method for determining the raw material composition of a cosmetic product is known, as described in Patent Document 1. In this method, supervised machine learning using training data is used to learn a prediction model (machine learning model) in which the blending ratios of each raw material classification of the cosmetic product's raw materials are used as feature quantities, and evaluation items such as the remaining mass after evaporation are used as objective variables. Then, using this prediction model, evaluation items of the cosmetic product, such as the remaining mass after evaporation, are determined (predicted), and the raw material composition of the cosmetic product is determined based on the determination results. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2023-10289 A Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, it has become desirable to predict the level of risk to the human body, such as the presence or absence of skin irritation when used, for cosmetics. For example, when the above-mentioned conventional prediction model is used to predict the level of risk to the human body of a cosmetic instead of the cosmetic evaluation items, there is a problem that the prediction cannot be made appropriately. That is, since cosmetics contain a very large number of ingredients, it is difficult to obtain ingredient data of high-risk cosmetics and ingredient data of low-risk cosmetics in a balanced manner. Therefore, when learning of a prediction model is performed using such ingredient data, there is a risk that overlearning or underlearning occurs in the learning of the prediction model.

[0005] The present invention has been made to solve the above-mentioned problems, and aims to provide a method for learning a predictive model that can improve the prediction accuracy when predicting the level of risk of cosmetics to the human body using a predictive model. [Means for solving the problem]

[0006] In order to achieve the above object, the invention of claim 1 is a method for learning a predictive model, which uses an information processing device to learn a predictive model that predicts the level of risk to the human body, including the presence or absence of skin irritation, when multiple cosmetics are used. The information processing device executes a learning data acquisition step of acquiring learning data by performing a predetermined process on unprocessed data in which raw material blend data representing the blending state of raw materials in the formulations of the multiple cosmetics is used as a feature and the level of risk in the formulations of the multiple cosmetics is used as an objective variable; and a learning step of using the learning data to learn a predictive model, which is a classification model, by any one of a predetermined ensemble learning method, One-Class Support Vector Machine, Local Outlier Factor, and Isolation Forest.

[0007] According to this method of training a predictive model, raw material blend data that represents the blending state of raw materials in multiple cosmetic formulations is used as a feature, and training data is obtained by performing a predetermined processing on unprocessed data in which the level of risk in multiple cosmetic formulations is used as the objective variable. Using such training data, training of a predictive model, which is a classification model, is performed using one of a predetermined ensemble learning method, One-Class Support Vector Machine, Local Outlier Factor, and Isolation Forest.

[0008] In the case of cosmetic prescription data, it is difficult to accumulate data in a state that can express a range of numerical data of risk to the human body, and as a result, when a regression model type prediction model is used in learning a prediction model that predicts the level of risk to the human body when using a cosmetic, overlearning or underlearning may occur. In contrast, when a prediction model that is a classification model with the level of risk as the objective variable as in the present invention is used, overlearning or underlearning can be suppressed in the learning of the prediction model compared to when a regression model type prediction model is used, due to the objective variable being binary. In addition, by using a predetermined ensemble learning method in the learning of the classification model, overlearning and underlearning can be suppressed compared to a learning method that executes learning of a single model.

[0009] In addition, the anomaly detection algorithms One-Class Support Vector Machine, Local Outlier Factor, and Isolation Forest have the property of being able to appropriately distinguish abnormal data from binary learning data with a very large bias by precisely learning the characteristics of a large amount of normal data. Therefore, when any of One-Class Support Vector Machine, Local Outlier Factor, and Isolation Forest is used in learning a prediction model as in the present invention, the above-mentioned property makes it possible to improve the prediction accuracy of the prediction model even when the learning data is imbalanced data in which low-risk data and high-risk data are in an imbalanced state, such as cosmetic prescription data. In this specification, "learning a prediction model" means learning model parameters of the prediction model.

[0010] In the present invention, it is preferable that the specified processing includes a first deletion processing for deleting data having a value of 0 from unprocessed data when the data having a value of 0 is contained in a specified proportion or more in all of the formulations of multiple cosmetic products among the raw material blend data.

[0011] According to this method for learning a predictive model, when learning data is obtained from unprocessed data, a first deletion process is performed, so that if the raw material blend data contains data with a value of 0 at a predetermined rate or more in all of the formulations of a plurality of cosmetics, the data with a value of 0 is deleted from the unprocessed data. This makes it possible to obtain learning data that includes information suitable for learning, and makes it possible to improve the prediction accuracy of the predictive model with the learning data.

[0012] In the present invention, the specified processing is characterized in that, when there are multiple data whose correlation coefficient between raw material blending data is equal to or greater than a specified value, the second deletion processing is performed to delete data other than one of the multiple data from the unprocessed data.

[0013] According to this method of learning a predictive model, when learning data is acquired, if there are multiple pieces of data with a correlation coefficient between raw material blend data that is equal to or greater than a predetermined value, only one of the multiple pieces of data is acquired as a feature, thereby making it possible to delete highly correlated data that becomes noise during learning, thereby further improving the predictive accuracy of the predictive model.

[0014] In the present invention, the predetermined ensemble learning method preferably includes a Balanced Random Forest.

[0015] In the case of cosmetics, it is necessary to design the formulation taking into consideration the level of risk to the human body when used and international regulations, and as a result, low-risk formulations are designed in the prescription data of many cosmetics. Therefore, inevitably, the accumulated prescription data of cosmetics becomes imbalanced data in which low-risk data and high-risk data are in an imbalanced state. In contrast, when Balanced Random Forest is used as a learning method as in the present invention, a prediction model is constructed in a state in which low-risk learning data and high-risk learning data are sampled in equal numbers. As a result, compared to the case of using a general Random Forest, it is possible to further suppress overlearning and underlearning, and to further improve the learning accuracy of the prediction model. [Brief description of the drawings]

[0016] [Figure 1] FIG. 1 is a diagram illustrating an information processing device that executes a prediction model learning method according to an embodiment of the present invention. [Diagram 2] FIG. 13 is a diagram illustrating an example of learning data. [Diagram 3] 13 is a flowchart showing a process of creating learning data. [Figure 4] 13 is a flowchart showing a first learning process. [Diagram 5] 13 is a flowchart showing a second learning process. [Figure 6] FIG. 13 is a diagram showing the prediction accuracy of a trained prediction model. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0017] Hereinafter, a method for learning a predictive model according to an embodiment of the present invention will be described with reference to the drawings. In this embodiment, a predictive model (classification model) that predicts the level of risk to the human body, such as the presence or absence of skin irritation, when multiple cosmetics are used is learned by a learning method described later.

[0018] The learning method of this embodiment is specifically executed by an information processing device 1 shown in Fig. 1. This information processing device 1 is a personal computer type, and includes a display 1a, a device main body 1b, and an input interface 1c. The device main body 1b includes a storage such as an HDD, a processor, and a memory (RAM, E2PROM, ROM, etc.) (none of which are shown).

[0019] The memory of the device main body 1b stores the learning data shown in Fig. 2 and unprocessed data (not shown) on which the learning data is based. Here, the unprocessed data was obtained by the applicant's experiments, and the learning data was created by applying a check process to the unprocessed data, which will be described later.

[0020] The learning data shown in Fig. 2 is data that represents the level of risk (i.e., the level of risk) in the prescription data of n cosmetics (n is several hundreds), and is configured to include the total amount (indicated as "total amount" in Fig. 2) of each of m ingredient categories (m is several tens) as a feature. As shown in the figure, the ingredient categories include moisturizing ingredients, oils, surfactants, etc. In this embodiment, the total amount of each ingredient category corresponds to the ingredient combination data.

[0021] Furthermore, in the learning data, the value 1 and the value 0 are set as objective variables (labels) for five types of risks 1 to 5. In this case, the value 1 indicates that any of the five types of risks 1 to 5 is high, and the value 0 indicates that any of the five types of risks 1 to 5 is low.

[0022] Risk 1 represents eye irritation, risk 2 represents primary skin irritation caused by the cosmetic, and risk 3 represents objective symptoms when using the cosmetic. Risk 4 represents subjective symptoms when using the cosmetic, and risk 5 represents minor skin trouble when using the cosmetic. In this embodiment, risks 1 to 5 correspond to risks.

[0023] On the other hand, application software for executing a learning data acquisition process, which will be described later, etc. is installed in the storage of the device main body 1b. The input interface 1c is composed of a keyboard, a mouse, etc. for operating the information processing device 1.

[0024] Next, the learning data acquisition process will be described with reference to Fig. 3. This process is for acquiring learning data from unprocessed data as described below.

[0025] As shown in Fig. 3, in this learning data acquisition process, first, a process of reading unprocessed data is executed (Fig. 3 / STEP 1). In this reading process, unprocessed data (not shown) stored in the memory of the information processing device 1 is read. This unprocessed data includes total blend amounts for each raw material category as raw material blend data for many cosmetic formulations.

[0026] Next, a high correlation check process is executed (FIG. 3 / STEP 2). In this high correlation check process, the correlation coefficient between the raw material blending data is calculated. If the correlation coefficient between the raw material blending data is equal to or greater than a predetermined value (e.g., 0.99), one raw material blending data is left, and the other raw material blending data is deleted from the unprocessed data. This is to improve the prediction accuracy of the prediction model by avoiding the occurrence of multicollinearity.

[0027] Next, a sparseness check process is performed (FIG. 3 / STEP 3). In this sparseness check process, if raw material composition data with a value of 0 is included in the entire recipe data at a predetermined ratio R or more in the unprocessed data after the high correlation check process, the raw material composition data is deleted. In this case, the predetermined ratio R is set to an appropriate value between 80% and 100%.

[0028] As described above, the high correlation check process and the sparseness check process are executed on the unprocessed data to obtain the learning data shown in Fig. 2. Then, the learning data obtained as described above is stored in the memory of the information processing device 1 (Fig. 3 / STEP 4).

[0029] In this embodiment, the high correlation check process and the sparseness check process correspond to the predetermined process and the learning data acquisition step, the high correlation check process corresponds to the second deletion process, and the sparseness check process corresponds to the first deletion process.

[0030] Next, the first learning process will be described with reference to Fig. 4. As described below, the first learning process is a process in which model parameters of a prediction model (classification model) are learned by a Balanced Random Forest using the above-mentioned learning data.

[0031] As shown in Fig. 4, in this first learning process, first, a data setting process is executed (Fig. 4 / STEP 10). In this data setting process, in the learning data stored in the memory of the information processing device 1, the learning data is set so that the number of prescription data for high-risk cosmetics is equal to the number of prescription data for low-risk cosmetics for each of the risks 2 to 4 described above, risks 2_3 and 4_5 described below, and a total risk described below.

[0032] Here, since the number of prescription data for high-risk cosmetics is smaller than the number of prescription data for low-risk cosmetics in the learning data stored in the memory of the information processing device 1, the learning data is set so that the number of both is equal by thinning out the prescription data for low-risk cosmetics.

[0033] Next, a model parameter learning process is executed (FIG. 4 / STEP 11). In this model parameter learning process, model parameters of the prediction model are learned by a Balanced Random Forest. Specifically, multiple learning devices (prediction models) are constructed by repeatedly growing (branching) a decision tree using data randomly sampled from the learning data set as described above. Then, model parameters of the prediction model are learned by repeatedly constructing the above decision tree.

[0034] Next, the second learning process will be described with reference to Fig. 5. As described below, the second learning process uses the above-mentioned learning data to learn model parameters of a prediction model (classification model) by a One-Class Support Vector Machine, a Local Outlier Factor, and an Isolation Forest.

[0035] In this second learning process, a model parameter learning process is executed (FIG. 5 / STEP 20). In this model parameter learning process, model parameters of a prediction model are learned by a One-Class Support Vector Machine using the above-mentioned learning data.

[0036] Furthermore, model parameters of the prediction model are learned by Local Outlier Factor using the above-mentioned learning data. Also, model parameters of the prediction model are learned by Isolation Forest using the above-mentioned learning data.

[0037] In this embodiment, the model parameter learning processes in STEPs 11 and 20 correspond to a learning step.

[0038] Next, the prediction accuracy of a prediction model (referred to as "learned model" in FIG. 5) in which learning of model parameters has been completed by executing the learning process as described above will be described with reference to FIG. 5. The data in FIG. 5 is data of a prediction model that showed the best prediction accuracy when predicting each risk using a prediction model trained with Balanced Random Forest, One-Class Support Vector Machine (referred to as One-Class SVM in FIG. 5), Local Outlier Factor, and Isolation Forest.

[0039] In addition, these prediction accuracies are specifically the results of verification by 5-fold cross-validation using hundreds of prescription data. In other words, hundreds of prescription data were divided into five groups, and one of the groups was used as test data, and the remaining group was used as learning data to derive the prediction accuracy of the prediction model. This process was repeated five times with different test data, and the five prediction accuracies were averaged to derive the prediction accuracy of the prediction model.

[0040] Furthermore, the Risk 2_3 data represents the verification results of the combined data of Risk 2 and Risk 3 because the amount of data for Risk 2 is small, and the Risk 4_5 data represents the verification results of the combined data of Risk 4 and Risk 5 because the amount of data for Risk 5 is small. Furthermore, the overall risk data represents the verification results of the combined data of Risks 1 to 5.

[0041] As is clear from Figure 5, the recall and accuracy are good for all data. Furthermore, in the case of the risk 1 and total risk data, the Kappa coefficient is good in addition to the recall and accuracy.

[0042] As described above, according to the learning method of the prediction model of the present embodiment, when learning data is obtained from unprocessed data, high correlation check processing and sparseness check processing are performed. In this high correlation check processing, the correlation coefficient between raw material blending data is calculated, and when the correlation coefficient of multiple raw material blending data is a predetermined value (for example, 0.99) or more, one raw material blending data is left, and the other raw material blending data is deleted from the unprocessed data. Thereby, the highly correlated data that becomes noise during learning can be deleted from the learning data, and the learning data can be created as one that has information suitable for learning.

[0043] In addition, in the sparseness check process, if data of an ingredient category whose ingredient blend data has a value of 0 is included in all of the formulations of multiple cosmetics at a predetermined rate R or more, the data of this ingredient category is deleted from the unprocessed data. This makes it possible to create learning data that includes information more suitable for learning.

[0044] Furthermore, using the above-mentioned learning data, a classification model type prediction model is learned by the learning methods Balanced Random Forest, One-Class Support Vector Machine, Local Outlier Factor, and Isolation Forest. In this way, when Balanced Random Forest is used as the learning method, a prediction model is constructed with the same number of low-risk data and high-risk learning data set. As a result, even when using cosmetic prescription data that is prone to becoming imbalanced data, overlearning and underlearning can be further suppressed compared to the general Random Forest, and the learning accuracy can be further improved.

[0045] Furthermore, when one of the One-Class Support Vector Machine, Local Outlier Factor, and Isolation Forest is used as the learning method, due to the above-mentioned characteristics, overlearning and underlearning can be further suppressed even when cosmetic prescription data that is likely to be imbalanced data is used. As a result, the learning accuracy of the prediction model can be improved.

[0046] In the embodiment, the total blend amount for each raw material category is used as the raw material blend data in the formulation of multiple cosmetic products. However, instead of this, the raw material blend ratio for each raw material category or the blend ratio of each raw material may be used as the raw material blend data.

[0047] In addition, in the embodiment, risks 1 to 5 are used as risks to the human body of cosmetics, but risk items further subdivided according to symptoms may be used.

[0048] Furthermore, in the embodiment, a personal computer type is used as the information processing device, but the information processing device of the present invention is not limited to this, and may be any device capable of executing learning of a prediction model. For example, one server or a combination of multiple servers may be used as the information processing device. Furthermore, a combination of multiple personal computers may be used as the information processing device, or a combination of a personal computer and a server may be used.

[0049] In addition, the embodiment is an example in which both the high correlation check process and the sparseness check process are executed in the learning data acquisition process of Figure 4, but instead, the learning data may be acquired by executing only one of the high correlation check process and the sparseness check process.

[0050] Furthermore, although the embodiment is an example in which Balanced Random Forest is used as the predetermined ensemble learning method, XGBoost may be used as the predetermined ensemble learning method instead of or in addition to this. Furthermore, Random Forest, Gradient Boosting, AdaBoost, LightGBM, CatBoost, Stacked Ensemble, etc. may be used as the predetermined ensemble learning method. [Explanation of symbols]

[0051] 1. Information processing device R Predetermined ratio

Claims

1. A method for learning a prediction model, which uses an information processing device to learn a prediction model that predicts the level of risk to the human body, including the presence or absence of skin irritation, when using a plurality of cosmetics, comprising: The information processing device includes: a learning data acquisition step of acquiring learning data by performing a predetermined process on unprocessed data in which raw material blend data representing the blending state of raw materials in the formulations of the plurality of cosmetic products is used as a feature amount and the level of the risk in the formulations of the plurality of cosmetic products is used as a target variable; A learning step of learning the prediction model, which is a classification model, by using the learning data and any one of a predetermined ensemble learning method, a One-Class Support Vector Machine, a Local Outlier Factor, and an Isolation Forest; A method for learning a predictive model, comprising:

2. 2. The method for learning a prediction model according to claim 1, A method for learning a predictive model, characterized in that the specified processing includes a first deletion process for deleting data with a value of 0 from the unprocessed data when the raw material blend data contains data with a value of 0 at a specified rate or more in all of the formulations of the plurality of cosmetics.

3. 3. The method for learning a prediction model according to claim 1, A method for learning a predictive model, characterized in that the specified processing further includes a second deletion processing for deleting data other than one of the plurality of data from the unprocessed data when there is a plurality of data whose correlation coefficient between the raw material blend data is a predetermined value or more.

4. 3. The method for learning a prediction model according to claim 1, A method for learning a predictive model, wherein the predetermined ensemble learning method includes a Balanced Random Forest.

Citation Information

Patent Citations

  • Cosmetic development support method

    JP2023010289A