Method, device, and computer program for creating training data for predicting chemical composition
By aggregating and ranking raw material usage frequency to create training data, the method addresses inefficiencies in machine learning models for chemical composition prediction, resulting in improved model training and accuracy.
Patent Information
- Application Number
- JP2020191597
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-11-18
- Publication Date
- 2025-08-28
- Estimated Expiration
- 2040-11-18
AI Technical Summary
Existing machine learning models for predicting chemical compositions face inefficiencies due to a large number of zero values in training data, leading to suboptimal model training and reduced prediction accuracy, especially when dealing with a small number of required raw materials for achieving target properties.
A method to create training data by aggregating and ranking raw material usage frequency, identifying a predetermined number of key materials, and correlating them with characteristic values to reduce zero values and improve model training.
This approach significantly reduces zero values in training data, enabling the creation of an optimal trained model and enhancing prediction efficiency for chemical composition formulations.
Smart Images

Figure 0007730633000010 
Figure 0007730633000011 
Figure 0007730633000012
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method, an apparatus, and a computer program for generating training data for predicting chemical compositions. [Background technology]
[0002] Patent Document 1 describes a technique for machine learning related to predicting vulcanized rubber compositions, in which if the number of learning input data before machine learning is less than a predetermined number, the learning input data is increased by interpolating samples created by regression analysis.Patent Document 2 describes a technique for estimating factor values and characteristic values using machine learning, in which if the learning data contains data that includes blank fields, the blank data is ignored and learning is performed using only valid data.
[0003] For chemical industry customers, predicting the properties of targeted chemical compositions is important. Typical target properties of chemical compositions include thermal conductivity, particle size, particle distribution, shape, filler treatment, Morse hardness, cure rate, hardness, bond line thickness, and electrical conductivity. For example, a thermally conductive silicone composition is manufactured by blending raw materials such as a thermally conductive filler and a silicone matrix. Traditionally, the blending amounts of a composition that meet the desired properties have been determined through trial and error based on existing compositions. While the relationship between these blending amounts and the properties of a targeted chemical composition can be calculated to some extent through modeling, the modeling process is extremely complex and time-consuming.
[0004] Currently, there is a technology that simplifies property prediction based on blend amounts by incorporating machine learning into this modeling. However, while there are many candidate raw materials for preparing a chemical composition, it is often the case that only a small number of raw materials are required to create a chemical composition that meets the target properties. When a neural network is trained to learn training data that indicates the relationship between the properties of a chemical composition and its raw materials, the training data will show zero values for raw materials that are not used, resulting in the training data containing many zero values. In this case, the large number of zero values hinders the learning of the causal relationship between the property values and the raw materials. This results in the problem of not creating an optimal trained model and reducing prediction efficiency. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 2020-38495 [Patent Document 2] Japanese Patent Application Laid-Open No. 2003-58582 Summary of the Invention [Problem to be solved by the invention]
[0006] The present invention aims to provide a method, device, and computer program for creating training data for predicting chemical compositions that solves the above-mentioned problems by reducing the amount of calculation in machine learning and improving prediction accuracy. [Means for solving the problem]
[0007] According to a first aspect of the present invention, a method for creating training data to be used in machine learning includes the steps of: aggregating the number of times raw materials are used in a dataset of multiple chemical compositions, the dataset including identification names of the multiple chemical compositions, characteristic value information indicating characteristic values of each of the multiple chemical compositions, and raw material information about the raw materials that constitute each of the multiple chemical compositions; ranking the raw materials based on the aggregated number of times they are used; identifying a predetermined number of raw materials from the ranked raw materials; and creating training data for machine learning by correlating the identified raw materials with the characteristic value information.
[0008] According to a second aspect of the present invention, the trained model creation method further includes a step of performing machine learning on the training data created by the first aspect to create a trained model.
[0009] According to a third aspect of the present invention, a method for predicting a blending material of a chemical composition includes inputting a desired property value of a desired chemical composition into a trained model created by the second aspect, and outputting blending materials and proportions to satisfy the desired property value.
[0010] According to a fourth aspect of the present invention, the method for producing a chemical composition further comprises the step of producing a desired chemical composition based on the formulation ingredients and proportions output by the third aspect.
[0011] According to a fifth aspect of the present invention, the training data is created by the training data creation method for use in machine learning according to the first aspect.
[0012] According to a fifth aspect of the present invention, a computer program causes a processor to execute the methods according to the first to fourth aspects.
[0013] According to a sixth aspect of the present invention, a teacher data creation device used in machine learning includes: an aggregation unit that aggregates the number of times raw materials are used in a dataset of multiple chemical compositions, the dataset including identification names of the multiple chemical compositions, characteristic value information indicating characteristic values of each of the multiple chemical compositions, and raw material information about the raw materials that constitute each of the multiple chemical compositions; an aggregation unit that ranks the raw materials based on the aggregated number of times they are used; an identification unit that identifies a predetermined number of raw materials from the ranked raw materials; and a creation unit that correlates the identified raw materials with the characteristic value information to create teacher data for machine learning.
[0014] According to a seventh aspect of the present invention, the trained model creation device further includes a learning unit that performs machine learning on the teacher data created by the teacher data creation device of the sixth aspect to create a trained model.
[0015] According to an eighth aspect of the present invention, a chemical composition formulation material prediction device further includes an output unit that inputs desired characteristic values of a desired chemical composition into a trained model created by the trained model creation device of the seventh aspect, and outputs formulation materials and proportions to satisfy the desired characteristic values.
[0016] According to a ninth aspect of the present invention, a system for producing a chemical composition further includes a production unit that produces a desired chemical composition based on the blending ingredients and their proportions output by the blending ingredient prediction device of the eighth aspect. [Brief explanation of the drawings]
[0017] [Figure 1] FIG. 1 is a diagram showing the minimum configuration of a teacher data creation device according to an embodiment of the present invention. [Figure 2] 1 is a diagram showing the configuration of a chemical composition manufacturing system according to an embodiment of the present invention. [Figure 3] FIG. 2 is a diagram showing a data set according to the present embodiment of the present invention. [Figure 4] FIG. 2 is a diagram showing a processing flow of a teacher data creation method according to the present embodiment of the present invention. [Figure 5] FIG. 1 illustrates ranked raw materials according to an embodiment of the present invention. [Figure 6] FIG. 2 is a diagram showing training data according to the embodiment of the present invention. [Figure 7] FIG. 1 is a diagram showing a process flow for producing a chemical composition using training data according to an embodiment of the present invention. [Figure 8] FIG. 2 is a diagram showing a processing flow of a teacher data creation method according to a first modified example of the present embodiment of the present invention. [Figure 9] FIG. 10 is a diagram showing a processing flow of a teacher data creation method according to a second modified example of the present embodiment of the present invention. [Figure 10] FIG. 10 is a diagram illustrating the configuration of a teacher data creation device according to a third modified example of the present embodiment of the present invention. [Figure 11] FIG. 10 is a diagram showing a processing flow for determining optimal training data according to a third modified example of the present embodiment of the present invention. [Figure 12] FIG. 10 is a diagram illustrating the configuration of a teacher data creation device according to a fourth modified example of the present embodiment of the present invention. [Figure 13] FIG. 10 is a diagram showing a processing flow for determining optimal training data according to a fourth modified example of the present embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0018] Hereinafter, the teacher data creation device according to this embodiment of the present invention will be described with reference to FIGS.
[0019] (Minimum configuration of the teacher data creation device) FIG. 1 is a diagram showing the minimum configuration of a teacher data creation device 1 according to this embodiment of the present invention. The teacher data creation device 1 includes a control unit 11 and a recording unit 12. The teacher data creation device 1 may be configured using a computer such as a server or a personal computer. The recording unit 12 records a dataset 121 for multiple chemical compositions related to a desired chemical composition, as shown in FIG. 3. The dataset 121 includes identification names of the multiple chemical compositions, characteristic value information indicating characteristic values of each of the multiple chemical compositions, and raw material information for the raw materials constituting each of the multiple chemical compositions. The control unit 11 includes a counting unit 111, a ranking unit 112, an identifying unit 113, and a creating unit 114. The counting unit 111 counts the number of times each raw material is used in the dataset 121 based on the raw material information. The ranking unit 112 ranks the raw materials based on the counted number of times each raw material is used. The identifying unit 113 identifies a predetermined number of raw materials from the ranked raw materials. The creation unit 114 creates training data for machine learning by correlating the identified raw materials with the characteristic value information.
[0020] For example, the chemical composition may be a curable composition or a silicone polymer composition, the curable composition may be a curable silicone composition, or the chemical composition may be a non-curable silicone composition.
[0021] Alternatively, a program for implementing the functions of the control unit 11 in FIG. 1 may be recorded on a computer-readable recording medium within the teacher data creation device 1, and the program may be loaded into a computer system and executed by a processor, thereby executing the operations of the teacher data creation device 1. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. Furthermore, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into a computer system. Furthermore, the term "computer-readable recording medium" also includes devices that retain a program for a certain period of time, such as volatile memory (RAM) within a computer system that serves as a server or client when a program is transmitted via a network such as the Internet or a communication line such as a telephone line. The program may also be a program that implements part of the functions described above. Furthermore, the program may be a so-called differential file (differential program) that can implement the functions described above in combination with a program already stored in the computer system.
[0022] 1 is a recording medium such as a portable medium such as a flexible disk, a magneto-optical disk, a ROM, or a CD-ROM, a storage device such as a hard disk built into a computer system, or a RAM. Some or all of the data in the recording unit 12 may be recorded in an external device different from the teacher data creation device 1, and the control unit 11 may read the data from the external device in accordance with the operation of the control unit 11 and perform the operation.
[0023] (Manufacturing system configuration) FIG. 2 is a diagram showing the configuration of a chemical composition manufacturing system 5 according to this embodiment of the present invention. The manufacturing system 5 includes a teacher data creation device 1, a learning unit 2, an output unit 3, and a manufacturing unit 4. The learning unit 2 performs machine learning on the teacher data 122 created by the teacher data creation device 1 to create a trained model. A trained model creation device may be configured that includes the teacher data creation device 1 and the learning unit 2. The output unit 3 inputs desired property values of a desired chemical composition into the trained model created by the trained model creation device, and outputs the material formulation and proportions required to satisfy the desired property values. A chemical composition composition prediction device may be configured that includes the teacher data creation device 1, the learning unit 2, and the output unit 3. The manufacturing unit 4 manufactures the desired chemical composition based on the material formulation and proportions output by the material formulation prediction device. The chemical composition manufacturing system 5 may be configured that includes the teacher data creation device 1, the learning unit 2, the output unit 3, and the manufacturing unit 4. The output unit 3 may include a display that allows a user to view the output content. The manufacturing section 4 may be a device that performs the selection of raw materials, heating, kneading (mixing machine), rolling, pressing, curing, etc. fully automatically, or some of the manufacturing steps may be performed manually.
[0024] 3 is a diagram showing a dataset 121 according to this embodiment of the present invention. The dataset 121 includes identification names of a plurality of chemical compositions, characteristic value information indicating characteristic values of each of the plurality of chemical compositions, and raw material information about raw materials constituting each of the plurality of chemical compositions. The raw material information includes raw materials consisting of only one type of substance (single component) and raw materials consisting of two or more types of substances. The raw material information may include only raw materials consisting of two or more types of substances.
[0025] Although omitted, the dataset 121 shown in FIG. 3 includes identification names of products 1 to 400 and raw material information of raw materials 1 to 52. The identification names, characteristic value information, and raw material information in dataset 121 are information collected by a user. Dataset 121 may be the raw dataset collected by the user as is, i.e., a raw dataset. Alternatively, if there is a raw material that is not used for all the collected products in the raw dataset, dataset 121 may be a dataset from which the column for that raw material has been removed, i.e., a cleaned dataset.
[0026] For example, a chemical composition identified as "Product 1" is manufactured from "Raw Material 1" and "Raw Material 3" in a blending ratio (mass ratio) of "Raw Material 1":"Raw Material 3" = 26:74. "Product 1" has a density of 1.1 g / cm 3 The chemical composition has the following characteristics: hardness (Type A): 30, tensile strength (tensile strength at break): 6.4 MPa. The chemical composition, identified as "Product 2," is manufactured from "Raw Material 1," "Raw Material 2," and "Raw Material 3" in a blending ratio (mass ratio) of "Raw Material 1": "Raw Material 2": "Raw Material 3" = 58:33:9. "Product 2" has a density of 1.1 g / cm 3 , hardness (Type A): 31, tensile strength: 5.2 MPa. As shown in FIG. 3, products 3 to 400 also have corresponding raw material information and corresponding property value information.
[0027] A blank cell in the dataset 121 (for example, a cell with row: product 1 and column: raw material 2) indicates that raw material 2 is not used to produce product 1. The property value information may be the "elongation (elongation at break)" of the product, modulus value (MPa), or oil resistance, and is not particularly limited as long as it is information that indicates the properties of the chemical composition.
[0028] (Process flow for creating training data) 4 is a diagram showing the processing flow of the teacher data creation method according to this embodiment of the present invention. The processing from recording the data set 121 to creating the teacher data 122 will be described in order with reference to FIGS.
[0029] A user of the training data creation device 1 specifies the type of chemical composition desired by the customer. The type of chemical composition desired by the customer may be, for example, a curable composition, a curable silicone composition, a silicone polymer composition, a non-curable silicone composition, or, more specifically, silicone rubber, HCR (high-consistency rubber, millable rubber), LSR (liquid silicone rubber), or the like. The user collects information on multiple chemical compositions related to the specified type of desired chemical composition, i.e., product information. Based on the collected product information, the recording unit 12 records a dataset 121 for multiple chemical compositions related to the desired chemical composition. The dataset 121 includes identification names of the multiple chemical compositions, characteristic value information indicating characteristic values of each of the multiple chemical compositions, and raw material information on the raw materials that make up each of the multiple chemical compositions. If the collected product information contains a raw material that is not used in all of the collected products, the recording unit 12 may clean the collected product information, i.e., record a dataset excluding the column for that raw material.
[0030] The counting unit 111 counts the number of times each ingredient is used in the data set 121 based on the ingredient information (step S401). The ranking unit 112 ranks the ingredients based on the counted number of times each ingredient is used (step S402).
[0031] Referring to FIG. 5, the aggregation of the number of uses and an example of ranking will be described. FIG. 5 is a diagram showing the ranked raw materials according to the present embodiment of the present invention. The ranked raw materials in FIG. 5 are generated based on the dataset 121 in FIG. 3. The raw material 1 at the uppermost stage (rank 1) in FIG. 5 indicates that it is used in 302 products out of products 1 to 400 in the dataset 121 in FIG. 3, showing that it has the highest number of uses among products 1 to 400. The next raw material 2 indicates that it is used in 120 products out of products 1 to 400, showing that it has the second highest number of uses after raw material 1. The same applies to ranks 3 to 52.
[0032] Next, the specifying unit 113 specifies a predetermined number of raw materials from the ranked raw materials (step S403). For example, the specifying unit 113 calculates a predetermined number n that satisfies the following formula (1).
Equation
Equation
[0033] For example, assume that the user has previously set S = 0.7. At this time, S n-1 <0.7 ≤ S n to calculate n that satisfies. Specifically, the values of Sn when n = 1, 2, 3,... are calculated by the above formula (2), and the first value of n when S n becomes 0.7 or more is determined. Using the ranked raw materials shown in FIG. 5,
Equation
number
number
[0034] Next, the creation unit 114 creates training data 122 for machine learning by correlating the identified raw materials with the characteristic value information (step S404).
[0035] Next, FIG. 6 is a diagram showing training data according to this embodiment of the present invention. An example of training data to be created will be described with reference to FIG. 6. The training data in FIG. 6 is information obtained by removing information on ingredients other than the identified top 14 ingredients and their identification names from dataset 121 in FIG. 3. If there is an ingredient other than the identified top 14 ingredients that the learning unit 2 wants to learn, the ingredient may be selected by the user to remain in the training data.
[0036] (Action, effect) As described above, the teacher data creation method according to this embodiment is a teacher data creation method used in machine learning. The method also includes a step of aggregating the number of times raw materials are used in a dataset of multiple chemical compositions. The dataset includes identification names of the multiple chemical compositions, characteristic value information indicating characteristic values of each of the multiple chemical compositions, and raw material information for the raw materials constituting each of the multiple chemical compositions. The method also includes a step of ranking the raw materials based on the aggregated number of times they are used. The method also includes a step of identifying a predetermined number of raw materials from the ranked raw materials. The method also includes a step of correlating the identified raw materials with the characteristic value information to create teacher data for machine learning.
[0037] This significantly reduces the number of zeros contained in the training data for learning the causal relationships between characteristic values and raw materials, making it possible to create training data for creating an optimal trained model and improving the prediction efficiency of chemical composition formulations.
[0038] (Processing flow after creating training data) The process up to the creation of the training data has been described in detail above. Next, with reference to Fig. 7, the process from training the training unit 2 with the created training data to producing a chemical composition will be described in order.
[0039] The teacher data creation device 1 transmits the created teacher data 122 to the learning unit 2. The learning unit 2 performs machine learning on the received teacher data 122 to create a trained model (step S701). Here, the learning unit 2 performs machine learning using the raw material information of the teacher data 122 as a correct label and the characteristic value information as a predictive material.
[0040] Next, the user inputs desired property values of a desired chemical composition into the trained model of the training unit 2 (step S702). The output unit 3 outputs the ingredients and proportions required to satisfy the desired property values (step S703). The output unit 3 may present the output ingredients and proportions to the user via a display or the like provided on the output unit 3.
[0041] The output unit 3 transmits the outputted blended ingredients and their proportions to the production unit 4. The production unit 4 produces the chemical composition based on the received blended ingredients and their proportions (step S704). The above is the processing flow from the learning unit 2 learning the created teacher data 122 to the production of the chemical composition.
[0042] (Modification of this embodiment) The above describes in detail the teacher data creation device 1 according to this embodiment, but the specific aspects of the teacher data creation device 1 are not limited to those described above, and various design changes can be made within the scope of the gist.
[0043] (First Modification of the Present Embodiment) For example, as a first modification of this embodiment, the raw material information of the dataset 121 recorded by the recording unit 12 may be composed of information on raw materials consisting of only one type of substance (single component information). That is, the dataset 121 includes identification names of multiple chemical compositions, characteristic value information indicating characteristic values of each of the multiple chemical compositions, and single component information on a single component that constitutes each of the multiple chemical compositions.
[0044] (Processing flow for creating training data in the first modified example) 8 is a diagram showing the processing flow of the teacher data creation method according to the first modified example of this embodiment of the present invention. The processing from recording the data set 121 to creating the teacher data 122 will be explained in order.
[0045] A user of the training data creation device 1 specifies the type of chemical composition desired by a customer. The user collects information on multiple chemical compositions related to the specified desired type of chemical composition, i.e., product information. The recording unit 12 records a dataset 121 for multiple chemical compositions related to the desired chemical composition based on the collected product information. The dataset 121 includes identification names of the multiple chemical compositions, characteristic value information indicating characteristic values of each of the multiple chemical compositions, and single-component information about the single components that make up each of the multiple chemical compositions. If the collected product information contains a single component that is not used in all of the collected products, the recording unit 12 may clean the collected product information, i.e., record a dataset excluding the column for that single component.
[0046] The single component information of the dataset 121 in the first variant may be created by converting the raw material into only one type of substance after the dataset 121 containing the raw material consisting of two or more types of substances is recorded in step 401 of FIG. 4.
[0047] The counting unit 111 counts the number of times each single component is used in the data set based on the single component information (step S801). The ranking unit 112 ranks the single components based on the counted number of times each single component is used (step S802). Once the single components are ranked, a table in which the "raw material" in FIG. 5 is "single component" is created.
[0048] Next, the identifying unit 113 identifies a predetermined number of single components from the ranked single components (step S803). For example, the identifying unit 113 may identify the predetermined number n of single components using equations (1) and (2) in this embodiment.
[0049] Next, the creation unit 114 creates training data for machine learning by correlating the identified single component with the characteristic value information (step S804). If there is a single component other than the identified single component that the learning unit 2 wants to learn, the user may select that single component to remain in the training data.
[0050] (Action, effect) As described above, the training data creation method according to the first modification of this embodiment is a training data creation method used for machine learning. The method also includes a step of aggregating the number of times each single component is used in a dataset of multiple chemical compositions. The dataset includes identification names of the multiple chemical compositions, characteristic value information indicating characteristic values of each of the multiple chemical compositions, and single component information for each single component constituting each of the multiple chemical compositions. The method also includes a step of ranking the single components based on the aggregated number of times each component is used. The method also includes a step of identifying a predetermined number of single components from the ranked single components. The method also includes a step of correlating the identified single components with the characteristic value information to create training data for machine learning.
[0051] This significantly reduces the number of zeros contained in the training data for learning the causal relationships between characteristic values and single components, making it possible to create training data for creating an optimal trained model and improving the prediction efficiency of chemical composition formulations.
[0052] (Second Modification of the Present Embodiment) As a second modification of this embodiment, the raw material information in the dataset 121 recorded by the recording unit 12 may be composed of characteristic information about the characteristics of the raw materials. That is, the dataset 121 includes identification names of multiple chemical compositions, characteristic value information indicating the characteristic values of each of the multiple chemical compositions, and characteristic information about the characteristics of the raw materials constituting each of the multiple chemical compositions. Here, the characteristic information of the raw materials includes at least one of the vinyl content, molecular weight, viscosity, weight of silicon-bonded hydrogen atoms, type of organosiloxane units constituting the raw material, functional groups on the organosiloxane units constituting the raw material, number or ratio of organosiloxane units constituting the raw material, particle size of the filler, surface area of the filler, chemical composition of the filler, density of the filler, electrical conductivity of the filler, thermal conductivity of the filler, shape and aspect ratio of the filler, glass transition temperature, and type, amount, and bonding position of the functional groups.
[0053] Here, the organosiloxane unit constituting the raw material will be further explained. When the raw material is an organosilicon compound, particularly an organopolysiloxane compound, the raw material has a siloxane bond represented by Si—O, and RSiO 1 / 2 (wherein R is independently a group selected from a monovalent organic group, a hydroxyl group, an alkoxy group, a hydrogen atom, and a halogen atom), monofunctional siloxane units (also referred to as M units) represented by the formula RSiO 2 / 2 (wherein R is the same group as above), a difunctional siloxane unit (also referred to as a D unit) represented by the formula: RSiO 3 / 2 (wherein R is the same group as above), and a trifunctional siloxane unit (also referred to as a T unit) represented by the formula 4 / 2The raw material is composed of organosiloxane units selected from tetrafunctional siloxane units (also referred to as Q units) represented by the formula: Here, the type of organosiloxane units constituting the raw material, the functional groups (corresponding to R above) on the organosiloxane units constituting the raw material, and the number or ratio of the organosiloxane units constituting the raw material significantly affect the reactivity (including reaction rate) and properties of the reactants (including physical properties such as hardness of the cured product). Therefore, by creating training data using characteristic information including these features, it becomes possible to design chemical compositions with desired properties. The monovalent organic group represented by R may contain heteroatoms such as oxygen, nitrogen, and sulfur. Monovalent organic groups include reactive functional groups such as alkenyl, epoxy, and (meth)acryloxy groups, as well as non-reactive functional groups such as alkyl and aryl groups, and the type is not particularly limited depending on the type of raw material. Of the above characteristics, the vinyl content, molecular weight, viscosity, amount of silicon-bonded hydrogen atoms, glass transition temperature, and type and bonding position of functional groups are, of course, affected by the characteristics of the organosiloxane units that make up the raw material.
[0054] The molecular weight characteristic may be the weight average molecular weight or the number average molecular weight, or the polydispersity of the polymer, and the viscosity characteristic may be the viscosity measured by at least one of a capillary viscometer, a falling ball viscometer, and a rotational viscometer, or may be the kinematic viscosity, depending on the chemical composition.
[0055] (Processing flow for creating training data in the second modified example) 9 is a diagram showing the processing flow of a teacher data creation method according to a second modified example of this embodiment of the present invention. The processing from recording the data set 121 to creating the teacher data 122 will be described in order.
[0056] A user of the training data creation device 1 specifies the type of chemical composition desired by a customer. The user collects information on multiple chemical compositions related to the specified desired type of chemical composition, i.e., product information. The recording unit 12 records a dataset 121 for multiple chemical compositions related to the desired chemical composition based on the collected product information. The dataset 121 includes identification names of the multiple chemical compositions, characteristic value information indicating the characteristic values of each of the multiple chemical compositions, and characteristic information on characteristic points of the raw materials that make up each of the multiple chemical compositions. If unrelated characteristic points exist for all of the collected products, the recording unit 12 may clean the collected product information, i.e., record a dataset excluding the column of the unrelated characteristic point.
[0057] The feature information of the dataset 121 in the second modified example may be created by converting the ingredient information into feature points for each ingredient after the dataset 121 including the ingredient information is recorded in step 401 of FIG. 4.
[0058] The counting unit 111 counts the number of times that feature points appear in a data set based on the feature information (step S901). The ranking unit 112 ranks the feature points based on the counted number of times that feature points appear (step S902). After the feature points are ranked, a table in FIG. 5 is created in which "feature points" are used as "raw materials."
[0059] Next, the identification unit 113 identifies a predetermined number of feature points from the ranked feature points (step S903). For example, the identification unit 113 may identify the predetermined number n of feature points using equations (1) and (2) in this embodiment.
[0060] Next, the creation unit 114 creates training data for machine learning by correlating the identified feature points with the characteristic value information (step S904). If there are feature points other than the identified feature points that the learning unit 2 wants to learn, the feature points may be left in the training data by user selection.
[0061] (Action, effect) As described above, the teacher data creation method according to the second modified example of this embodiment is a teacher data creation method used for machine learning. The method also includes a step of aggregating the frequency of occurrence of feature points in a dataset of multiple chemical compositions. The dataset includes identification names of the multiple chemical compositions, characteristic value information indicating characteristic values of each of the multiple chemical compositions, and characteristic information about feature points of raw materials constituting each of the multiple chemical compositions. The method also includes a step of ranking the feature points based on the aggregated frequency of occurrence. The method also includes a step of identifying a predetermined number of feature points from the ranked feature points. The method also includes a step of correlating the identified feature points with the characteristic value information to create teacher data for machine learning.
[0062] This significantly reduces the number of zeros contained in the training data used to learn the causal relationships between characteristic values and feature points, making it possible to create training data for creating an optimal trained model and improving the prediction efficiency of chemical composition formulations.
[0063] (Configuration of the third modified example of this embodiment) 10 is a diagram showing the configuration of a teacher data creation device according to a third modified example of this embodiment of the present invention. The teacher data creation device 1 may further include a determination unit 115. The determination unit 115 determines whether the teacher data created by the teacher data creation method of this embodiment, the teacher data creation method of the first modified example, or the teacher data creation method of the second modified example is optimal. Furthermore, if the determination unit 115 determines that the teacher data created by the used method is not optimal, the determination unit 115 creates teacher data by a creation method that has not yet been used.
[0064] (Processing flow of the third modified example of this embodiment) 11 is a diagram showing a processing flow for determining optimal training data according to a third modified example of this embodiment of the present invention. The processing flow from creating training data to determining optimal training data will be described below.
[0065] The teacher data creation device 1 creates teacher data 122 using the teacher data creation method according to this embodiment. The teacher data creation device 1 transmits the created teacher data 122 to the learning unit 2. The learning unit 2 performs machine learning on the received teacher data 122 to create a trained model (step S1101). Next, the user inputs desired property values of a desired chemical composition to the trained model of the learning unit 2 (step S1102). The output unit 3 outputs the blended ingredients and proportions required to satisfy the desired property values (step S1103). The output unit 3 transmits the output blended ingredients and proportions to the production unit 4. The production unit 4 produces a chemical composition based on the received blended ingredients and proportions (step S1104). Steps S1101 to S1104 up to this point are the same as steps S701 to S704 in FIG. 7.
[0066] Next, the determination unit 115 determines whether the created teacher data is optimal. Specifically, the determination unit 115 measures the characteristic values of the produced chemical composition (step S1105). Next, the determination unit 115 determines whether the measured characteristic values match the desired characteristic values (step S1106). If the measured characteristic values match the desired characteristic values (step S1106: Yes), the determination unit 115 determines that the created teacher data is optimal (step S1107). If the measured characteristic values do not match the desired characteristic values (step S1106: No), the determination unit 115 determines that the created teacher data is not optimal, and creates teacher data using a teacher data creation method that has not yet been used, i.e., the creation method according to the first or second modified example (step S1108).
[0067] For example, if the next method to be used for creating teacher data is the method according to the first modified example, and it is determined that the teacher data created in the same manner is not optimal (step S1106: No), the teacher data is created using a method for creating teacher data that has not yet been used, i.e., the method according to the second modified example (step S1108). This process is repeated until a method for creating teacher data that is determined to be optimal is found.
[0068] If training data is created for all methods and the optimal training data creation method is not found, the training data that is considered to be the most effective may be selected. The above is the processing flow from creating training data to determining the optimal training data.
[0069] 11, whether the training data according to this embodiment is optimal is determined first, but whether the training data according to the first or second modification is optimal may be determined first. Which training data is determined first may be selected by the user, or may be automatically selected depending on the type of desired chemical composition, for example, whether it is millable rubber.
[0070] (Action, effect) As described above, the teacher data creation method according to the third modification of this embodiment includes a step of determining whether the teacher data created by the teacher data creation method according to this embodiment, the first modification, or the second modification is optimal. If the teacher data is determined to be not optimal, the method includes a step of creating teacher data by a creation method not previously used.
[0071] This allows optimal training data to be selected from multiple training data created by multiple methods, further improving the efficiency of predicting the formulation of chemical compositions.
[0072] (Configuration of the fourth modified example of this embodiment) FIG. 12 is a diagram showing the configuration of a teacher data creation device according to a fourth modified example of this embodiment of the present invention. The teacher data creation device 1 may further include a determination unit 115. The determination unit 115 determines whether the teacher data created by the teacher data creation method of this embodiment, the teacher data creation method of the first modified example, or the teacher data creation method of the second modified example is optimal. Furthermore, if the determination unit 115 determines that the teacher data created by the previously used method is not optimal, it creates teacher data by a creation method that has not yet been used. Furthermore, the recording unit 12 includes verification data 123 for verifying the usefulness of the teacher data 122. The verification data 123 is data independent of the dataset 121 and the teacher data 122. The verification data 123 includes characteristic values obtained by measuring a chemical composition previously manufactured, as well as the ingredients and proportions of the chemical composition.
[0073] (Processing flow of the fourth modified example of this embodiment) 13 is a diagram showing a processing flow for determining optimal training data according to a fourth modified example of the present embodiment of the present invention. The processing flow from creating training data to determining optimal training data will be described below.
[0074] The teacher data creation device 1 creates teacher data 122 using the teacher data creation method according to this embodiment. The teacher data creation device 1 transmits the created teacher data 122 to the learning unit 2. The learning unit 2 performs machine learning on the received teacher data 122 to create a trained model (step S1301). Next, the manufacturing system 5 inputs the property values in the verification data 123 into the trained model of the learning unit 2 as desired property values of a desired chemical composition (step S1302). The output unit 3 outputs the blended materials and proportions required to satisfy the desired property values (step S1303).
[0075] Next, the determination unit 115 determines whether the output ingredients and proportions match the ingredients and proportions corresponding to the input characteristic values in the verification data 123 (step S1304). If the output ingredients and proportions match the ingredients and proportions in the verification data 123 (step S1304: Yes), the determination unit 115 determines that the created training data is optimal (step S1305). If the output ingredients and proportions do not match the ingredients and proportions in the verification data 123 (step S1304: No), the determination unit 115 determines that the created training data is not optimal, and creates training data using a training data creation method that has not yet been used, i.e., the first or second modified example (step S1306).
[0076] For example, if the next method to be used for creating teacher data is the method according to the first modified example, and it is determined that the teacher data created in the same manner is not optimal (step S1304: No), the teacher data is created using a method for creating teacher data that has not yet been used, i.e., the method according to the second modified example (step S1306). This process is repeated until a method for creating teacher data that is determined to be optimal is found.
[0077] If training data is created for all methods and the optimal training data creation method is not found, the training data that is considered to be the most effective may be selected. The above is the processing flow from creating training data to determining the optimal training data.
[0078] In step S1304 of FIG. 13, the determination unit 115 determined whether the outputted ingredients and proportions match the ingredients and proportions of the verification data 123. Here, the determination unit 115 may determine whether there is a match by evaluating the degree of match between the ingredients and proportions of the outputted ingredients and proportions and the ingredients and proportions of the verification data 123. For example, the determination unit 115 may numerically evaluate the degree of deviation between the ingredients and proportions of the outputted ingredients and proportions and the ingredients and proportions of the verification data 123, and determine that there is a match when the evaluation value is less than a predetermined threshold.
[0079] 13, whether the training data according to this embodiment is optimal is determined first, but whether the training data according to the first or second modification is optimal may be determined first. Which training data is determined first may be selected by the user, or may be automatically selected depending on the type of desired chemical composition, for example, whether it is millable rubber.
[0080] (Action, effect) As described above, the teacher data creation method according to the fourth modification of this embodiment includes a step of determining whether the teacher data created by the teacher data creation method according to this embodiment, the first modification, or the second modification is optimal. The determining step includes a step of machine learning the created teacher data in a learning unit of a computer, and a step of inputting characteristic values in the verification data to the learning unit and outputting the ingredients and their proportions. The determining step also includes a step of determining whether the output ingredients and their proportions match the ingredients and their proportions in the verification data corresponding to the input characteristic values. The determining step also includes a step of determining that the created teacher data is optimal if they match, and a step of determining that the created teacher data is not optimal if they do not match.
[0081] This allows optimal training data to be selected from multiple training data created by multiple methods without manufacturing, further improving the efficiency of predicting the formulation of chemical compositions.
[0082] Although the present embodiment has been described above, the present embodiment and its modifications are presented as examples and are not intended to limit the scope of the invention. The present embodiment and its modifications can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. The present embodiment and its modifications are included in the scope of the invention and its equivalents as defined in the claims, as well as in the scope and spirit of the invention. [Explanation of symbols]
[0083] 1. Teacher data creation device 11 Control section 111 Counting Department 112 Ranking Department 113 Specific section 114 Creation Department 115 Judgment section 12 Recording section 121 datasets 122 Teacher Data 123 Verification Data 2. Learning Department 3 Output section 4 Manufacturing Department 5 Manufacturing System
Claims
1. A method for creating training data for use in machine learning by a computer, comprising: a step of aggregating the number of times raw materials are used in a dataset of a plurality of chemical compositions, the dataset including identification names of the plurality of chemical compositions, characteristic value information indicating characteristic values of each of the plurality of chemical compositions, and raw material information on raw materials constituting each of the plurality of chemical compositions; ranking the raw materials based on the counted number of uses; identifying a predetermined number of ingredients from the ranked ingredients; a step of correlating the identified raw materials with the characteristic value information to create training data for machine learning; A method for creating training data, including:
2. The step of identifying a predetermined number of raw materials from the ranked raw materials includes: A step of calculating a predetermined number n that satisfies the following formula (1): [Equation 1] S is a preset value that satisfies 0<S<1, S n is expressed by the following formula (2): [Equation 2] N k indicates the number of times the raw material is used, last indicates the total number of raw material types in the dataset, and n is a natural number; identifying a predetermined number n of top-ranked raw materials from the ranked raw materials; The teacher data creation method according to claim 1 , further comprising:
3. A method for creating training data for use in machine learning by a computer, comprising: a step of aggregating the number of times a single component is used in a dataset of a plurality of chemical compositions, the dataset including identification names of the plurality of chemical compositions, characteristic value information indicating characteristic values of each of the plurality of chemical compositions, and single component information for each of the single components constituting each of the plurality of chemical compositions; ranking the single components based on the aggregated number of uses; identifying a predetermined number of single components from the ranked single components; a step of correlating the identified single component with the characteristic value information to create training data for machine learning; A method for creating training data, including:
4. A method for creating training data for use in machine learning by a computer, comprising: a step of counting the number of occurrences of feature points in a dataset of a plurality of chemical compositions, the dataset including identification names of the plurality of chemical compositions, feature value information indicating feature values of each of the plurality of chemical compositions, and feature information on feature points of raw materials constituting each of the plurality of chemical compositions; ranking the feature points based on the counted number of occurrences; identifying a predetermined number of feature points from the ranked feature points; a step of correlating the identified feature points with the characteristic value information to create training data for machine learning; A method for creating training data, including:
5. The training data creation method of claim 4, wherein the characteristic points include at least one of the vinyl content, molecular weight, viscosity, weight of silicon-bonded hydrogen atoms, type of organosiloxane units constituting the raw material, functional groups on the organosiloxane units constituting the raw material, number or ratio of organosiloxane units constituting the raw material, particle size of the filler, surface area of the filler, chemical composition of the filler, density of the filler, electrical conductivity of the filler, thermal conductivity of the filler, shape and aspect ratio of the filler, glass transition temperature, and type, amount and bonding position of the functional groups.
6. The explanatory variables of the teacher data are characteristic value information, The objective variable of the training data is a raw material, a single component, or a feature point, a step of having a learning unit perform machine learning on the created teacher data; inputting desired characteristic values of a desired chemical composition into the learning unit, and outputting ingredients and proportions to satisfy the desired characteristic values; causing the device to produce a chemical composition based on the outputted formulation ingredients and proportions; measuring a characteristic value of the produced chemical composition using an apparatus and determining whether the measured characteristic value matches the desired characteristic value; If there is a match, determining that the created training data is optimal; If they do not match, determining that the created training data is not optimal; The teacher data creation method according to any one of claims 1 to 5, further comprising:
7. a step of creating teacher data by the teacher data creation method according to claim 1, 3, or 4, in which explanatory variables are characteristic value information and objective variables are raw materials, single components, or feature points; a step of causing a learning unit to perform machine learning on the created teacher data; inputting desired characteristic values of a desired chemical composition into the learning unit, and outputting ingredients and proportions to satisfy the desired characteristic values; causing the device to produce a chemical composition based on the outputted formulation ingredients and proportions; measuring a characteristic value of the produced chemical composition using an apparatus and determining whether the measured characteristic value matches the desired characteristic value; If there is a match, determining that the created training data is optimal; If they do not match, determining that the created training data is not optimal; If it is determined that the data is not optimal, a step of creating teacher data by the teacher data creating method of claim 1, 3, or 4 that has not been used before; A method for creating training data, including:
8. The explanatory variables of the teacher data are characteristic value information, The objective variable of the training data is a raw material, a single component, or a feature point, a step of causing a learning unit to perform machine learning on the created teacher data; inputting the characteristic values in the verification data to the learning unit and outputting the ingredients and their proportions; a step of determining whether the outputted ingredients and proportions match the ingredients and proportions in the verification data corresponding to the input characteristic values; If there is a match, determining that the created training data is optimal; If they do not match, determining that the created training data is not optimal; and The teacher data creation method according to any one of claims 1 to 5, further comprising:
9. a step of creating teacher data by the teacher data creation method of claim 1, 3, or 4, in which explanatory variables are characteristic value information and objective variables are raw materials, single components, or feature points; a step of causing a learning unit to perform machine learning on the created teacher data; inputting the characteristic values in the verification data to the learning unit and outputting the ingredients and their proportions; a step of determining whether the outputted ingredients and proportions match the ingredients and proportions in the verification data corresponding to the input characteristic values; If there is a match, determining that the created training data is optimal; If they do not match, determining that the created training data is not optimal; If it is determined that the data is not optimal, a step of creating teacher data by the teacher data creating method of claim 1, 3, or 4 that has not been used before; A method for creating training data, including:
10. The training data creation method according to any one of claims 1 to 9, wherein the chemical composition is a curable composition or a silicone polymer composition.
11. The training data generating method according to claim 10 , wherein the curable composition is a curable silicone composition.
12. The training data creation method according to any one of claims 1 to 9, wherein the chemical composition is a non-curable silicone composition.
13. A step of creating teacher data by the teacher data creation method according to any one of claims 1 to 12, wherein an explanatory variable is characteristic value information and a target variable is a raw material, a single component, or a feature point; A trained model creation method comprising a step of performing machine learning on the training data by a computer to create a trained model.
14. A step of creating a trained model by the trained model creation method according to claim 13; A method for predicting the formulation of chemical compositions, comprising the steps of inputting desired property values of a desired chemical composition into the trained model by a computer, and outputting formulation materials and proportions that will satisfy the desired property values.
15. A teacher data creation device used in machine learning, a counting unit that counts the number of times raw materials are used in a data set of a plurality of chemical compositions, the data set including identification names of the plurality of chemical compositions, characteristic value information indicating characteristic values of each of the plurality of chemical compositions, and raw material information about raw materials that constitute each of the plurality of chemical compositions; a ranking unit that ranks the raw materials based on the counted number of times of use; an identification unit that identifies a predetermined number of raw materials from the ranked raw materials; a creation unit that creates training data for machine learning by correlating the specified raw material with the characteristic value information; A teacher data creation device including:
16. The identification unit A predetermined number n that satisfies the following formula (1) is calculated, [Equation 3] S is a preset value that satisfies 0<S<1, and S n is expressed by the following formula (2): [Equation 4] N k indicates the number of times the raw material is used, last indicates the total number of types of raw materials in the dataset, and n is a natural number. Identifying a predetermined number n of top-ranked raw materials from the ranked raw materials; The teacher data creation device according to claim 15.
17. A teacher data creation device used in machine learning, a counting unit that counts the number of times a single component is used in a data set of a plurality of chemical compositions, the data set including identification names of the plurality of chemical compositions, characteristic value information indicating characteristic values of each of the plurality of chemical compositions, and single component information about a single component that constitutes each of the plurality of chemical compositions; a ranking unit that ranks the single components based on the counted number of times of use; an identifying unit that identifies a predetermined number of single components from the ranked single components; a creation unit that creates training data for machine learning by correlating the identified single component with the characteristic value information; A teacher data creation device including:
18. A teacher data creation device used in machine learning, a counting unit that counts the number of occurrences of feature points in a dataset of a plurality of chemical compositions, the dataset including identification names of the plurality of chemical compositions, feature value information indicating feature values of each of the plurality of chemical compositions, and feature information on feature points of raw materials that constitute each of the plurality of chemical compositions; a ranking unit that ranks the feature points based on the counted number of occurrences; an identifying unit that identifies a predetermined number of feature points from the ranked feature points; a creation unit that creates training data for machine learning by correlating the specified feature points with the characteristic value information; A teacher data creation device including:
19. The training data creation device of claim 18, wherein the characteristic points include at least one of vinyl content, molecular weight, viscosity, weight of silicon-bonded hydrogen atoms, type of organosiloxane units constituting the raw material, functional groups on the organosiloxane units constituting the raw material, number or ratio of organosiloxane units constituting the raw material, particle size of the filler, surface area of the filler, chemical composition of the filler, density of the filler, electrical conductivity of the filler, thermal conductivity of the filler, shape and aspect ratio of the filler, glass transition temperature, and type, amount and bonding position of the functional groups.
20. The explanatory variables of the teacher data are characteristic value information, The objective variable of the training data is a raw material, a single component, or a feature point, A determination unit, The created teacher data is subjected to machine learning by a learning unit of a computer, inputting desired characteristic values of a desired chemical composition into the learning unit, and outputting blending materials and proportions to satisfy the desired characteristic values; causing the device to produce a chemical composition based on the outputted ingredients and proportions; causing an apparatus to measure a characteristic value of the produced chemical composition and determining whether the measured characteristic value matches the desired characteristic value; If they match, it is determined that the created training data is optimal; The teacher data creation device according to any one of claims 15 to 19, further comprising a judgment unit configured to determine that the created teacher data is not optimal if there is no match.
21. The explanatory variables of the teacher data are characteristic value information, The objective variable of the training data is a raw material, a single component, or a feature point, A determination unit, The created teacher data is subjected to machine learning by a learning unit of a computer, Inputting the characteristic values in the verification data into the learning unit and outputting the blended materials and their proportions; determining whether the outputted ingredients and proportions match the ingredients and proportions in the verification data corresponding to the input characteristic values; If they match, it is determined that the created training data is optimal; The teacher data creation device according to any one of claims 15 to 19, further comprising a judgment unit configured to determine that the created teacher data is not optimal if there is no match.
22. The training data creation device according to any one of claims 15 to 21, wherein the chemical composition is a curable composition or a silicone polymer composition.
23. The training data creation device according to claim 22, wherein the curable composition is a curable silicone composition.
24. The teacher data creation device according to any one of claims 15 to 21, wherein the chemical composition is a non-curable silicone composition.
25. The teacher data creation device according to any one of claims 15 to 24, a learning unit that performs machine learning on training data created by the training data creation device, the training data including explanatory variables that are characteristic value information and objective variables that are raw materials, single components, or feature points, to create a trained model; A trained model creation device including:
26. The trained model creation device according to claim 25; an output unit that inputs desired characteristic values of a desired chemical composition into the trained model created by the trained model creation device and outputs blended materials and proportions to satisfy the desired characteristic values; A chemical composition formulation prediction device comprising:
Citation Information
Patent Citations
Simulation system of design / Combination
JP2003058582A
Method and device for predicting physical property data
JP2020038495A
Method and device for selecting informative features
KR1020190136969A
Systems and methods for estimating a condition from sensor data using random forest classification
US20190137539A1
Learning method, learning device, and learning program
WO2019225228A1