Information processing device, information processing method, and information processing program

By using an index to assess data contribution, the approach optimizes data collection in materials informatics, addressing the challenge of variable scarcity/excess and enhancing predictive accuracy in high-performance materials without data disclosure.

JP2025156154APending Publication Date: 2025-10-14ASAHI KASEI KOGYO KABUSHIKI KAISHA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025052761
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2025-03-27
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

The challenge in materials informatics is the lack or excess of explanatory variables required to predict physical properties in high-performance materials, particularly when data is collected across multiple processes by different entities, leading to unclear mechanisms of action and difficulty in understanding the essential factors for achieving target performance.

Method used

An index called the degree of contribution of experimental data is introduced to assess the importance of existing data, facilitating data collection plans and using a recommendation unit, contribution calculation unit, and control unit to optimize data collection based on machine learning models.

Benefits of technology

This approach allows for targeted data collection that enhances the understanding of data importance, improving the prediction of physical properties in materials informatics by generating accurate inference models without disclosing sensitive data across participants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025156154000001_ABST
    Figure 2025156154000001_ABST
Patent Text Reader

Abstract

To provide an information processing device, information processing method, and information processing program.SOLUTION: An information processing device 100 includes a recommendation unit and a contribution calculation unit. The recommendation unit has a first inference model that performs machine learning on first teacher data including teacher explanatory variables and teacher objective variables provided by each of a plurality of participants, to output, in response to input of an input explanatory variable, a first predicted value of an input objective variable corresponding to the input explanatory variable, and output, in response to input of a target value of the input objective variable, optimization data that is an explanatory variable corresponding to the target value. The contribution calculation unit calculates, based on the optimization data and the first teacher data, a contribution of the first teacher data corresponding to at least one participant to the optimization data.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]

[0002] Non-Patent Document 1 states, "Federated learning is an emerging distributed machine learning framework for privacy preservation." (Abstract) [Prior art document] [Non-patent literature] [Non-Patent Document 1] Federated Learning on Non-IID Data Summary of the Invention

[0003] When formulating a plan to collect experimental data for materials informatics, one of the challenges can be the lack or excess of explanatory variables required to predict physical properties. For high-performance materials manufactured through complex, multi-step processes, even the specialized knowledge of experienced personnel leaves the mechanism of action unclear. Consequently, it is impossible to identify the factors essential for achieving the target performance. This can lead to the challenges described above. When collecting data across multiple processes, such as material manufacturing, preparation, and processing, different data collectors may be responsible for each process. In such cases, it becomes difficult to fully understand the data across multiple processes, making the challenges described above particularly pronounced.

[0004] To prevent data collection overload or underload under conditions where it is difficult for each entity to grasp the entire experimental data, one embodiment of the present invention introduces an index called the degree of contribution of experimental data to each predicted data. The degree of contribution in this invention is an index that indicates the degree of importance of existing experimental data for, for example, recommending experimental conditions by computer using materials informatics.

[0005] Furthermore, in one embodiment of the present invention, a plan for collecting additional experimental data is facilitated based on the contribution of existing experimental data to a certain target variable, and data is collected from an experimental device based on the plan. The information processing device of the present invention may include at least one of a "recommendation unit," a "contribution calculation unit," a "planning unit," and a "control unit." The "recommendation unit" may have a first inference model, and when a target value of the objective variable is input, may output optimization data which is an explanatory variable corresponding to the target value. The "contribution calculation unit" may calculate the contribution of at least one experimental data included in the first training data to the optimization data based on the optimization data and the first training data (training data used for machine learning of the first inference model). The "planning unit" may develop and output additional experimental data collection plans based on the contributions. The "control unit" may acquire experimental data by controlling the experimental equipment based on the data collection plan.

[0006] A first aspect of the present invention provides an information processing device. The information processing device includes a recommendation unit having a first inference model that, when an input explanatory variable is input, outputs a first predicted value of an input objective variable corresponding to the input explanatory variable by machine learning first teacher data having teacher explanatory variables and teacher objective variables provided by each of a plurality of participants. When a target value of the input objective variable is input to the recommendation unit, the recommendation unit outputs optimization data that is an explanatory variable corresponding to the target value. The information processing device further includes a contribution calculation unit that calculates a contribution of the first teacher data corresponding to at least one participant to the optimization data based on the optimization data and the first teacher data.

[0007] When a target value is input, the recommendation unit may generate candidates for optimized data corresponding to the target value. The first inference model may output a first predicted value corresponding to the generated candidates. The recommendation unit may output the optimized data based on the first predicted value and the target value.

[0008] In any of the information processing devices described above, the recommendation unit may have a second inference model that outputs a second predicted value of the input objective variable when an input explanatory variable is input by machine learning second teacher data obtained by excluding, from the first teacher data, individual explanatory variables that are teacher explanatory variables corresponding to at least one of the multiple participants and individual objective variables that are teacher objective variables corresponding to at least one of the multiple participants. The contribution calculation unit may calculate the contribution based on the first predicted value and the second predicted value.

[0009] In any of the above information processing devices, the first inference model may output a first predicted value corresponding to the optimization data when the optimization data is input. The second inference model may output a second predicted value corresponding to the optimization data when the optimization data is input. The contribution calculation unit may calculate the contribution based on the first predicted value corresponding to the optimization data and the second predicted value corresponding to the optimization data.

[0010] In any of the above information processing devices, the recommendation unit may perform machine learning on second training data for each of the multiple participants, excluding individual explanatory variables and individual objective variables corresponding to one of the multiple participants. The second inference model may output a second predicted value for each of the multiple participants. The contribution calculation unit may calculate a contribution for each of the multiple participants based on the first predicted value and the second predicted value for each of the multiple participants. The information processing device may further include a selection unit that selects at least one participant based on the contribution, and prompts the recommendation unit to input additional individual training data, which is individual training data having individual explanatory variables and individual objective variables.

[0011] In any of the information processing devices described above, the recommendation unit may perform machine learning on second teacher data from which individual explanatory variables and individual objective variables corresponding to at least two of the plurality of participants have been excluded, and the contribution calculation unit may calculate contributions of the at least two participants based on the first predicted value and the second predicted value.

[0012] In any of the information processing devices described above, weights may be assigned in advance to the plurality of objective variables, and the contribution calculation unit may calculate the contribution based on the first predicted value, the second predicted value, and the weights.

[0013] In any of the information processing devices described above, the teacher explanatory variables may include a plurality of types of parameters. At least one type of parameter may be provided by at least one participant. The contribution calculation unit may calculate a contribution of the at least one parameter to the optimization data based on the optimization data and the first teacher data.

[0014] In any of the information processing devices described above, a plurality of types of parameters may be provided by at least one participant, and the contribution calculation unit may calculate a contribution of the at least one type to the optimization data for each of the plurality of types, and calculate a contribution of the at least one participant based on the calculated contribution of each of the plurality of types.

[0015] A second aspect of the present invention provides an information processing method. The information processing method includes a first inference model generation step in which a recommendation unit generates a first inference model that outputs a first predicted value, which is a predicted value of an input objective variable corresponding to an input explanatory variable, by machine learning first teacher data having teacher explanatory variables and teacher objective variables provided by each of a plurality of participants, when an input explanatory variable is input. When a target value of the input objective variable is input to the recommendation unit, the recommendation unit outputs optimization data, which is an explanatory variable corresponding to the target value. The information processing method further includes a contribution calculation step in which a contribution calculation unit calculates a contribution of the first teacher data corresponding to at least one participant to the optimization data based on the optimization data and the first teacher data.

[0016] In a third aspect of the present invention, there is provided an information processing program that causes a computer to function as an information processing device.

[0017] The above summary of the invention does not list all of the features of the present invention, and subcombinations of these features may also be inventions. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 is a diagram illustrating an example of an information providing system 500. [Figure 2] FIG. 2 is a diagram showing an example of training data 201 provided by a participant 200. [Figure 3] 10 is a diagram showing an example of first teacher data 210 provided by each of a plurality of participants 200. FIG. [Figure 4] 1 is a diagram illustrating an example of an information processing device 100 according to an embodiment of the present invention. [Figure 5] FIG. 2 is a diagram illustrating an example of a data space 212. [Figure 6] FIG. 2 is another diagram showing an example of the data space 212. [Figure 7] FIG. 10 is a diagram illustrating an example of learning by the recommendation unit 20. [Figure 8] FIG. 10 is a diagram showing another example of learning by the recommendation unit 20. [Figure 9] A figure showing another example of inference using the second inference model 24. [Figure 10] A figure showing another example of inference using the second inference model 24. [Figure 11] FIG. 10 is a diagram showing another example of the first teacher data 210 provided by each of the multiple participants 200. [Figure 12] 1 is a flowchart illustrating an example of an information processing method according to an embodiment of the present invention. [Figure 13] 1 is a flowchart illustrating an example of an information processing method according to an embodiment of the present invention. [Figure 14] 1 is a flowchart illustrating an example of an information processing method according to an embodiment of the present invention. [Figure 15] 1 is a flowchart illustrating an example of an information processing method according to an embodiment of the present invention. [Figure 16] 1 is a flowchart illustrating an example of an information processing method according to an embodiment of the present invention. [Figure 17] 1 is a flowchart illustrating an example of an information processing method according to an embodiment of the present invention. [Figure 18] FIG. 2 is a diagram illustrating another exemplary configuration of the information processing device 100. [Figure 19] An example of information on the contribution of each explanatory variable calculated by the contribution calculation section 60 is shown below. [Figure 20] Another example of information about the contribution of each explanatory variable calculated by the contribution calculation section 60 will be shown below. [Figure 21] 10 is a diagram illustrating an example of the operation of the plan formulation unit 72. FIG. [Figure 22] 10 is a diagram illustrating another example of the operation of the plan formulation unit 72. FIG. [Figure 23] FIG. 2 is a diagram illustrating another exemplary configuration of the information processing device 100. [Figure 24] FIG. 22 is a diagram showing an example of a computer 2200 in which an information processing device 100 according to an embodiment of the present invention may be implemented in whole or in part. DETAILED DESCRIPTION OF THE INVENTION

[0019] The present invention will be described below through embodiments of the invention, but the following embodiments do not limit the scope of the invention according to the claims. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.

[0020] FIG. 1 is a diagram showing an example of an information providing system 500. The information providing system 500 includes an information processing device 100, a network 400, and a plurality of terminals 300. The information processing device 100 in this example is a server that provides services provided by the information providing system 500. The terminals 300 in this example are terminals of participants 200 who receive services provided by the information providing system 500. In this example, participants 200-1 to 200-n (n is an integer of 2 or greater) have terminals 300-1 to 300-n, respectively. The network 400 may be a general-purpose network such as the Internet, or may be a network using dedicated lines.

[0021] The terminal 300 transmits the individual teacher data 202 provided by the participant 200 to the information processing device 100 via the network 400. The information processing device 100 receives the individual teacher data 202 transmitted from each of the multiple terminals 300 and stores the received individual teacher data 202. One individual teacher data 202 provided by one participant 200 and stored in the information processing device 100 may be kept secret from the other participants 200.

[0022] The information processing device 100 may generate an estimation model 21 using individual training data 202 transmitted from each of the multiple terminals 300. The estimation model 21 may be a regression model such as an arithmetic expression generated by regression analysis or the like, or may be a model generated by machine learning.

[0023] For example, the individual training data 202 includes explanatory variables that indicate the state of a material or the like before some processing is performed, and a response variable that indicates the state of a product or the like after the processing is performed. The estimation model 21 is a model that indicates the relationship between the explanatory variables and the response variable. The estimation model 21 is a model that can estimate the corresponding response variable when an explanatory variable is input.

[0024] The information processing device 100 provides services to each terminal 300 using the estimation model 21. When explanatory variables are input to one terminal 300 and the information processing device 100 receives the explanatory variables transmitted by the one terminal 300, the information processing device 100 may generate an objective variable corresponding to the explanatory variables using the estimation model 21. The information processing device 100 may transmit the generated objective variable to the one terminal 300 to which the explanatory variables were input. The information processing device 100 may conceal the estimation model 21 and the content of the processing by the estimation model 21 from the other terminals 300.

[0025] The information processing device 100 of this example can generate a highly accurate inference model by generating an inference model using individual training data 202 from multiple participants 200. Each participant 200 can receive services based on a highly accurate estimation model 21 without disclosing the individual training data 202 that he or she provided to other participants 200.

[0026] FIG. 2 is a diagram illustrating an example of individual teacher data 202 provided by a participant 200. The teacher data 201 includes teacher explanatory variables 204 and teacher objective variables 206 corresponding to the teacher explanatory variables 204. The teacher explanatory variables 204 may include multiple types of parameters. In this example, the teacher explanatory variables 204 include one or more physical quantities. In this example, the teacher objective variables 206 are physical quantities realized using the physical quantities indicated by the teacher explanatory variables 204. For example, the teacher objective variables 206 are physical property values ​​of an object to be manufactured. In this example, the teacher explanatory variables 204 include at least one of a parameter indicating the material of the object to be manufactured and a parameter indicating the processing conditions in a manufacturing process using the material. The teacher explanatory variables 204 shown in FIG. 2 include parameters related to a reaction process R and a molding process M. In this example, the parameters related to the reaction process R include a composition Co and a process condition P1.

[0027] The composition Co may include the types and composition ratios of one or more raw materials. In the example of FIG. 2, the composition Co includes the types and composition ratios of the raw materials of the main material Sm, the secondary material Sb, and the additive Sa. The numerical values ​​shown in FIG. 2 are an example of the composition ratios. The process conditions P1 may include reaction conditions for the material having the composition Co. In the example of FIG. 2, the process conditions P1 include a process temperature and a process time in the reaction treatment of the material having the composition Co. In this example, the molding process M includes process conditions P2. The process conditions P2 may include a process temperature and a process pressure in the process of molding the substance produced in the reaction process R into a predetermined shape.

[0028] The teacher explanatory variable 204 may have evaluation conditions Ev in the evaluation process E. The evaluation process E in this example is a process for evaluating the physical property values ​​of the molded product molded in the molding process M. The evaluation conditions Ev are conditions for evaluating the molded product molded in the molding process M in the evaluation process E. The evaluation conditions Ev may include an evaluation temperature and an evaluation humidity, which are the temperature and humidity of the measurement environment in which the physical property values ​​of the molded product are measured.

[0029] The teacher objective variable 206 is a parameter related to the evaluation result when the molded product molded in the molding process M is evaluated under the evaluation condition Ev. In this example, the teacher objective variable 206 includes the evaluation results of the hardness and elongation of the molded product molded in the molding process M. The teacher data 201 in this example is actual data obtained when the participant 200 actually creates a molded product. At least a part of the teacher data 201 may be data obtained by calculation such as a simulation.

[0030] FIG. 3 is a diagram showing an example of first teacher data 210 provided by each of multiple participants 200. The multiple participants 200 are, for example, multiple companies belonging to the same industry. Individual teacher data 202 is teacher data 201 (see FIG. 1) provided by one participant 200. In this example, individual teacher data 202-1 to individual teacher data 202-n (n is an integer equal to or greater than 2) are provided by participants 200-1 to 200-n, respectively.

[0031] The individual teacher data 202 has an individual explanatory variable 203 and an individual objective variable 205. The individual explanatory variable 203 is a teacher explanatory variable 204 corresponding to at least one of the multiple participants 200. The individual objective variable 205 is a teacher objective variable 206 corresponding to at least one of the multiple participants 200. The individual teacher data 202-u includes an individual explanatory variable 203-u and an individual objective variable 205-u (u is an integer from 1 to n). The first teacher data 210 is data obtained by aggregating the individual teacher data 202-1 to 202-n. The individual teacher data 202-1 to 202-n may have the same parameters.

[0032] The first teacher data 210 may include an identification parameter Id for identifying multiple participants 200. In this example, the identification parameter Id includes the name of the data provider and a data number. In this example, the individual teacher data 202-1 to 202-n provided by participants 200-1 to 200-n are labeled "Company A" to "Company N," respectively.

[0033] One data number may correspond to one set of teacher explanatory variables 204 and teacher objective variables 206. In this example, the three sets included in the individual teacher data 202-1 are assigned data numbers "A1" to "A3", the three sets included in the individual teacher data 202-2 are assigned data numbers "B1" to "B3", and the three sets included in the individual teacher data 202-n are assigned data numbers "N1" to "N3".

[0034] 4 is a diagram showing an example of an information processing device 100 according to an embodiment of the present invention. The information processing device 100 includes a recommendation unit 20 and a contribution calculation unit 60. The information processing device 100 may include an identification unit 10, an input unit 30, an output unit 40, a storage unit 50, and a selection unit 70.

[0035] A part or the whole of the information processing device 100 may be realized by a computer. The recommendation unit 20 may be a CPU (Central Processing Unit) of the computer. When the information processing device 100 is realized by a computer, a program for causing the computer to function as the information processing device 100 may be installed on the computer, and a program for causing the computer to execute an information processing method described below may be installed on the computer.

[0036] The input unit 30 receives input of a teacher explanatory variable 204 and a teacher objective variable 206. The input unit 30 may communicate with the terminals 300 of the participants 200 via a network 400. The input unit 30 may receive individual teacher data 202 from the terminals 300 of the participants 200. The input unit 30 may read the individual teacher data 202 from a storage unit 50 that stores the individual teacher data 202. The input unit 30 may have a mouse, keyboard, etc. for inputting the individual teacher data 202. The storage unit 50 may store at least one of the individual teacher data 202 and first teacher data 210.

[0037] The recommendation unit 20 generates optimization data Dt, which are explanatory variables corresponding to the target value Vt of the input objective variable, based on the target value Vt of the input objective variable input from the terminal 300 of one of the participants 200. For example, the target value Vt of the input objective variable is a physical property value that a newly generated objective object should have. The recommendation unit 20 estimates a combination of explanatory variable values ​​that can achieve the target value Vt of the input objective variable, using the individual training data 202 stored in the storage unit 50. The recommendation unit 20 may generate the optimization data Dt by inputting the target value Vt of the input objective variable into the estimation model 21. The output unit 40 outputs the generated optimization data Dt to the terminal 300 of the participant 200 that provided the target value Vt of the input objective variable. The output unit 40 may communicate with the terminal 300 of the participant 200 via the network 400.

[0038] The identification unit 10 identifies a data space 212 (described below) for each of the multiple participants 200 based on the individual teacher data 202 provided by each of the multiple participants 200. The data space 212 (described below) for each participant 200 is a portion of a data space 214 (described below) for all participants. The data space 212 is a virtual space. As an example, the data space 212 (described below) is a multidimensional space having axes corresponding to each parameter included in the individual teacher data 202. Each individual teacher data 202 is placed at a position in the data space 214 (described below) for all participants depending on the value of each parameter. The data space 212 (described below) for each participant 200 may include all of the individual teacher data 202 provided by that participant 200. The data space 212 (described below) for each participant 200 may also include the area between any two individual teacher data 202 among the multiple individual teacher data 202 provided by that participant 200. In addition, the identification unit 10 may refer to the individual teacher data 202 provided by each of the multiple participants 200 and identify which participant 200 provided each individual teacher data 202.

[0039] The recommendation unit 20 generates optimized data Dt corresponding to the target value Vt of the input objective variable based on the target value Vt and data space information Is, which is information indicating the data space. For example, the recommendation unit 20 may generate the optimized data Dt based on the data space information Is of the participant 200 who input the target value Vt of the input objective variable. More specifically, the recommendation unit 20 may generate optimized data Dt that is not included in the data space 212 (described below) of the participant 200. If the recommendation unit 200 generates optimized data Dt that is not included in the data space 212 (described below) of the participant 200, the participant 200 will be more likely to obtain optimized data Dt that is outside the range of experiments that their company has already conducted.

[0040] FIG. 5 is a diagram showing an example of a data space 212. For ease of explanation, the case where n in FIG. 3 is 2 will be described. The identification unit 10 may identify the data space 212 of each of the multiple participants 200 based on the individual teacher data 202. In this example, the identification unit 10 identifies the data space 212-1 of the participant 200-1 (Company A in FIG. 3) based on the individual teacher data 202-1, and identifies the data space 212-2 of the participant 200-2 (Company B in FIG. 3) based on the individual teacher data 202-2. In FIG. 5, the data space 212-1 is indicated by a fine dashed line, and the data space 212-2 is indicated by a coarse dashed line. The identification unit 10 may identify the data space 212 based on the teacher explanatory variables 204, or may identify the data space 212 based on the teacher explanatory variables 204 and the teacher objective variable 206.

[0041] The identification unit 10 may identify a data space 214 based on all data of individual teacher data 202-1 to individual teacher data 202-n. For example, the data space 214 of all participants is a space having axes corresponding to all parameters appearing in at least one individual teacher data 202. When there are two individual teacher data 202 (individual teacher data 202-1 and individual teacher data 202-2), the data space 214 is the outer frame (thick frame in FIG. 5) of the union of data space 212-1 (fine dashed frame in FIG. 5) and data space 212-2 (coarse dashed frame in FIG. 5). Data space 212-1 to data space 212-n of each participant 200 exist within the data space 212. FIG. 5 shows an example where n is 2.

[0042] Data space 212-1 is designated as area D1, and data space 212-2 is designated as area D2. As shown in FIG. 5, there may be an overlapping area between data space 212-1 and data space 212-2. This overlapping area is designated as overlapping area D3. In FIG. 5, overlapping area D3 is indicated by hatching. The area of ​​area D1 excluding overlapping area D3 is an area not covered by data space 212-2 of participant 200-2 (Company B in FIG. 3). The area of ​​area D2 excluding overlapping area D3 is an area not covered by data space 212-1 of participant 200-1 (Company A in FIG. 3).

[0043] FIG. 6 is another diagram showing an example of the data space 212. As with FIG. 5, the case where n is 2 will be described. For ease of illustration, FIG. 6 shows the data space 212 when there are two types of teacher explanatory variables 204 (the composition ratio of raw material A in FIG. 3 and the process temperature of process condition P1). In this case, the data space 212 has an axis indicating the composition ratio of raw material A and an axis indicating the process temperature. Each of the three white circles in FIG. 6 corresponds to data numbers A1 to A3 in FIG. 3, respectively, and each of the three black circles corresponds to data numbers B1 to B3 in FIG. 3, respectively. Each data space 212 contains more individual teacher data 202, but these are omitted in FIG. 6. In the example of FIG. 6, the data space 212 is a planar region based on two types of teacher explanatory variables 204. In FIG. 6, data space 212-1 is indicated by a fine dashed line, and data space 212-2 is indicated by a coarse dashed line.

[0044] Each data space 212 includes all of the individual teacher data 202 provided by the corresponding participant 200. The data space 212 may also include a region enclosed by connecting with straight lines any of the multiple individual teacher data 202 provided by the corresponding participant 200. The data space 212 may also be a region enclosed by connecting with straight lines the multiple outermost individual teacher data 202 provided by the corresponding participant 200. The data space 212 may be a region estimated based on the individual teacher data 202 using a machine learning model such as the k-nearest neighbor method or one-class support vector machine (OCSVM) used for the region estimation problem. For example, in the k-nearest neighbor method, the recommendation unit 20 calculates the distance d as the average value of k distances d' between an arbitrary data point and each of multiple individual explanatory variables 203 in the data space 212. The recommendation unit 20 may select the k distances d' in ascending order. The recommendation unit 20 may specify an area where the distance d is within an arbitrary threshold as the data space 212. The distance d′ may be the Euclidean distance between an arbitrary data point and the individual explanatory variable 203 in the data space 212, or may be the Mahalanobis distance.

[0045] The recommendation unit 20 may recommend optimized data Dt based on the target value Vt of the input objective variable and the data space information Is. The target value Vt of the input objective variable is a target value desired by the participant 200. In the examples of FIGS. 2 and 3, the target value Vt of the input objective variable is a physical property value of the target object that one participant 200 wants to obtain, and is at least one of the values ​​of "hardness" and "stretchability" in the evaluation process E. The target value Vt of the input objective variable may be different for each participant 200. The participant 200 may input the target value Vt using the input unit 30 via the terminal 300. The recommended optimized data Dt may be transmitted to the terminal 300 by the output unit 40. The optimized data Dt may be displayed on the display unit of the terminal 300.

[0046] The data space information Is is information related to the data space 212. The data space information Is includes information such as the range occupied by each data space 212 in the data space 214, the size of the data space 212, the density of the teacher explanatory variables 204 in the data space 212, the machine learning model for estimating the data space 212, and the maximum, minimum, and average values ​​of the distances between each data space 212. The optimized data Dt is an explanatory variable corresponding to the target value Vt. A candidate for the optimized data Dt is referred to as a candidate Ve. Multiple candidates Ve may be generated for one target value Vt.

[0047] For example, the recommendation unit 20 may generate the optimized data Dt within a range that is not included in the specific data space 212. The recommendation unit 20 may generate the optimized data Dt within a range in which the distance from the specific data space 212 is equal to or greater than a predetermined value. The distance between the data space 212 and the specific explanatory variable will be described later. The specific data space 212 is, for example, the data space 212 of the participant 200 who inputs the target value Vt. The recommendation unit 20 may generate the optimized data Dt from an overlapping region D3 (see FIG. 5) within the specific data space 212.

[0048] The recommendation unit 20 may generate one or more candidates Ve for the optimization data Dt based on the first teacher data 210 (see FIG. 3). The recommendation unit 20 may generate the candidates Ve based on the first teacher data 210 stored in the storage unit 50. The candidates Ve are candidates for the optimization data Dt, which may be explanatory variables corresponding to the target value Vt that one participant 200 desires to obtain. The recommendation unit 20 may generate the candidates Ve based on the first teacher data 210 and the target value Vt. The recommendation unit 20 may set one or more explanatory variables obtained by inputting the target value Vt into the estimation model 21 generated from the first teacher data 210 as the candidates Ve. The candidates Ve may include all parameters in the teacher explanatory variables 204 (see FIGS. 2 and 3). The candidates Ve may or may not be output to the terminal 300 by the output unit 40.

[0049] For example, if the teacher explanatory variable 204 has p types of parameters (for example, 15 types from “raw material A” to “temperature” of the evaluation condition Ev in FIG. 3) and the teacher objective variable 206 has one type of parameter (for example, “hardness” in FIG. 3), the p types of parameters of the teacher explanatory variable 204 are denoted by x1 to x p and the parameter of the teacher objective variable 206 is y. In this example, all data from data numbers A1 to N3 in FIG. 3 are approximated by the following formula 1 through regression analysis. Formula 1 is an example of the estimation model 21 (see FIG. 1). In formula 1, a, b, p, and k are constants.

number

[0050] The recommendation unit 20 calculates the parameters x1 to x2 when y in Equation 1 is the target value Vt. p The calculated parameters x1 to x p The explanatory variables specified by the combination of the following may be the candidate Ve. When y is the target value Vt, the parameters x1 to x2 that satisfy Equation 1 p Therefore, the recommendation unit 20 uses the parameters x1 to x p Multiple sets of the following may be recommended. Equation 1 is for a linear function, but all of the data from data numbers A1 to N3 in FIG. 3 may be approximated by a q-th order function (q is an integer equal to or greater than 2). All of the data from data numbers A1 to N3 may be approximated by a machine learning model such as a Support Vector Machine (SVM), Gaussian process regression, or Neural Network (NN).

[0051] The recommendation unit 20 may recommend one of the candidates Ve as the optimized data Dt based on the distance d between the generated candidate Ve and the data space 212. The recommendation unit 20 may calculate the distance d (e.g., d1 or d2 in FIG. 6) between each candidate Ve and a specific data space 212. The specific data space 212 is, for example, the data space 212 of the participant 200 who input the target value Vt. The recommendation unit 20 may calculate the distances (e.g., d1 and d2 in FIG. 6) between each candidate Ve and multiple data spaces 212.

[0052] The distance between the candidate Ve and one individual explanatory variable 203 (one of the individual explanatory variables 203-1 to 203-n) in one data space 212 is defined as distance d'. The p types of parameters of the individual explanatory variables 203 in one data space 212 (for example, data space 212-1) are defined as x1 to x p Then, the parameters x1 to x calculated by Eq. p are the parameters x1'~x p Then, the distance d' is expressed by the following equation 2.

number

[0053] A distance d′ may be defined for the individual explanatory variables 203 in each of the multiple data spaces 212 (data spaces 212-1 to 212-n). The recommendation unit 20 may calculate the distance d′ for the individual explanatory variables 203 in each of the multiple data spaces 212.

[0054] The recommendation unit 20 may calculate the distance d from multiple distances d' between the candidate Ve and each of multiple individual explanatory variables 203 in the data space 212. The distance d may be the average, median, or minimum value of the multiple distances d'. The recommendation unit 20 may calculate the distance d using a machine learning model that estimates the data space 212. For example, in the k-nearest neighbor method, the recommendation unit 20 calculates the distance d as the average value of k distances d' between the candidate Ve and each of multiple individual explanatory variables 203 in the data space 212. The recommendation unit 20 selects the k distances d' in ascending order. The distance d' may be the Euclidean distance between the candidate Ve and each of the individual explanatory variables 203 in the data space 212 or may be the Mahalanobis distance. The recommendation unit 20 may recommend, as the optimized data Dt, a candidate Ve whose distance d from a specific data space 212 satisfies a predetermined condition, from among one or more candidates Ve. For example, the recommendation unit 20 may recommend, as the optimized data Dt, the candidate Ve that has the longest distance d from the data space 212 of the participant 200 who input the target value Vt. The white star and black star in Fig. 6 are examples of the candidate Ve of the optimized data Dt, and the black star is an example of the optimized data Dt.

[0055] FIG. 7 is a diagram showing an example of learning by the recommendation unit 20. The recommendation unit 20 performs machine learning on first teacher data 210. The recommendation unit 20 may generate a machine learning model using the first teacher data 210. The recommendation unit 20 performs machine learning on the relationship between the teacher explanatory variables 204 and the teacher objective variable 206. The recommendation unit 20 may generate a machine learning model that approximates the relationship between the teacher explanatory variables 204 and the teacher objective variable 206. The recommendation unit 20 may generate a first inference model 22 by machine learning using the first teacher data 210. The first inference model 22 is an example of the estimation model 21 (see FIG. 1 ).

[0056] The recommendation unit 20 has a first inference model 22. When an input explanatory variable is input, the first inference model 22 outputs a first predicted value Vp1 of the input objective variable corresponding to the input explanatory variable. The input explanatory variable may be a candidate Ve, or may be an explanatory variable other than the candidate Ve provided by one of the participants 200. The input explanatory variable is input to the first inference model 22, which has undergone machine learning. The first inference model 22 has machine-learned the relationship between the teacher explanatory variable 204 and the teacher objective variable 206. Therefore, when the input explanatory variable is input to the first inference model 22, a first predicted value Vp1 of the input objective variable can be inferred. When a target value Vt is input to the recommendation unit 20, the recommendation unit 20 outputs optimization data Dt, which is an explanatory variable corresponding to the target value Vt.

[0057] The first inference model 22 may predict a first predicted value Vp1. The recommendation unit 20 may recommend one of the candidates Ve as the optimized data Dt based on the first predicted value Vp1 and the target value Vt. The recommendation unit 20 may recommend the optimized data Dt based on the difference between the first predicted value Vp1 and the target value Vt. If the difference between the first predicted value Vp1 and the target value Vt is less than a predetermined threshold, the recommendation unit 20 may recommend the candidate Ve corresponding to the first predicted value Vp1 as the optimized data Dt. If the difference between the first predicted value Vp1 and the target value Vt is equal to or greater than a predetermined threshold, the recommendation unit 20 may not recommend the optimized data Dt. If the first predicted value Vp1 exceeds the target value Vt, the recommendation unit 20 may recommend the optimized data Dt. If the target value Vt has multiple parameters, the recommendation unit 20 may recommend at least one of the Pareto optimal solutions for the multiple parameters as the optimized data Dt. The Pareto optimal solution is a solution that is not dominated by any other candidate Ve other than the Pareto optimal solution.

[0058] The contribution calculation unit 60 calculates the contribution of the first teacher data 210 corresponding to at least one participant 200 to the optimized data Dt based on the optimized data Dt and the first teacher data 210. The first teacher data 210 corresponding to at least one participant 200 refers to any of the individual teacher data 202-1 to 202-n.

[0059] For example, the contribution calculation unit 60 calculates the contribution of one individual teacher data 202 to the optimized data Dt based on the distance d between the optimized data Dt and the data space 212 of the individual teacher data 202. Each of the multiple individual teacher data 202 (individual teacher data 202-1 to individual teacher data 202-n) corresponds to a respective one of the multiple data spaces 212 (data space 212-1 to data space 212-n). The contribution calculation unit 60 may calculate the contribution of each of the multiple individual teacher data 202 (individual teacher data 202-1 to individual teacher data 202-n) to the optimized data Dt based on the distance d between the optimized data Dt and each of the multiple data spaces 212 (data space 212-1 to data space 212-n).

[0060] The distance d between the optimized data Dt and the data space 212-n is defined as the distance d n The contribution calculation unit 60 calculates the distance d n The smaller the distance d is, the higher the contribution of the individual training data 202-n corresponding to the data space 212-n to the optimized data Dt is calculated. n For example, the contribution calculation unit 60 may calculate the contribution C of the individual teaching data 202-u (u is an integer from 1 to n) to the optimized data Dt using the following formula 3: u Calculate [%].

number

[0061] The individual explanatory variables 203 (see Figure 3) are p types of parameters x1 to x p For example, the individual explanatory variable 203-1 in FIG. 3 has 15 types of parameters x1 to x 15 The individual explanatory variables 203 in one individual training data set 202 are p types of parameters x1 to x pIn the case where there are a plurality of sets of p types of parameters x1 to x2 in the individual explanatory variables 203, the recommendation unit 20 p For each of the pairs, the distance d′ may be calculated based on the above formula 2. The distance d is calculated based on the p types of parameters x1 to x2 in the individual explanatory variables 203. p The distance d′ may be the average, median, or minimum value of the distances d′ calculated for each of the sets.

[0062] When a target value Vt is input, the recommendation unit 20 may generate one or more candidates Ve of the optimized data Dt corresponding to the target value Vt. The first inference model 22 may output a first predicted value Vp1 of the input objective variable corresponding to the candidate Ve generated by the recommendation unit 20. When a candidate Ve is input, the first inference model 22 outputs the corresponding first predicted value Vp1. The first predicted value Vp1 does not necessarily match the target value Vt. The recommendation unit 20 may output one of the candidates Ve as the optimized data Dt based on the first predicted value Vp1 and the target value Vt. Of the multiple first predicted values ​​Vp1 corresponding to the multiple candidates Ve, the one closest to the target value Vt is determined to be the most likely predicted value. The recommendation unit 20 may recommend, as the optimized data Dt, the candidate Ve for which the first inference model 22 output the most likely predicted value from among the candidates Ve input to the first inference model 22.

[0063] 8 is a diagram showing another example of learning by the recommendation unit 20. The second teacher data 220 is teacher data obtained by excluding, from the first teacher data 210, the individual explanatory variables 203 and the individual objective variables 205 corresponding to at least one participant 200 out of the multiple participants 200. The second teacher data 220 is teacher data obtained by excluding, for example, the individual explanatory variables 203-1 and the individual objective variables 205-1 of participant 200-1 (Company A, see FIG. 3) from the first teacher data 210, and is teacher data into which the individual explanatory variables 203 and the individual objective variables 205 of participants 200-2 to 200-n (Company B to Company N, see FIG. 3) are aggregated.

[0064] The recommendation unit 20 may perform machine learning on the second teacher data 220. The recommendation unit 20 may generate a machine learning model using the second teacher data 220. The recommendation unit 20 performs machine learning on the relationship between the individual explanatory variables 203 and the individual objective variable 205. The recommendation unit 20 may generate a machine learning model that approximates the relationship between the individual explanatory variables 203 and the individual objective variable 205. The recommendation unit 20 may generate a second inference model 24 by performing machine learning on the second teacher data 220. The recommendation unit 20 may generate the second inference model 24 by machine learning using the second teacher data 220. The recommendation unit 20 may include the second inference model 24.

[0065] When an input explanatory variable is input, the second inference model 24 outputs a second predicted value Vp2 of the input objective variable corresponding to the input explanatory variable. The contribution calculation unit 60 may calculate the contribution of one individual training data 202 corresponding to one participant 200 to the optimization data Dt based on the first predicted value Vp1 and the second predicted value Vp2.

[0066] When the second inference model 24 is generated by machine learning the second teacher data 220 from which one individual teacher data 202-u (1≦u≦n) is excluded, the second predicted value Vp2 output by the second inference model 24 is Vp2 u For example, the contribution calculation unit 60 calculates the contribution C of one individual training data 202-u to the optimized data Dt using the following formula 4: u Calculate.

number

[0067] As shown in Equation 4, the contribution calculation section 60 may calculate the contribution to the optimized data Dt based on the difference between the first predicted value Vp1 and the second predicted value Vp2. The contribution calculation section 60 may evaluate the contribution to the optimized data Dt as greater the greater the difference.

[0068] When the difference between the first predicted value Vp1 and the second predicted value Vp2 is relatively small, the effect on the first predicted value Vp1 of excluding the individual teacher data 202-1 from the first teacher data 210 (see FIG. 3) is relatively small. Therefore, the contribution calculation unit 60 may calculate the contribution of the individual teacher data 202-1 to the optimized data Dt to be smaller as the difference between the first predicted value Vp1 and the second predicted value Vp2 is smaller.

[0069] When the teacher objective variable 206 and the individual objective variable 205 have multiple types of parameters, the contribution calculation unit 60 may calculate the contribution of the individual teacher data 202 to the optimization data Dt for each of the multiple types of parameters. The contribution calculation unit 60 may calculate the average, median, minimum, or maximum of the contribution calculated for each of the multiple types of parameters as the contribution of the individual teacher data 202 to the optimization data Dt.

[0070] In the example of Fig. 3, the teacher objective variable 206 and the individual objective variable 205 have two types of parameters. In the example of Fig. 3, the contribution calculation unit 60 may calculate the contribution of the individual teacher data 202 to the optimization data Dt for each of the two types of parameters ("stiffness" and "stretchability"). The contribution calculation unit 60 may calculate the average, median, minimum, or maximum of the contribution calculated for one parameter (stiffness) and the contribution calculated for the other parameter (stretchability) as the contribution of the individual teacher data 202 to the optimization data Dt.

[0071] The first inference model 22 may output first predicted values ​​Vp1 of multiple input objective variables. In the example of Figure 3, the first inference model 22 may output first predicted values ​​Vp1 of "hardness" and "stretchability" in the "evaluation step E." Similarly, the second inference model 24 may output second predicted values ​​Vp2 of multiple input objective variables.

[0072] Weights may be assigned in advance to the multiple objective variables. The weights may be assigned in advance by each of the multiple participants 200. Each of the multiple participants 200 may assign a different weight to the multiple objective variables. Each of the multiple objective variables may have an objective variable that they place more importance on. In the example of FIG. 3, for example, participant 200-1 assigns a weight of 70% to the objective variable "hardness" and a weight of 30% to the objective variable "stretchability." A single weight agreed upon by the multiple participants 200 may be assigned to the multiple objective variables. The assigned weights may be stored in the storage unit 50.

[0073] The contribution calculation unit 60 may calculate the contribution of the individual teacher data 202 to the optimized data Dt based on the first predicted value Vp1, the second predicted value Vp2, and weights assigned to the multiple objective variables. The contribution calculation unit 60 may calculate the contribution for each of the multiple objective variables based on the first predicted value Vp1 and the second predicted value Vp2, and may calculate the contribution of the individual teacher data 202 to the optimized data Dt by weighted averaging the calculated contributions. In the example of FIG. 3, the contribution calculation unit 60 calculates the contribution of the individual teacher data 202-1 to the optimized data Dt for each objective variable, “stiffness” and “stretchability.” The contribution calculation unit 60 calculates the contribution of the individual teacher data 202-1 to the optimized data Dt by weighted averaging the contribution calculated for “stiffness” and the contribution calculated for “stretchability” at a ratio of, for example, 70% and 30%, respectively.

[0074] The recommendation unit 20 may generate the second inference model 24 by machine learning the second teacher data 220 from which the individual explanatory variables 203 and the individual objective variables 205 corresponding to at least two of the multiple participants 200 have been excluded. For example, in the example of Fig. 3, when n is 3 or greater, the recommendation unit 20 performs machine learning on the second teacher data 220 from which the individual explanatory variables 203-1 and the individual objective variables 205-1 corresponding to participant 200-1, and the individual explanatory variables 203-2 and the individual objective variables 205-2 corresponding to participant 200-2 have been excluded.

[0075] When the input explanatory variables are input, the second inference model 24 outputs a second predicted value Vp2 of the input objective variable. The contribution calculation unit 60 may calculate the contributions of at least two of the multiple participants 200 to the optimized data Dt based on the first predicted value Vp1 and the second predicted value Vp2. For example, the contribution calculation unit 60 calculates the contributions of the participants 200-1 and 200-2 to the optimized data Dt based on the first predicted value Vp1 and the second predicted value Vp2.

[0076] When the recommendation unit 20 performs machine learning on the second teacher data 220 from which one of the individual teacher data 202-1 provided by the participant 200-1 and the individual teacher data 202-2 provided by the participant 200-2 has been excluded, the second predicted value Vp2 output by the second inference model 24 is referred to as the second predicted value Vp2-1. When the recommendation unit 20 performs machine learning on the second teacher data 220 from which the other of the individual teacher data 202-1 and the individual teacher data 202-2 has been excluded, the second predicted value Vp2 output by the second inference model 24 is referred to as the second predicted value Vp2-2. When the recommendation unit 20 performs machine learning on the second teacher data 220 from which both the individual teacher data 202-1 and the individual teacher data 202-2 have been excluded, the second predicted value Vp2 output by the second inference model 24 is referred to as the second predicted value Vp2-3.

[0077] When the individual teacher data 202-1 and the individual teacher data 202-2 are similar and the difference between the second predicted value Vp2-3 and the first predicted value Vp1 is relatively large, the difference between the second predicted value Vp2-1 and the first predicted value Vp1 and the difference between the second predicted value Vp2-2 and the first predicted value Vp1 are likely to be relatively small. Therefore, the contribution calculation unit 60 can calculate the contributions of the participants 200-1 and 200-2 to the optimization data Dt.

[0078] The recommendation unit 20 may calculate the similarity between the individual teacher data 202-1 and the individual teacher data 202-2 based on the distance ds between the individual teacher data 202-1 and the individual teacher data 202-2. In the example of FIG. 3, p types of parameters of the individual explanatory variable 203-1 are set as x 11 ~x 1p The r types of parameters of the individual objective variable 205-1 are expressed as y 11 ~y 1r Similarly, let p types of parameters of individual explanatory variables 203-2 be x 21 ~x 2p The r types of parameters of the individual objective variable 205-2 are set as y 21 ~y 2r The distance ds is expressed by the following equation 5.

number

[0079] The recommendation unit 20 may calculate the similarity between the individual teacher data 202-1 and the individual teacher data 202-2 to be greater as the distance ds becomes smaller. A threshold value may be set in advance for the distance ds. The recommendation unit 20 may determine that the individual teacher data 202-1 and the individual teacher data 202-2 are similar when the distance ds is smaller than the threshold value.

[0080] When the optimized data Dt is input, the first inference model 22 may output a first predicted value Vp1 corresponding to the optimized data Dt. The first predicted value Vp1 corresponding to the optimized data Dt is defined as the first predicted value Vp1'. When the optimized data Dt is input, the second inference model 24 may output a second predicted value Vp2 corresponding to the optimized data Dt. The second predicted value Vp2 corresponding to the optimized data Dt is defined as the first predicted value Vp2'.

[0081] The contribution calculation unit 60 may calculate the contribution of the individual teacher data 202 to the optimized data Dt based on the first predicted value Vp1' and the second predicted value Vp2'. For example, the contribution calculation unit 60 calculates the contribution of the individual teacher data 202 to the optimized data Dt based on the difference between the first predicted value Vp1' and the second predicted value Vp2'. The contribution calculation unit 60 may evaluate the contribution of the individual teacher data 202 to the optimized data Dt to be higher the larger the difference between the first predicted value Vp1' and the second predicted value Vp2'.

[0082] 9 is a diagram showing another example of inference using the second inference model 24. The recommendation unit 20 may perform machine learning on the second teacher data 220 for each of the multiple participants 200, from which the individual explanatory variables 203 and the individual objective variables 205 corresponding to one of the multiple participants 200 have been excluded. In the example of FIG. 9, the recommendation unit 20 performs machine learning on the second teacher data 220-u for each of n participants 200, from which the individual explanatory variables 203-u and the individual objective variables 205-u corresponding to participant 200-u (1≦u≦n) have been excluded, thereby generating a second inference model 24 (second inference model 24-1 to second inference model 24-n) for each of the multiple participants 200.

[0083] Figure 10 is a diagram showing another example of inference by the second inference model 24. The second inference model 24 may output a second predicted value Vp2 for each of multiple participants 200. In the example of Figure 10, the second inference model 24-1 outputs second predicted values ​​Vp2-1 to Vp2-n for participants 200-1 to 200-n, respectively.

[0084] The contribution calculation section 60 may calculate the contribution of each of the multiple participants 200 to the optimized data Dt based on the first predicted value Vp1 and the second predicted value Vp2 of each of the multiple participants 200. In the example of Fig. 10, when n is 2, the contribution calculation section 60 calculates the contribution of the participant 200-1 based on the first predicted value Vp1 and the second predicted value Vp2-1, and calculates the contribution of the participant 200-2 based on the first predicted value Vp1 and the second predicted value Vp2-2.

[0085] For example, the contribution calculation unit 60 calculates the difference between the first predicted value Vp1 and the second predicted value Vp2 of each of the multiple participants 200, and evaluates the contribution of the participant 200 to the optimized data Dt higher the larger the calculated difference. The contribution calculation unit 60 may rank the contributions of the participants 200 based on the magnitude of the difference between the first predicted value Vp1 and the second predicted value Vp2 of each of the multiple participants 200.

[0086] The selection unit 70 (see FIG. 4) may select at least one participant 200 to be prompted to input additional individual teacher data 202 to the recommendation unit 20 based on the participant 200's degree of contribution to the optimization data Dt. For example, if the contribution calculation unit 60 ranks the participants 200's degrees of contribution, the selection unit 70 may select one or more participants 200 with the lowest degree of contribution as participants 200 to be prompted to input individual teacher data 202. A contribution threshold may be set in advance for the degree of contribution. The threshold may be stored in the memory unit 50. The selection unit 70 may select one or more participants 200 whose degree of contribution is less than the contribution threshold as participants 200 to be prompted to input individual teacher data 202.

[0087] The output unit 40 (see FIG. 4) may output information to the terminal 300 of one or more participants 200 whose contribution level is the lowest or below the contribution level threshold, prompting them to input individual training data 202. The information may be displayed on the terminal 300 of the one or more participants 200. The first inference model 22 performs machine learning on the first training data 210. The first training data 210 is data obtained by aggregating the individual training data 202 provided by each of the multiple participants 200. Therefore, the quality of the optimized data Dt may be affected by the quality of the individual training data 202 provided by each of the multiple participants 200. Therefore, it is preferable for participants 200 with low contribution levels to provide higher quality individual training data 202.

[0088] FIG. 11 is a diagram showing another example of the first teacher data 210 provided by each of the multiple participants 200. When the teacher explanatory variable 204 has multiple types of parameters (15 types in this example, from "raw material A" to "teacher objective variable 206" of evaluation condition E), at least one type of parameter may be provided by at least one participant 200. The individual teacher data 208 may be provided by each of the multiple participants 200. In this example, the individual teacher data 208-1 to 208-6 are provided by the participants 200-1 to 200-6, respectively, and the individual teacher data 208-7 is provided by the participant 200-6. In this example, the first teacher data 210 is data obtained by aggregating the individual teacher data 208-1 to 208-7. This example differs from the example in FIG. 3 in this respect.

[0089] In this example, the teacher explanatory variables 204 are provided by participants 200-1 to 200-6, and the teacher objective variable 206 is provided by participant 200-6. In this example, the composition Co of the main material Sm is provided by participant 200-1, the composition Co of the secondary material Sb is provided by participant 200-2, and the composition Co of the additional material Sa is provided by participant 200-3. In this example, the process conditions P1 are provided by participant 200-4, the process conditions P2 are provided by participant 200-5, and the evaluation conditions Ev and the teacher objective variable 206 are provided by participant 200-6.

[0090] Multiple types of parameters may be provided by at least one participant 200. In this example, participant 200-1 provides three types of parameters, namely, the composition Co of raw material A to the composition Co of raw material C. Similarly, in this example, participants 200-1 to 200-6 each provide multiple types of parameters.

[0091] The recommendation unit 20 (see FIG. 4) may generate the first inference model 22 by machine learning the teacher explanatory variables 204 and the teacher objective variables 206. When a target value Vt of the input objective variable is input to the recommendation unit 20, the recommendation unit 20 outputs optimized data Dt, which is an explanatory variable corresponding to the target value Vt.

[0092] The contribution calculation unit 60 may calculate the contribution of at least one parameter to the optimized data Dt based on the optimized data Dt and the first teacher data 210. In this example, the contribution calculation unit 60 calculates the contribution of each of the individual teacher data 208-1 to 208-6 to the optimized data Dt based on the optimized data Dt and the first teacher data 210. The contribution calculation unit 60 may calculate the contribution h of at least one species to the optimized data Dt for each of the multiple species. For example, the contribution calculation unit 60 may calculate the contribution h of each of the multiple species of parameters to the optimized data Dt based on the optimized data Dt and the first teacher data 210.

[0093] The contribution h may be calculated based on the average predicted value Pa of the first inference model. The predicted value Pa may be the average value of the objective variable in multiple optimization data output by the first inference model, or may be the average value of the objective variable in the first teacher data 210. Specifically, data is generated in which all explanatory variables except for some of the explanatory variables xi in the optimization data Dt are replaced with the average values ​​of the explanatory variables included in the first teacher data 210, and this data is input to the first inference model to obtain the predicted value Ph of the objective variable. Then, the contribution h of xi may be calculated based on the difference between Ph and Pa.

[0094] For simplicity, we will explain an example where there are three explanatory variables (x1, x2, x3) and one objective variable (y). In this case, the explanatory variables of the optimized data are (x1o, x2o, x3o), the predicted value of the objective variable when the optimized data is input to the first inference model is yo, the average values ​​of the explanatory variables in the first training data are (x1a, x2a, x3a), and the average value of the objective variable y is Pa.

[0095] When yo=50 and Pa=30, the difference of 20 is the total contribution of the optimization data. Here, when calculating the contribution of each explanatory variable, we will further explain using x1 as an example. Suppose that by inputting (x1o, x2a, x3a) into the first inference model, the predicted value Ph of the objective variable is 35. In this case, the contribution of x1 can be calculated as h=5, which is the difference between Ph=35 and Pa=30.

[0096] Specifically, the Shapley value or its approximate value, the SHAP (SHapley Additive exPlanations) value, may be calculated as the contribution h. The Shapley value h is a distribution algorithm used when a reward earned through the cooperation of multiple participants is distributed according to the contributions of each of the multiple participants. The contribution calculation unit 60 may calculate the Shapley value h based on the optimization data Dt and the first teacher data 210 using a known method (for example, paragraphs

[0007] to

[0010] of Japanese Patent No. 7086497).

[0097] The contribution calculation unit 60 calculates the contribution h of the p types of parameters of the teacher explanatory variables 204 to the optimization data Dt. s (1≦s≦p) for each of the plurality of species. s For example, the contribution calculation unit 60 calculates the contribution C of each of the plurality of individual training data 208 to the optimized data Dt according to the following 6.

number

[0098] In Equation 6, p is the total number of parameters in the teacher explanatory variables 204 (15 in this example). The denominator of Equation 6 is the sum of the contributions of all parameters in the teacher explanatory variables 204. In Equation 6, i is the number of parameters provided by each participant 200. For example, i for participant 200-1 is 3. The numerator of Equation 6 is the sum of the contributions of parameters in one or more teacher explanatory variables 204 provided by each participant 200.

[0099] 12 is a flowchart showing an example of an information processing method according to an embodiment of the present invention. The information processing method according to an embodiment of the present invention will be described using the information processing device 100 shown in FIG. 4 as an example. The information processing method may include a data providing and aggregating step. The data providing and aggregating step may include a data providing step S70, a data aggregating step S84, a data excluding step S88, and a preprocessing step S90. The preprocessing step S90 may include a determining step S92 and a missing value adding step S94.

[0100] The data provision step S70 is a step in which each of the multiple participants 200 (participants 200-1 to 200-n) provides individual teacher data 202 (individual teacher data 202-1 to individual teacher data 202-n, see FIG. 3) or individual teacher data 208 (individual teacher data 208-1 to individual teacher data 208-n', see FIG. 11). In the example of FIG. 11, n' is 7.

[0101] Determination step S80 is a step in which the recommendation unit 20 determines whether at least one of the parameters is missing in at least one of the individual teacher data 202-1 to 202-n or at least one of the individual teacher data 208-1 to 208-n. In the example of FIG. 3, the recommendation unit 20 determines whether at least one of 15 parameters from the composition Co of raw material A to the temperature of the evaluation condition Ev is missing. In the example of FIG. 11, the recommendation unit 20 determines whether at least one of three parameters in the composition Co of raw material A is missing, for example. If it is determined that at least one of the parameters is missing, the information processing method proceeds to missing parameter addition step S82. If it is determined that at least one of the parameters is not missing, the information processing method proceeds to data aggregation step S84.

[0102] The data aggregation step S84 is a step in which the recommendation unit 20 aggregates the individual teacher data 202-1 to 202-n, or aggregates the individual teacher data 208-1 to 208-n'. In the data aggregation step S84, the recommendation unit 20 aggregates the individual teacher data 202-1 to 202-n, or aggregates the individual teacher data 208-1 to 208-n', to generate first teacher data 210 (see FIGS. 3 and 11).

[0103] The determination step S92 is a step in which the recommendation unit 20 determines whether there are any parameters with missing values ​​in the first training data 210 (see FIGS. 3 and 11) generated in the data aggregation step S84. If it is determined that there are any parameters with missing values, the information processing method proceeds to the missing value addition step S94. If it is not determined that there are any parameters with missing values, the information processing method ends the data provision and aggregation steps.

[0104] 13 is a flowchart showing an example of an information processing method according to an embodiment of the present invention. The information processing method of this example includes a first inference model generation step. The first inference model generation step may include a first inference model generation step S200, a model saving step S202, and a determination step S204. The first inference model generation step S200 is a step in which the recommendation unit 20 performs machine learning on the first teacher data 210 to generate a first inference model 22 that outputs a first predicted value Vp1 of an input objective variable when an input explanatory variable is input. In the first inference model generation step S200, the recommendation unit 20 may generate multiple first inference models 22.

[0105] The model saving step S202 is a step in which the recommendation unit 20 stores the first inference model 22 in the memory unit 50. If multiple first inference models 22 are generated in the first inference model generating step S200, the recommendation unit 20 may store the first inference model 22 with the highest inference accuracy in the memory unit 50 in the model saving step S202. The recommendation unit 20 may determine the accuracy of the first inference model 22 using an index such as the mean absolute error or coefficient of determination, which represents the error between the true value and the predicted value.

[0106] The judgment step S204 is a step in which the recommendation unit 20 judges whether it has performed machine learning for all the objective variables (in the examples of Figures 3 and 11, "hardness" and "stretchability" in evaluation step E). If it is not judged that machine learning has been performed for all the objective variables, the information processing method returns to the first inference model generation step S200. If it is judged that machine learning has been performed for all the objective variables, the information processing method ends the first inference model generation step.

[0107] 14 is a flowchart showing an example of an information processing method according to an embodiment of the present invention. The information processing method of this example includes a second inference model generation step. The second inference model generation step may include a data exclusion step S88, a second inference model generation step S240, a model storage step S242, and a judgment step S244. The data exclusion step S88 is a step in which the recommendation unit 20 excludes individual training data 202 corresponding to at least one of the multiple participants 200 from the first training data 210. By excluding one individual training data 202 from the first training data 210 in the data exclusion step S88, second training data 220 is created.

[0108] The second inference model generation step S240 is a step in which the recommendation unit 20 performs machine learning on the second teacher data 220 to generate a second inference model 24 that outputs a second predicted value Vp2 of the input objective variable when an input explanatory variable is input. In the data exclusion step S88 (see FIG. 16 ), when the second teacher data 220 is created by excluding individual teacher data 202 corresponding to at least two of the multiple participants 200 from the first teacher data 210, the second inference model generation step S240 may be a step in which the recommendation unit 20 performs machine learning on the second teacher data 220 to generate the second inference model 24. The model storage step S242 is a step in which the recommendation unit 20 stores the second inference model 24 in the storage unit 50.

[0109] Determination step S244 is a step in which recommendation unit 20 determines whether machine learning has been performed for all objective variables (in the examples of Figures 3 and 11, "hardness" and "stretchability" in evaluation step E). If it is determined that machine learning has not been performed for all objective variables, the information processing method returns to second inference model generation step S240. If it is determined that machine learning has been performed for all objective variables, the information processing method terminates the second inference model generation step.

[0110] 15 is a flowchart showing an example of an information processing method according to an embodiment of the present invention. The information processing method may include a determination step S100. The determination step may include a spatial information saving step S102 and a determination step S104.

[0111] The identification step S100 is a step in which the identification unit 10 identifies the data space 212 of each of the multiple participants 200 based on individual teacher data 202 having teacher explanatory variables 204 and teacher objective variables 206 provided by each of the multiple participants 200. The spatial information saving step S102 is a step in which the recommendation unit 20 stores the spatial information of the data space 212 of each of the multiple participants 200 in the storage unit 50. The determination step S104 is a step in which the recommendation unit 20 determines whether the data spaces 212 of all participants 200 have been identified. If it is determined that the data spaces 212 of all participants 200 have been identified, the information processing method proceeds to an output step (described below). If it is not determined that the data spaces 212 of all participants 200 have been identified, the information processing method returns to the identification step S100.

[0112] 16 is a flowchart showing an example of an information processing method according to an embodiment of the present invention. The information processing method includes an output step. The output step includes an output step S128. The output step may further include a target value acquisition step S108, a candidate generation step S112, a distance d calculation step S116, a predicted value prediction step S120, and a judgment step S124.

[0113] The target value acquisition step S108 is a step in which the recommendation unit 20 acquires the target value Vt of the input objective variable. Each participant 200 may input the target value Vt via the input unit 30. The recommendation unit 20 may acquire the target value Vt input by each participant 200.

[0114] The candidate generation step S112 is a step of generating a candidate Ve of optimized data Dt corresponding to the target value Vt when the target value Vt is input to the recommendation unit 20. The distance d calculation step S116 is a step in which the recommendation unit 20 calculates the distance d between the candidate Ve generated in the candidate generation step S112 and the data space 212 identified in the identification step S100. The output step S128 may be a step in which the recommendation unit 20 outputs one of the candidates Ve as optimized data Dt based on the candidate Ve generated in the candidate generation step S112 and the distance d calculated in the distance d calculation step S116.

[0115] The predicted value prediction step S120 is a step in which the first inference model 22 predicts a predicted value Vp of an input objective variable corresponding to a candidate Ve in the optimized data Dt. The determination step S124 may be a step in which the recommendation unit 20 determines not to generate a new candidate Ve when the difference between the predicted value Vp and the target value Vt is less than a predetermined threshold. The determination step S124 may also be a step in which the recommendation unit 20 determines not to generate a new candidate Ve when the predicted value Vp exceeds the target value Vt. The determination step S124 may also be a step in which the recommendation unit 20 determines not to generate a new candidate Ve when the number of times candidates have been generated in the candidate generation step has reached a predetermined number.

[0116] The determination step S124 is a step in which the recommendation unit 20 determines whether to generate a new candidate Ve. The recommendation unit 20 may determine not to generate a new candidate Ve when the difference between the predicted value Vp and the target value Vt is less than a predetermined threshold. The recommendation unit 20 may determine not to generate a new candidate Ve when the predicted value Vp exceeds the target value Vt. The recommendation unit 20 may determine not to generate a new candidate Ve when the number of times candidates have been generated in the candidate generation step S112 reaches a predetermined number. Furthermore, the determination step S124 may determine whether the recommendation unit 20 generates a new candidate Ve based on the data space information Is. The determination step S124 may determine whether the recommendation unit 20 generates a new candidate Ve based on the candidate Ve generated in the candidate generation step S112 and the distance d calculated in the distance d calculation step S116. The recommendation unit 20 may determine not to generate a new candidate Ve when the distance d exceeds a predetermined value.

[0117] The output step S128 may be a step in which, when a target value Vt of an input objective variable is input to the recommendation unit 20, the recommendation unit 20 outputs optimized data Dt corresponding to the target value Vt. The output step S128 may be a step in which the recommendation unit 20 recommends optimized data Dt, which is an explanatory variable corresponding to the target value Vt, based on the target value Vt and the data space information Is. The output step S128 may be a step in which the recommendation unit 20 recommends one of the candidates Ve as optimized data Dt based on the candidates Ve generated in the candidate generation step S112 and the distance d calculated in the distance d calculation step S116. The optimized data Dt recommended in the output step S128 may be output by the output unit 40 to the terminal 300 of the participant 200.

[0118] The output step S128 is a step in which, when a target value Vt of the input objective variable is input to the recommendation unit 20, the recommendation unit 20 outputs optimized data Dt corresponding to the target value Vt. The output step S128 may be a step in which the recommendation unit 20 outputs the optimized data Dt based on the predicted value Vp predicted in the predicted value prediction step S120 and the target value Vt acquired in the target value acquisition step S108.

[0119] The identification step S100 (see FIG. 15) may be a step in which the identification unit 10 identifies the first data space 212-1 of the first participant 200-1. The candidate generation step S112 may be a step in which the recommendation unit 20 generates a first candidate Ve1 for the optimized data Dt in the first participant 200-1. The distance d calculation step S116 may be a step in which the recommendation unit 20 calculates a first distance d1 between the first candidate and the first data space 212-1. The recommendation step S128 may be a step in which the recommendation unit 20 recommends the optimized data Dt based on the first distance d1.

[0120] The identification step S100 (see FIG. 15) may be a step in which the identification unit 10 identifies the second data space 212-2 of the second participant 200-2. The distance d calculation step S116 may be a step in which the recommendation unit 20 calculates a second distance d2 between the first candidate Ve1 of the optimized data Dt for the first participant 200-1 and the second data space 212-2. The recommendation step S128 may be a step in which the recommendation unit 20 recommends the optimized data Dt further based on the second distance d2.

[0121] The output step S128 may be a step in which the recommendation unit 20 determines the optimized data Dt such that the first distance d1 is equal to or greater than the first threshold distance. The output step S128 may be a step in which the recommendation unit 20 determines the optimized data Dt such that the second distance d2 is equal to or less than the second threshold distance. The output step S128 may be a step in which the recommendation unit 20 determines the optimized data Dt such that the first distance d1 is equal to or greater than the first threshold distance and the second distance d2 is equal to or less than the second threshold distance.

[0122] The determination step S129 is a step in which the recommendation unit 20 determines whether the recommendation of the optimized data Dt to all participants 200 has been completed. If it is not determined that the recommendation of the optimized data Dt has been completed, the information processing method returns to the target value acquisition step S108. In the target value acquisition step S108, the recommendation unit 20 may acquire the target values ​​Vt of the participants 200 for whom the recommendation of the optimized data Dt has not been completed. If it is determined that the recommendation of the optimized data Dt has been completed, the information processing method ends the output step.

[0123] 17 is a flowchart showing an example of an information processing method according to an embodiment of the present invention. The information processing method includes a contribution calculation step. The contribution calculation step includes a contribution calculation step S130. The contribution calculation step may further include a selection step S131, a determination step S132, and a contribution calculation step S134.

[0124] The contribution calculation step S130 is a step in which the contribution calculation unit 60 calculates the contribution of the first teacher data 210 (i.e., one individual teacher data 202) corresponding to at least one participant 200 to the optimized data Dt based on the optimized data Dt and the first teacher data 210. The contribution calculation step S130 may be a step in which the contribution calculation unit 60 calculates the contribution of one individual teacher data 202 to the optimized data Dt based on the first predicted value Vp1 and the second predicted value Vp2. The contribution calculation step S130 may be a step in which the contribution calculation unit 60 calculates the contribution of one individual teacher data 202 to the optimized data Dt based on the first predicted value Vp1 corresponding to the optimized data Dt and the second predicted value Vp2 corresponding to the optimized data Dt.

[0125] When the recommendation unit 20 machine-learns the second teacher data 220, which is the individual teacher data 202 corresponding to at least two of the multiple participants 200 and is excluded from the first teacher data 210, the contribution calculation step S130 may be a step of calculating the contribution of at least two of the participants to the optimized data Dt based on the first predicted value Vp1 and the second predicted value Vp2.

[0126] The contribution calculation step S130 may be a step in which the contribution calculation unit 60 calculates the contribution of each of the multiple participants 200 to the optimized data Dt based on the first predicted value Vp1 and the second predicted value Vp2 of each of the multiple participants 200. The selection step S131 is a step in which the selection unit 70 selects at least one participant 200 to prompt the recommendation unit 20 to input additional individual teacher data 202 or individual teacher data 208 (see FIG. 11 ) based on the contribution to the optimized data Dt.

[0127] When the teacher explanatory variable 204 has multiple types of parameters (15 types in the example of Figure 11, from ``raw material A'' to ``teacher objective variable 206'' of evaluation condition E) and at least one type of parameter is provided by at least one participant 200, the contribution calculation step S130 may be a step in which the contribution calculation unit 60 calculates the contribution of at least one parameter to the optimization data Dt based on the optimization data Dt and the first teacher data 210.

[0128] The contribution calculation step S130 may be a step in which the contribution calculation unit 60 calculates the contribution h of at least one species to the optimization data Dt for each of the multiple species, calculates the contribution C of at least one participant based on the calculated contribution h for each of the multiple species, and calculates the contribution C of at least one participant based on the calculated contribution h for each of the multiple species. Furthermore, in a case in which weights are assigned in advance to the multiple objective variables, the contribution calculation step S130 may be a step in which the contribution of one individual training data 202 to the optimization data Dt is calculated based on the first predicted value Vp1, the second predicted value Vp2, and the weights.

[0129] Determination step S132 is a step in which the recommendation unit 20 determines whether calculation of the contribution to the optimized data Dt has been completed for all participants 200. If it is determined that calculation of the contribution has not been completed, the information processing method returns to target value acquisition step S108. In target value acquisition step S108, the recommendation unit 20 may acquire target values ​​Vt for participants 200 for which calculation of the contribution has not been completed. If it is determined that calculation of the contribution has been completed for all participants 200, the information processing method ends the information processing.

[0130] 18 is a diagram showing another example of the configuration of the information processing device 100. The information processing device 100 of this example includes a recommendation unit 20, a contribution calculation unit 60, a plan formulation unit 72, and a control unit 74. The information processing device 100 may further include one or more, or all, of an identification unit 10, an input unit 30, an output unit 40, a storage unit 50, and a selection unit 70. Components in FIG. 18 that are assigned the same reference numerals as those in FIG. 4 may have the same functions and configurations as any of the examples described in FIG. 1 to FIG. 17.

[0131] The information processing device 100 of this example facilitates planning the collection of additional experimental data based on the contribution of existing experimental data to a certain objective variable, and collects data from the experimental equipment based on the plan. The recommendation unit 20 has a first inference model 22, and when a target value of the objective variable is input, outputs optimized data that is an explanatory variable corresponding to the target value. The operation of the recommendation unit 20 may be similar to any of the examples described in Figures 1 to 17.

[0132] The contribution calculation unit 60 may calculate the contribution of at least one piece of experimental data included in the first training data to the optimization data based on the optimization data and the first training data (training data used in machine learning of the first inference model). The operation of the contribution calculation unit 60 may be similar to any of the examples described with reference to FIGS. 1 to 17.

[0133] The contribution calculation unit 60 may calculate the contribution of each type of explanatory variable included in the experimental data. The type of explanatory variable includes the type, shape, and size of the input material, the temperature in each process, the time of each process, etc. The type of explanatory variable for which the contribution is to be calculated is set as the target type. The contribution calculation unit 60 can calculate the contribution of the explanatory variable of the target type based on the distance between optimized data generated from training data that excludes the explanatory variables of the target type and optimized data generated from training data that does not exclude the explanatory variables of the target type. The specific method for calculating the contribution is the same as the examples in FIGS. 1 to 17. The contribution calculation unit 60 may calculate the contribution of each type of explanatory variable by processing each type of explanatory variable in turn as the target type.

[0134] The plan formulation unit 72 formulates and outputs an additional experimental data collection plan based on the contribution calculated by the contribution calculation unit 60. For example, the plan formulation unit 72 formulates an experimental data collection plan for conducting an additional experiment that includes a type of explanatory variable with a high contribution. The plan formulation unit 72 may formulate an experimental data collection plan for conducting multiple additional experiments in which the values ​​of the type of explanatory variable are changed. The experimental data collection plan may specify the type of at least one explanatory variable in the additional experiment to be performed, and the value or tolerance range of the explanatory variable. The experimental data collection plan may also specify the types of all explanatory variables in the additional experiment to be performed, and the value or tolerance range of each explanatory variable.

[0135] The control unit 74 may acquire experimental data by controlling the experimental equipment based on the data collection plan. The control unit 74 may generate control data for controlling the experimental equipment based on the data collection plan and send it to the experimental equipment. The experimental equipment may have an execution unit that causes the experimental equipment to execute an experiment in accordance with the control data. The control unit 74 may cause the experimental equipment to execute an experiment in accordance with the data collection plan by sending the data collection plan to the experimental equipment. The experimental equipment may have an execution unit that causes the experimental equipment to execute an experiment in accordance with the data collection plan. In this case, the execution unit converts the data collection plan into control data for the experimental equipment.

[0136] The control unit 74 acquires experimental data, which is the result of an experiment performed by the experimental equipment based on the data collection plan. The experimental data may include explanatory variables and target variables. The control unit 74 may provide the newly acquired experimental data to the recommendation unit 20. The recommendation unit 20 may generate each inference model based on the training data to which the new experimental data has been added. This type of control makes it possible to acquire additional experimental data with a high degree of contribution, making it easy to generate more accurate inference models.

[0137] The contribution calculation unit 60 can specify any conditions for output based on the experimental data collection strategy. In this example, the contribution calculation unit 60 calculates the contribution of each explanatory variable to determine whether the level of a factor obtained through an experiment is excessive or insufficient. The contribution can be output in a format (described later) that satisfies any privacy level.

[0138] The following are examples of specific methods for calculating and presenting the contribution of explanatory variables. To maximize disclosure to each user, the contribution of all explanatory variables may be presented to each user. The type of explanatory variables for which contributions are to be disclosed may be selected depending on the privacy level set. For example, if the required privacy level is high, only the values ​​related to the explanatory variables provided by each user may be presented, and other explanatory variables may be kept confidential. In order to utilize the relative importance of each explanatory variable, the contribution calculation unit 60 may calculate the relative ranking or deviation of the contributions of each explanatory variable. A more specific example of the calculation algorithm is as follows: (1) One or more types of explanatory variables for acquiring information about the degree of contribution are determined. The types of explanatory variables may be determined by one or more participants 200. In addition, the scope of disclosure to each participant 200 may be set from among the following (A) to (D). (A) The contribution values ​​of all explanatory variables are disclosed to all participants 200. (B) The processed values ​​of the contributions of all explanatory variables are disclosed to all participants 200. The processed values ​​are relative values ​​such as rankings, for example. (C) For each participant 200, the processed values ​​of the contributions of the explanatory variables included in the experimental data provided by that participant 200 are disclosed. Furthermore, the processed values ​​of the contributions of the explanatory variables whose names are kept confidential are disclosed. (D) For each participant 200, the processed value of the contribution level is disclosed only for the explanatory variables included in the experimental data provided by that participant 200. (2): Each participant 200 provides experimental data to be used as training data (first training data). The recommendation unit 20 generates a first inference model from the experimental data. (3): Obtain input objective variables from participants 200 who wish to obtain optimization data. (4): Using the first inference model generated in (2), generate optimization data corresponding to the input objective variables in (3). (5) Calculate the contribution of each explanatory variable to the optimization data. (6): The contribution value of each explanatory variable is processed according to the type and scope of the explanatory variables to be disclosed set in (1). For example, the contribution value may be processed by at least one of the following processes (E) to (G). (E) Ranking process: The contributions are sorted in descending order, and ordinal numbers are set to 1, 2, 3, etc., starting from the top. (F) Deviation processing: Each contribution is fitted to a normal distribution, which may have a mean of 50 and a deviation of 10. (G) Percentile processing: Based on a preset standard, the magnitude of the contribution is converted into a percentile ranging from 0 to 100. The magnitude of the contribution corresponding to percentiles 0 and 100 may be set, for example, based on past actual values. For example, of the contributions calculated in the past, the maximum value may correspond to percentile 100, and the minimum value may correspond to percentile 0. (7): The names of the explanatory variables set in (1) that are not to be disclosed are anonymized. The anonymization of the names may be a process of replacing the names of the explanatory variables with blanks or with predefined common names.

[0139] FIG. 19 shows an example of information about the contribution of each explanatory variable calculated by the contribution calculation unit 60. FIG. 19 shows information presented to one of the participants 200. In the example of FIG. 19, the contribution of each explanatory variable is converted into an ordinal number indicating a rank. The greater the contribution of an explanatory variable, the lower the rank it is given. Furthermore, the names of some explanatory variables are kept secret (masked data).

[0140] Fig. 20 shows another example of information on the contribution of each explanatory variable calculated by the contribution calculation unit 60. In the example of Fig. 20, the contribution of each explanatory variable is replaced with a percentile ranging from 0 to 100. For example, the greater the contribution, the closer the percentile is to "100," and the smaller the contribution, the closer it is to "0."

[0141] The planning unit 72 recommends the creation of explanatory variables to be included in the additionally acquired experimental data or the deletion of explanatory variables unnecessary for the experimental data, based on the contribution of each explanatory variable calculated by the contribution calculation unit 60. This reduces the uncertainty of predictions in the optimization data. By controlling the selection of appropriate explanatory variables in the additionally executed experiments, the planning unit 72 can acquire experimental data that can contribute to reducing the uncertainty of predictions in the first inference model used to create the optimization data.

[0142] An example of the operation of the planning unit 72 is as follows. (1): Set a standard value for the evaluation value of the accuracy of the optimization data creation for the inference model. The evaluation value will be described later in (7). (2): Set the number or distance of the target subset. (3): Set the input objective variable. The input objective variable may be set as a value or a range. For example, the value of the objective variable "heat resistance" is set to "150°C or more, 200°C or less." The input objective variable may be input by the participant 200. (4): Create a list of each explanatory variable included in the optimization data output by the first inference model according to the input objective variable, sorted in descending order by contribution. In other words, sort each explanatory variable in descending order by contribution. (5): From the list of explanatory variables, a set number of explanatory variables are selected in order of their contribution. The number of selections may be set in advance by the participant 200, etc. An inference model is created using only the selected explanatory variables in the entire first training data. (6): From the first training data, training data in the vicinity of the input objective variable is extracted based on the criterion set in (2). For example, all training data for which the distance of the objective variable from the input objective variable is equal to or less than the distance set in (2) may be extracted. Alternatively, the number of training data set in (2) may be extracted in order of the shortest distance of the objective variable from the input objective variable. (7): The inference model created in (5) is applied to each of the training data extracted in (6). For example, the explanatory variables of each training data are input into the inference model to calculate an inference objective variable. An evaluation value of the inference model is calculated based on the difference between the calculated inference objective variable and the objective variable included in the training data. For example, the smaller the difference, the higher the evaluation value. The evaluation value may be the difference between the value of the inference objective variable and the value of the objective variable of the training data, or an evaluation value for the prediction such as a coefficient of determination. The coefficient of determination may be, for example, the square of the correlation coefficient between the value of the inference objective variable and the value of the objective variable of the training data. (8): Compare the evaluation value calculated in (7) with the reference value set in (1). If the evaluation value of the inference model is lower than the reference value, the number of selections in (5) is increased by one, an inference model is created, and the evaluation value is calculated. The number of selections is increased by one, and the creation and evaluation of the inference model is repeated until the evaluation value is equal to or greater than the reference value. When the evaluation value is equal to or greater than the reference value, explanatory variables not selected in (5) are subject to an overprovision warning. In (5), if the evaluation value of the inference model created with all explanatory variables selected is smaller than the reference value, a warning about insufficient explanatory variables may be displayed. If the evaluation value of the inference model is higher than the reference value, the number of selections in (5) is reduced by one, an inference model is created, and the evaluation value is calculated. The number of selections is reduced by one, and the creation and evaluation of the inference model are repeated until the evaluation value becomes smaller than the reference value. Explanatory variables that were not selected in (5) in the processing cycle immediately before the processing cycle in which the evaluation value became smaller than the reference value may be subject to an overprovision warning.

[0143] The plan formulation unit 72 may generate a data collection plan to acquire additional experimental data excluding explanatory variables that are the subject of an overprovision warning. The plan formulation unit 72 may generate an experimental plan for the experimental device in which fixed values ​​are set for explanatory variables that are the subject of an overprovision warning. When a warning for insufficient explanatory variables has occurred, the plan formulation unit 72 may generate an experimental plan for the experimental device that includes new explanatory variables.

[0144] 21 is a diagram illustrating an example of the operation of the plan development unit 72. As shown in FIG. 19, the plan development unit 72 in this example determines whether or not each explanatory variable ranked by contribution is subject to an overprovision warning. In the example of FIG. 21, explanatory variables ranked 84th and above are subject to an overprovision warning. The plan development unit 72 may create a data collection plan so that additional experimental data showing diverse values ​​can be acquired for explanatory variables that are not subject to an overprovision warning.

[0145] FIG. 22 is a diagram illustrating another example of the operation of the planning unit 72. In this example, the contribution level is converted into a percentile, as shown in FIG. 20 . In this example, the evaluation value of the inference model using all explanatory variables does not reach the reference value. In this case, the planning unit 72 may issue an insufficient provision warning for all explanatory variables. The planning unit 72 may create a data collection plan so that additional experimental data can be acquired. The planning unit 72 may create a warning as part of the data collection plan, indicating that experimental data is insufficient and should be added. The planning unit 72 may issue a warning indicating that experimental data including a new explanatory variable should be added, or may specify an explanatory variable with a high contribution level for existing explanatory variables and issue a warning indicating that experimental data including the explanatory variable should be added. An explanatory variable with a high contribution level may be, for example, one whose index, such as a contribution level value or ranking, exceeds a set threshold.

[0146] Fig. 23 is a diagram showing another example configuration of the information processing device 100. The information processing device 100 of this example has a data acquisition unit 76 instead of the control unit 74 in the configuration described in Fig. 18. The configuration other than the data acquisition unit 76 may be the same as the example described in Fig. 18.

[0147] The planning unit 72 in this example outputs the formulated experimental data collection plan to the experimental equipment and causes the experimental equipment to execute an experiment in accordance with the experimental data collection plan. The experimental data collection plan may be the same as the example described in FIG. 18. The planning unit 72 may output control data for each device of the experimental equipment to the experimental equipment in order to conduct an experiment in accordance with the experimental data collection plan. In response to a request from a terminal of any of the participants 200, the planning unit 72 may transmit the experimental data collection plan or the control data to the terminal. For example, the participant 200 may request the experimental data collection plan or the control data from the information processing device 100 by selecting a data request button displayed on the terminal.

[0148] The data acquisition unit 76 acquires additional experimental data from the experimental equipment that has operated based on the data collection plan or control data. The additional experimental data is similar to the example described in Figure 18. With this configuration, it is possible to acquire additional experimental data with a high degree of contribution, and to easily generate a more accurate inference model.

[0149] 24 is a diagram showing an example of a computer 2200 in which the information processing device 100 according to an embodiment of the present invention may be embodied, in whole or in part. A program installed on the computer 2200 can cause the computer 2200 to perform operations associated with the information processing device 100 according to an embodiment of the present invention, or to function as one or more sections of the information processing device 100, or to execute the operations or one or more sections, or to execute each step of the information processing method of the present invention (see FIGS. 1 to 23). The program can be executed by the CPU 2212 to cause the computer 2200 to perform specific operations associated with some or all of the blocks in the flowcharts ( FIGS. 12 to 17 ) and block diagrams ( FIGS. 4 , 18 , and 23 ) described herein.

[0150] A computer 2200 according to one embodiment of the present invention includes a CPU 2212, a RAM 2214, a graphics controller 2216, and a display device 2218. The CPU 2212, the RAM 2214, the graphics controller 2216, and the display device 2218 are interconnected by a host controller 2210. The computer 2200 further includes input / output units such as a communication interface 2222, a hard disk drive 2224, a DVD-ROM drive 2226, and an IC card drive. The communication interface 2222, the hard disk drive 2224, the DVD-ROM drive 2226, and the IC card drive are connected to the host controller 2210 via an input / output controller 2220. The computer further includes legacy input / output units such as a ROM 2230 and a keyboard 2242. The ROM 2230, the keyboard 2242, and the like are connected to the input / output controller 2220 via an input / output chip 2240.

[0151] The CPU 2212 controls each unit by operating in accordance with programs stored in the ROM 2230 and the RAM 2214. The graphics controller 2216 acquires image data generated by the CPU 2212 into a frame buffer or the like provided in the RAM 2214 or into the RAM 2214, thereby causing the image data to be displayed on the display device 2218.

[0152] The communication interface 2222 communicates with other electronic devices via a network. The hard disk drive 2224 stores programs and data used by the CPU 2212 in the computer 2200. The DVD-ROM drive 2226 reads programs or data from the DVD-ROM 2201 and provides the read programs or data to the hard disk drive 2224 via the RAM 2214. The IC card drive reads programs and data from an IC card or writes programs and data to an IC card.

[0153] The ROM 2230 stores a boot program or the like that is executed by the computer 2200 upon activation, or a program that depends on the hardware of the computer 2200. The input / output chip 2240 may connect various input / output units to the input / output controller 2220 via a parallel port, a serial port, a keyboard port, a mouse port, or the like.

[0154] The programs are provided by a computer-readable medium such as a DVD-ROM 2201 or an IC card. The programs are read from the computer-readable medium, installed in the hard disk drive 2224, RAM 2214, or ROM 2230, which are also examples of computer-readable media, and executed by the CPU 2212. Information processing described in these programs is read by the computer 2200, and brings about cooperation between the programs and the various types of hardware resources described above. An apparatus or method may be configured by realizing information manipulation or processing according to the use of the computer 2200.

[0155] For example, when communication is performed between the computer 2200 and an external device, the CPU 2212 may execute a communication program loaded into the RAM 2214 and instruct the communication interface 2222 to perform communication processing based on the processing described in the communication program. Under the control of the CPU 2212, the communication interface 2222 reads transmission data stored in a transmission buffer processing area provided in the RAM 2214, the hard disk drive 2224, the DVD-ROM 2201, or a recording medium such as an IC card, and transmits the read transmission data to the network, or writes reception data received from the network to a reception buffer processing area or the like provided on the recording medium.

[0156] The CPU 2212 may read all or a necessary portion of a file or database stored on an external recording medium such as the hard disk drive 2224, the DVD-ROM drive 2226 (DVD-ROM 2201), an IC card, etc. into the RAM 2214. The CPU 2212 may perform various types of processing on the data on the RAM 2214. The CPU 2212 may then write the processed data back to the external recording medium.

[0157] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and processed. The CPU 2212 may perform various types of processing on data read from the RAM 2214, including various types of operations, information processing, conditional decisions, conditional branches, unconditional branches, information search or replacement, etc., specified by the instruction sequences of the programs described in this disclosure. The CPU 2212 may write the results back to the RAM 2214.

[0158] CPU 2212 may search for information in a file, database, etc. in the recording medium. For example, if a plurality of entries each having an attribute value of a first attribute associated with an attribute value of a second attribute are stored in the recording medium, CPU 2212 may search the plurality of entries for an entry that matches a condition specified by the attribute value of the first attribute, read the attribute value of the second attribute stored in the entry, and by reading the second attribute value, obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.

[0159] The above-described programs or software modules may be stored on the computer 2200 or in a computer-readable medium of the computer 2200. A recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can be used as the computer-readable medium. The programs may be provided to the computer 2200 by the recording medium.

[0160] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications and improvements can be made to the above embodiments. It is clear from the claims that such modifications and improvements can also be included within the technical scope of the present invention.

[0161] It should be noted that the execution order of each process, such as operations, procedures, steps, and stages, in the devices, systems, programs, and methods shown in the claims, specifications, and drawings is not specifically stated as "before," "prior to," etc., and that the processes can be performed in any order unless the output of a previous process is used in a subsequent process. Even if the operational flow in the claims, specifications, and drawings is described using "first," "next," etc. for convenience, this does not mean that the processes must be performed in this order. [Explanation of symbols]

[0162] 10 Identification unit, 20 Recommendation unit, 22 First inference model, 24 Second inference model, 30 Input unit, 40 Output unit, 50 Memory unit, 60 Contribution calculation unit, 70 Selection unit, 72 Planning unit, 74 Control unit, 76 Data acquisition unit, 100 Information processing device, 200 Participant, 202 Individual training data, 203 Individual explanatory variable, 204 Training explanatory variable, 205 Individual objective variable, 206 Training objective variable, 208 Individual training data, 210 First training data, 212 Data space, 214 Data space, 2 20···Second training data, 300···Terminal, 400···Network, 500···Information provision system, 2200···Computer, 2201···DVD-ROM, 2210···Host controller, 2212···CPU, 2214···RAM, 2216···Graphics controller, 2218···Display device, 2220···Input / output controller, 2222···Communication interface, 2224···Hard disk drive, 2226···DVD-ROM drive, 2230···ROM, 2240···Input / output chip, 2242···Keyboard

Claims

1. a recommendation unit having a first inference model that performs machine learning on first teacher data having teacher explanatory variables and teacher objective variables provided by each of a plurality of participants, and that outputs a first predicted value of an input objective variable corresponding to an input explanatory variable when an input explanatory variable is input; When a target value of the input objective variable is input to the recommendation unit, the recommendation unit outputs optimization data which is an explanatory variable corresponding to the target value; a contribution calculation unit that calculates a contribution of the first teacher data corresponding to at least one of the participants to the optimization data based on the optimization data and the first teacher data; Information processing device.

2. When the target value is input, the recommendation unit generates candidates for the optimization data corresponding to the target value; the first inference model outputs the first predicted value corresponding to the candidate generated by the recommendation unit; the recommendation unit outputs one of the candidates as the optimized data based on the first predicted value and the target value. The information processing device according to claim 1 .

3. the recommendation unit includes a second inference model that performs machine learning on second teacher data obtained by excluding, from the first teacher data, individual explanatory variables that are the teacher explanatory variables corresponding to at least one of the plurality of participants and individual objective variables that are the teacher objective variables corresponding to the at least one of the participants, and that outputs a second predicted value of the input objective variable when the input explanatory variable is input; the contribution calculation unit calculates the contribution based on the first predicted value and the second predicted value. The information processing device according to claim 1 .

4. the first inference model, when receiving the optimization data, outputs the first predicted value corresponding to the optimization data; the second inference model, when receiving the optimization data, outputs the second predicted value corresponding to the optimization data; the contribution calculation unit calculates the contribution based on the first predicted value corresponding to the optimized data and the second predicted value corresponding to the optimized data. The information processing device according to claim 3 .

5. the recommendation unit performs machine learning on the second training data, from which the individual explanatory variables and the individual objective variables corresponding to one of the plurality of participants have been excluded, for each of the plurality of participants; the second inference model outputs the second predicted value for each of the plurality of participants; the contribution calculation unit calculates the contribution of each of the plurality of participants based on the first predicted value and the second predicted value of each of the plurality of participants; a selection unit that selects at least one of the participants based on the degree of contribution, and prompts the participant to input additional individual teacher data, the individual teacher data having the individual explanatory variables and the individual objective variables, to the recommendation unit; The information processing device according to claim 3 .

6. the recommendation unit performs machine learning on the second teacher data from which the individual explanatory variables and the individual objective variables corresponding to at least two of the plurality of participants have been excluded; the contribution calculation unit calculates the contributions of the at least two parties based on the first predicted value and the second predicted value. The information processing device according to claim 3 .

7. Weights are assigned in advance to the plurality of input objective variables; the contribution calculation unit calculates the contribution based on the first predicted value, the second predicted value, and the weighting. The information processing device according to claim 3 .

8. The teacher explanatory variables have a plurality of types of parameters, and at least one type of the parameters is provided by at least one of the participants; the contribution calculation unit calculates the contribution of the at least one parameter to the optimization data based on the optimization data and the first teacher data. The information processing device according to claim 1 .

9. A plurality of said parameters are provided by at least one of said participants; the contribution calculation unit calculates a contribution of the at least one species to the optimization data for each of the plurality of species, and calculates the contribution of at least one of the participants based on the calculated contribution of each of the plurality of species. The information processing device according to claim 8 .

10. the first teacher data includes at least one piece of experimental data; a planning unit that plans and outputs a data collection plan for collecting additional experimental data based on the contribution calculated by the contribution calculation unit. The information processing device according to claim 1 .

11. a control unit for acquiring the additional experimental data by controlling an experimental device based on the data collection plan; The information processing device according to claim 10.

12. The data acquisition unit further acquires the additional experimental data from the experimental equipment that has operated based on the data collection plan. The information processing device according to claim 10.

13. a first inference model generation step in which a recommendation unit generates a first inference model by machine learning first teacher data having teacher explanatory variables and teacher objective variables provided by each of a plurality of participants, the first inference model outputting a first predicted value, which is a predicted value of an input objective variable corresponding to an input explanatory variable, when an input explanatory variable is input; When a target value of an input objective variable is input to the recommendation unit, the recommendation unit outputs optimization data which is an explanatory variable corresponding to the target value; The method further includes a contribution calculation step of calculating a contribution of the first teacher data corresponding to at least one of the participants to the optimization data based on the optimization data and the first teacher data by a contribution calculation unit. Information processing methods.

14. An information processing program for causing a computer to function as the information processing device according to any one of claims 1 to 6 and 10 to 12.