TEST PRODUCTION CONDITION PROPOSAL SYSTEM AND TEST PRODUCTION CONDITION PROPOSAL PROCEDURE
Patent Information
- Application Number
- DE112023004076
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-07-25
- Publication Date
- 2025-08-21
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a system and a method for proposing test production conditions for (a) material(s) to a material developer. STATE OF THE ART
[0002] In the field of materials science for materials research and development, a method called "Materials Informatics (MI)" for efficiently predicting the physical properties, structures, etc., of materials using information technology (informatics) such as statistical analysis and machine learning is now widely used. Regarding materials research and development using this materials informatics, for example, a technology is known from PTL 1. PTL 1 discloses a system for estimating manufacturing conditions for substances with optimal physical properties and structures from a data set including manufacturing conditions for each of a plurality of sample substances and substance information indicating physical properties and structures of the respective substances. LITERATURE LISTPATENT LITERATURE
[0003] PTL 1: WO 2021 / 044913 SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0004] At the materials development site(s) where materials informatics is used, various types of evaluation tests are usually conducted to actually measure the properties of a material(s) being developed, and data indicating the actual measurement results are obtained (hereinafter referred to as "measured property data"). Then, various types of learned machine learning models are constructed by inputting the obtained measured property data into a computer and learning them. The condition(s) for test production of the material (hereinafter referred to as "test production condition(s)") are estimated using the constructed learned machine learning model.Under these circumstances, the accuracy of the computer to estimate the test production condition(s) (hereinafter referred to as "estimation accuracy") is determined by the accuracy of the learned machine learning model used for processing to estimate the test production condition(s). Then, the accuracy of the learned machine learning models is generally determined according to, for example, the number of individual measured property data learned when constructing the machine learning model (hereinafter referred to as the "number of samples") or the degree of variation in their distribution.
[0005] Under these circumstances, the measured property data input to the computer to construct the machine learning model(s) (hereinafter also referred to as "learning data") are generally characterized by an extremely large number of explanatory variables, while the number of samples is often small. Therefore, if the number of individual learning data is small or if there is a bias in the distribution of the learning data, the computer cannot construct highly accurate machine learning models. Therefore, in such a case, there is a concern that sufficient estimation accuracy may not be obtained even if the computer is instructed to perform the processing to estimate the test production condition(s) using the constructed machine learning models.
[0006] In view of the problem described above, it is an object of the present invention to provide a technology capable of accurately proposing an optimal test production condition for the material even when the number of individual learning data used for constructing a machine learning model(s) is small or when there is a distortion in the distribution of the learning data. MEANS TO SOLVE THE PROBLEMS
[0007] A test production condition suggestion system according to the present invention proposes a test production condition for a material to a material developer and includes a regression model creation processing unit and a test production condition suggestion processing unit. The regression model creation processing unit executes regression model creation processing on measured property data indicating an actual measurement result of properties of the material. The test production condition suggestion processing unit searches for an optimal test production condition for the material using the created regression model and executes test production condition suggestion processing based on a search result.The regression model creation processing includes: processing for calculating a weight reference, which is a reference for weighting measured property data; and processing for performing weighting of the measured property data based on the calculated weight reference.
[0008] Furthermore, a test production condition suggestion method according to the present invention is a method for suggesting a test production condition for a material to a material developer using a computer. This test production condition suggestion method causes the computer to execute regression model creation processing and test production condition suggestion processing. The regression model creation processing represents processing for creating a regression model for measured property data indicating an actual measurement result of the properties of the material. The test production condition suggestion processing represents processing for searching for an optimal test production condition for the material using the created regression model and proposing a test production condition for the material based on a search result.The regression model creation processing includes: processing for calculating a weight reference, which is a reference for weighting measured property data; and processing for performing weighting of the measured property data based on the calculated weight reference.
[0009] Incidentally, the problems and their solutions disclosed in this application are explained in more detail in the DESCRIPTION OF THE EMBODIMENTS section and in the descriptions of the drawings. ADVANTAGES OF THE INVENTION
[0010] According to the present invention, the optimal test production condition for the material can be suggested with excellent accuracy even when the number of individual learning data used for constructing the machine learning model(s) is small or when there is a distortion in the distribution of the learning data. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 is a schematic diagram illustrating the structure of a test production condition suggestion system according to an embodiment of the present invention; Fig. 2 is a diagram illustrating the functional blocks of the test production condition suggestion system according to an embodiment of the present invention; Fig. 3 is a flowchart illustrating a flow of the entire processing of the test production condition suggestion system according to an embodiment of the present invention; Fig. 4 is a flowchart illustrating the details of preprocessing of measured property data; Fig. Figure 5 is a flowchart showing the details of regression model building processing; Fig. Figure 6 is a figure showing a comparison of the estimation accuracy before and after weighting; and Fig. Figure 7 is a flowchart showing the details of test production condition proposal processing. DESCRIPTION OF EMBODIMENTS
[0011] This embodiment is described in detail below. Fig. 1 is a schematic diagram showing the structure of a test production condition suggestion system according to an embodiment of the present invention. Fig. The test production condition suggestion system 1 shown in FIG. 1 is configured to optimize each of the various test production conditions to be considered in the test production of a material, such as material compositions and firing conditions, and to propose the results to a materials developer who is a user of this system. This optimization of the test production conditions is performed using various types of machine learning algorithms based on measured property data that indicate the actual measurement results of the properties of the relevant material.Furthermore, various types of regression models such as Gaussian process regression (GPR) and linear regression, regression trees (including one case by an ensemble method), regression by a neural network (neural network regression), support vector regression (SVR), logistic regression, LASSO regression (Least Absolute Shrinkage and Selection Operator Regression) are used as prediction models.
[0012] As in Fig. As shown in Figure 1, when the user of the test production condition suggestion system 1 performs test production of a material under the test production condition(s) suggested by the test production condition suggestion system 1 and evaluates the properties of the material that is a product of the test production, measured property data indicating an evaluation value of the properties of that product are newly generated. After causing the test production condition suggestion system 1 to learn these measured property data as new learning data, the test production condition suggestion system 1 proposes more optimized test production condition(s) to the user. Each time the user repeats this workflow, the test production condition suggestion system 1 according to this embodiment can propose the test production condition(s) with a better predicted value of the properties.Incidentally, the test production condition suggestion system may include a function for performing the test production of the material and a function for actually measuring the properties of the material that is the subject of the test production, and may be integrally configured with these functions.
[0013] The test production condition suggestion system 1 according to this embodiment is realized by a general-purpose computer device as shown in Fig. 1. In the following explanation, it is assumed that the test production condition suggestion system 1 is implemented by a general-purpose computer device including one or more processor devices, one or more storage devices, one or more input / output devices, and wired or wireless communication lines connecting these components (both are not shown in the drawing).
[0014] This computer device is installed, for example, as a terminal within a laboratory and is connected via a communication network such as the Internet 400 and leased lines to various other terminals installed inside and outside the laboratory and used by each of the users (hereinafter referred to as "user terminals"), as well as to other devices such as server device(s). Incidentally, the computer device and the Internet 400 are connected by wires via known communication devices (not shown in the drawing), but they may also be connected wirelessly.
[0015] Next, an explanation of the various functions of the test production condition suggestion system 1 is provided by Fig. 2 is referred to. Fig. 2 is a diagram illustrating the functional blocks of the test production condition suggestion system according to an embodiment of the present invention. Incidentally, the respective blocks explained below denote functional unit blocks, but not hardware unit components. The test production condition suggestion system 1 according to this embodiment is configured as shown in Fig. 2, includes a control unit 11, a storage unit 12, a user interface unit 13 and a communication unit 14.
[0016] The control unit 11 performs various types of data processing based on the user inputs acquired by the user interface unit 13, the data acquired by the communication unit 14, and the programs and data stored in the storage unit 12. The control unit 11 also serves as an interface for the user interface unit 13, the communication unit 14, and the storage unit 12.
[0017] The control unit 11 includes respective functional blocks of a measured property data preprocessing unit 111, a regression model creation processing unit 112, and a test production condition suggestion processing unit 113. The control unit 11 is configured, for example, using processor devices such as a CPU (Central Processing Unit) and various types of coprocessors (hereinafter also referred to simply as "processors"), and can implement these functional blocks by executing specific programs. Incidentally, the control unit 11 can also be configured using, for example, logic circuits such as an FPGA (Field Programmable Gate Array) instead of the processors. Further, the control unit 11 can be configured by a combination of the processors and the logic circuits.
[0018] The programs to be executed by the control unit 11 can be installed from program source(s). The program source(s) can be, for example, a recording medium or the like that can be read by program distribution computer(s) or computer(s). Furthermore, the programs executed by the control unit 11 can be configured by a device driver, an operating system, various types of application programs located at an upper layer thereof, and a library that provides functions common to these programs. Furthermore, two or more programs can be implemented as one program, and one program can be implemented as two or more programs.
[0019] The measured property data preprocessing unit 111 applies preprocessing to the measured property data in a so-called raw data state immediately after its recording. This processing performed by the measured property data preprocessing unit 111 is hereinafter referred to as measured property data preprocessing.
[0020] The regression model creation processing unit 112 executes processing to create a regression model(s) with respect to the measured characteristic data to which the preprocessing has been applied. This processing executed by the regression model creation processing unit 112 is referred to as regression model creation processing.
[0021] The test production condition suggestion processing unit 113 performs processing to search the measured property data to which preprocessing has been applied for optimal test production condition(s) for the material using the regression model creation processing unit 112, and suggests the test production condition(s) for the material to the user based on the search result. This processing by the test production condition suggestion processing unit 113 is hereinafter referred to as test production condition suggestion processing.
[0022] The specific content of these processing sequences will be described later.
[0023] The storage unit 12 is configured, for example, using storage devices such as RAM(s) and flash memory(s), and stores programs for providing various types of processing instructions to the control unit 11 and data indicating various types of information to be used for the processing performed by the control unit 11. For example, the measured characteristic data to which preprocessing has been applied by the measured characteristic data preprocessing unit 111 (hereinafter referred to as "preprocessed data"), data indicating regression models generated by the regression model generation processing unit 112, etc. are stored in the storage unit 12.The control unit 11 can implement the respective functional blocks of the aforementioned measured property data preprocessing unit 111, the regression model creation processing unit 112, and the test production condition suggestion processing unit 113 by reading / writing these individual pieces of information from / to the storage unit 12.
[0024] In addition to accepting user inputs, the user interface unit 13 is responsible for processing related to the user interface, such as image display and sound output. The user interface unit 13 has respective functional blocks of an input unit 131 and an output unit 132. The input unit 131 recognizes various types of user operations. The input unit 131 is configured, for example, using a keyboard, a pointing device, a touch panel, etc. The output unit 132 performs, for example, the screen display and sound output for the user. The output unit 132 is configured, for example, using a liquid crystal display and a touchscreen.
[0025] The communication unit 14 is responsible for communication processing performed over the Internet 400 with other devices such as each user's user terminal(s), the server device(s), and so on. The communication unit 14 is configured, for example, using a NIC (Network Interface Card) and an HBA (Host Bus Adapter).
[0026] This embodiment has been explained by describing that the respective functions of the test production condition suggestion system 1 are integrally implemented by one computer device. However, these respective functions may be implemented by a plurality of interconnected computer devices or server devices. The test production condition suggestion system 1 may also be configured by including a general-purpose computer device, such as a laptop, and a web browser installed in this general-purpose computer device, or it may be configured by including a web server and various types of portable devices.
[0027] Furthermore, the explanations of each function are examples, and a plurality of functions can be combined into one function or a function can be divided into a plurality of functions.
[0028] Next, an explanation of an entire flow of the test production condition suggestion system 1 will be given with reference to Fig. 3 provided. Fig. 3 is a flowchart illustrating the flow of the entire processing of the test production condition suggestion system according to an embodiment of the present invention. Incidentally, in the following explanation, there is a case where the processing is explained by referring to each aforementioned function or program as the subject matter; however, the processing explained by referring to the function or program as the subject matter may be processing performed by a processor or a device including that processor.
[0029] In step S310, the control unit 11 causes the measured property data preprocessing unit 111 to perform the preprocessing of measured property data. Accordingly, the preprocessing is applied to the measured property data, which then becomes the preprocessed data, so that it becomes possible to perform any subsequent processing normally. Incidentally, the details of the preprocessing of measured property data performed in step S310 will be explained later with reference to a flowchart in Fig. 4. When the preprocessing of measured property data is completed, the control unit 11 proceeds to step S320.
[0030] In step S320, the control unit 11 causes the regression model creation processing unit 112 to execute the regression model creation processing. Accordingly, a regression model is created with respect to the preprocessed data. Incidentally, the details of the regression model creation processing performed in step S320 will be explained later with reference to a flowchart in Fig. 5. When the regression model creation processing is completed, the control unit 11 proceeds to step S330.
[0031] In step S330, the control unit 11 executes the regression model evaluation processing. This regression model evaluation processing is used to evaluate the generalization performance, which is an index indicating the prediction accuracy of the relevant regression model with respect to each of the plurality of regression models created as a result of the respective processing sequences before and in step S320. This evaluation is performed, for example, through cross-validation with other regression models. The evaluation result is visualized, for example, by a chart such as a scatter plot or a box-and-whisker plot. As a result, the user receives a suggestion for the test production condition(s) based on the regression model with good generalization performance. When the regression model evaluation processing is completed, the control unit 11 proceeds to step S340.
[0032] In step S340, the control unit 11 causes the test production condition suggestion processing unit 113 to execute the test production condition suggestion processing. With this test production condition suggestion processing unit, the user of the test production condition suggestion system 1 can modify the test production condition(s) for the material suggested by the test production condition suggestion system 1 as needed to make them more preferable. The control unit 11 determines a predicted value for the properties of the material when the material is to be test-produced under the user-modified test production condition(s) by applying it to the selected regression model and presenting the predicted value to the user. More specifically, the user can perform this work interactively to modify the test production condition(s) while verifying the predicted value.Accordingly, the test production condition suggestion system 1 is configured as a system capable of incorporating the knowledge of the user, who is a developer of the relevant material, into the test production condition(s) suggested to the user for the material. Incidentally, the details of the test production condition suggestion processing performed in step S340 will be explained later with reference to a flowchart in FIG. Fig. 7. When the test production condition suggestion processing is completed, the control unit 11 ends the process shown in the flowchart in Fig. 3 processing shown once.
[0033] The test production condition suggestion system 1 according to this embodiment performs each processing in steps S310 to S340 in Fig. 3 and suggests good test production condition(s) to the user. More specifically, in the test production condition suggestion system 1 according to this embodiment, in step S310, the required preprocessing is automatically applied to the measured property data in a raw data state. Thus, each processing in steps S320 to S340 can be performed without the need for complicated manual preprocessing of measured property data. Furthermore, with the test production condition suggestion system 1 according to this embodiment, the user can select a regression model with good generalization performance. Therefore, the test production condition suggestion system 1 can suggest to the user the test production condition(s) for which the predicted property value is good.
[0034] Furthermore, as already mentioned with reference to Fig. 1, the user of the test production condition suggestion system 1, after performing the test production of the material under the test production conditions suggested by the test production condition suggestion system 1, actually measuring the properties of the product, and causing the test production condition suggestion system 1 to learn data indicating the actual measurement result as new measured property data, can cause the test production condition suggestion system 1 to perform each processing of steps S310 to S340 in Fig. 3 again. In this case, the test production condition suggestion system 1 may suggest more optimized test production condition(s) to the user. More specifically, the test production condition suggestion system 1 according to this embodiment may, each time it executes steps S310 to S340 in Fig. 3 with respect to the same test production object, suggesting test production condition(s) with a better predicted value of the properties.
[0035] Fig. Figure 4 is a flowchart showing the details of preprocessing of measured property data.
[0036] In step S410, the control unit 11 causes the measured property data preprocessing unit 111 to receive an input of measured property data from the user via the input unit 131 or the communication unit 14. The measured property data input to the test production condition suggestion system 1 may be, for example, category data, continuous data, or discrete data. Further, a specific data format of the measured property data to be input to the test production condition suggestion system 1 may be determined as appropriate. When the processing in step S410 is completed, the control unit 11 proceeds to step S420.
[0037] In step S420, the control unit 11 causes the measured property data preprocessing unit 111 to specify a variable type with respect to the measured property data whose input was received from the user in step S410. Under these circumstances, either an explanatory variable(s) or a target variable(s) are specified. The explanatory variable(s) is / are a variable that serves as a basis for determining a predicted value of the properties. In this embodiment, the material composition, firing conditions, etc., which constitute the test production conditions, correspond to the explanatory variables. Furthermore, the target variable(s) is / are a variable that indicates a property value(s) of the material to be test-produced, which becomes a prediction object.An example of concrete processing in step S420 is that the explanatory variable can be set as the default, and a setting operation can be accepted by the user who wishes to change it to the target variable. When the processing in step S420 is completed, the control unit 11 proceeds to step S430.
[0038] In step S430, the control unit 11 causes the measured characteristic data preprocessing unit 111 to judge whether or not there is an abnormal value with respect to the measured characteristic data for which either the explanatory variable or the target variable was set in step S420. This processing for judging whether or not there is an abnormal value is performed, for example, by displaying the measured characteristic data as a histogram and judging whether or not there is an outlier outside the range of an average value ±2σ, and then judging whether or not there is a data input error or a failure of an evaluation test device in generating the measured characteristic data, in which the presence of the outlier is determined.On the one hand, if it is determined that the measured characteristic data contains an abnormal value, the control unit 11 deletes this abnormal value and recognizes it as a missing value, and proceeds to step S440. Further, if the type of abnormal value included in the measured characteristic data is one of the variables, that is, a target variable in which explanatory variables are completely duplicated and there is a standard with different characteristics, the control unit 11 deletes (a) sample(s) related to this abnormal value and proceeds to step S440. In such a case where all explanatory variables are duplicated and the standard with different characteristics is treated as an abnormal value, the standard itself must be deleted, unlike the case where part of the explanatory variable is an abnormal value, for example, due to an input error.On the other hand, if it is determined that no abnormal value is included in the measured characteristic data, the control unit 11 proceeds directly to step S440.
[0039] Incidentally, if all explanatory variables are duplicated and there is a standard with different properties, as described above, duplication may be considered reasonable, and it may sometimes be desirable to retain all explanatory variables. In such a case, the test production condition suggestion system 1 according to this embodiment may skip the processing in step S430.
[0040] In step S440, the control unit 11 causes the measured characteristic data preprocessing unit 111 to judge whether or not any missing value(s) are / are missing with respect to the measured characteristic data. This judgment is made because the measured characteristic data sometimes includes the missing value(s) in advance. On the one hand, if it is judged that the missing value(s) are / are included in the measured characteristic data, the control unit 11 supplements the missing value(s) and proceeds to step S450. The missing value supplement processing is performed, for example, using an average value, a median value, a minimum value, a maximum value, etc. of the measured characteristic data, excluding an abnormal value(s), as values to supplement the missing value(s). Further, the missing value(s) may be supplemented using linear interpolation.Incidentally, when the missing value(s) is / are supplemented in step S440, the measured property data preprocessing unit 111 may display the supplemented value, for example, in red font to facilitate identification of the supplemented value. Furthermore, the measured property data preprocessing unit 111 may delete standards itself, for example, without supplementing the relevant missing value in step S440. Furthermore, in the test production condition suggestion system 1 according to this embodiment, when a large proportion of missing explanatory variables is present, for example, when 50% or more of the data set is missing, the measured property data preprocessing unit 111 may delete the relevant explanatory variables itself. On the other hand, when it is judged that the missing value(s) is / are not included in the measured property data, the control unit 11 proceeds directly to step S450.
[0041] If the explanatory variable is not a continuous value but a categorical value, in step S450, the control unit 11 causes the measured characteristic data preprocessing unit 111 to perform coding processing on the relevant explanatory variable and convert it into numerical value data. The measured characteristic data preprocessing unit 111 performs this coding processing by, for example, referring to a record in a table indicating the correspondence relationship between the categorical data and the numerical value data, which is stored in the storage unit 12. When the processing in step S450 is completed, the control unit 11 proceeds to step S460.
[0042] In step S460, the control unit 11 causes the measured characteristic data preprocessing unit 111 to judge whether or not redundant explanatory variables are included in the measured characteristic data. This judgment is made based on whether a combination of explanatory variables with a correlation coefficient equal to or greater than a specific number, for example, 0.8 or more, can be extracted. On the one hand, if it is judged that redundant explanatory variables are included in the relevant measured characteristic data, the control unit 11 deletes one of the redundant explanatory variables and proceeds to step S470.Incidentally, in the test production condition suggestion system 1 according to this embodiment, the combination of explanatory variables whose correlation coefficient is equal to or greater than 0.8 is visualized to the user via the output unit 132, so that the explanatory variable to be deleted can be selected via the input unit 131. On the other hand, if it is judged that no redundant explanatory variable(s) is / are included in the relevant measured characteristic data, the control unit 11 proceeds directly to step S470.
[0043] In step S470, the control unit 11 causes the measured property data preprocessing unit 111 to perform standardization processing of the measured property data, if necessary. The standardization processing is to convert the scale of the measured property data so that an average = 0 and a standard deviation (variance) = 1 are obtained. When the processing in step S470 is completed, the control unit 11 stores the preprocessed data, that is, the measured property data to which the preprocessing in steps S410 to S470 was applied, in Fig. 4 was applied, in the storage unit 12 and terminates the preprocessing of measured property data shown in the flow chart in Fig. 4. Furthermore, the control unit 11 may cause the measured characteristic data preprocessing unit 111 to perform normalization processing on the measured characteristic data if necessary. In such a case, any of the subsequent processing may be performed without performing the standardization processing in S470.
[0044] Fig. Figure 5 is a flowchart showing the details of regression model building processing.
[0045] In step S510, the control unit 11 causes the regression model creation processing unit 112 to select (a) condition(s) for implementing cross-validation to evaluate a regression model to be created with respect to the preprocessed data. Incidentally, the test production condition suggestion system 1 according to this embodiment evaluates each regression model using K-fold cross-validation. Furthermore, K=10 is set as the implementation condition as the default value. In this case, the test production condition suggestion system 1 evaluates the regression model using 10-fold cross-validation. Incidentally, in the test production condition suggestion system 1 according to this embodiment, the user can also select the cross-validation implementation condition(s).More specifically, the regression model creation processing unit 112 may accept the cross-validation implementation condition(s) from the user via the input unit 131 or the communication unit 14. When the processing in step S510 is completed, the control unit 11 proceeds to step S520.
[0046] In step S520, the control unit 11 causes the regression model creation processing unit 112 to select a candidate regression model to be used as the prediction model for searching for the test production conditions. When this happens, the regression model creation processing unit 112 selects as a candidate regression model(s) for which it has accepted the user's selection processing via the input unit 131 or the communication unit 14.With the test production condition suggestion system 1 according to this embodiment, the user can select a variety of regression models as candidates from various types of regression models, such as Gaussian process regression, the aforementioned linear regression, regression trees (including a case by an ensemble method), neural network regression, support vector regression, logistic regression, and LASSO regression. When the processing in step S520 is completed, the control unit 11 proceeds to step S530.
[0047] In step S530, the control unit 11 causes the regression model creation processing unit 112 to execute processing for calculating a weight reference, which is a reference for the weight of the measured property data. The weight reference is calculated based on the difference between the target variable included in the measured property data and a target property indicating a target value of the properties of the material, or a statistical quantity indicating the rarity of the explanatory variable included in the measured property data (the details will be described later). Incidentally, a specific example of the statistical quantity indicating the rarity of the explanatory variable may include an occurrence probability of the explanatory variable satisfying specific condition(s). When the processing in step S530 is completed, the control unit 11 proceeds to step S540.
[0048] In step S540, the control unit 11 causes the regression model creation processing unit 112 to execute processing for weighting the measured characteristic data based on the weight reference calculated in step S530. Furthermore, in the test production condition suggestion system 1 according to this embodiment, the regression model creation processing unit 112 may directly weight the regression model based on the weight reference calculated in step S530.Incidentally, the processing executed in step S540 by the regression model creation processing unit 112 based on the result of step S530 is specifically processing for setting a loss function, which is a function serving as a learning index, oversampling processing for amplifying rare or highly significant measured feature data and adding the amplified data as learning data, and undersampling processing for deleting redundant or less significant measured feature data from the learning data (the details will be described later). While these processing sequences are being executed, the weight of the above-described learning data is adjusted as needed, even when the amount of learning data to be used to create a machine learning model is small or when there is a bias in the distribution of the learning data.As a result, it is possible to improve the estimation accuracy and propose the test production condition(s) for a good material with excellent accuracy. When the processing in step S540 is completed, the control unit 11 proceeds to step S550.
[0049] In step S550, the control unit 11 causes the regression model creation processing unit 112 to search for and set an optimal hyperparameter for each regression model for each method selected as a candidate in step S520. In the test production condition suggestion system 1 according to this embodiment, the regression model creation processing unit 112 automatically searches for all parameters for each regression model and automatically sets a parameter that realizes the best generalization performance of the relevant regression model when the parameter is set as a hyperparameter when creating the regression model. When the processing in step S550 is completed, the control unit 11 proceeds to step S560.
[0050] In step S560, the control unit 11 causes the regression model creation processing unit 112 to perform processing to generate a regression model(s) for each method for which the optimal hyperparameter has been determined. When the regression model creation processing unit 112 generates the regression model for each method, it performs processing to select a regression model with the highest generalization performance from all the generated regression models and determine a final regression model. When the processing in step S560 is completed, the control unit 11 ends the process shown in the flowchart in Fig. 5 shows the processing used to create the regression model.
[0051] Incidentally, as described above, the weighting reference is calculated based on the difference between the target variable contained in the measured property data and the target property indicating the target value of the properties of the material or a statistical quantity indicating the rarity of the explanatory variable contained in the measured property data.
[0052] Of the principles described above, on the one hand, the concrete content of the processing executed in step S530 when the weight reference is calculated based on the difference between the target variable included in the measured property data and the target property indicating the target value of the properties of the material will be described below as steps S532 to S534.
[0053] In step S532, the control unit 11 causes the regression model creation processing unit 112 to calculate the difference between the target variable contained in the measured property data and the target property indicating the target value of the material's properties. When the processing in step S532 is completed, the control unit 11 proceeds to step S534.
[0054] In step S534, the control unit 11 causes the regression model creation processing unit 112 to determine the weighting reference using a function of the difference between the target variable and the target property calculated in step S532. Specific examples of this function may be the absolute value or the square value of the difference between the target variable and the target property, which indicates the distance between the target variable and the target property. Incidentally, the reason why the weighting reference is determined based on the distance between the target variable and the target property is that regression problems are often solved in material development, and therefore, the number of occurrences of the target variable, which is generally used to determine the weighting reference in classification problems, cannot be utilized.Therefore, in the test production condition suggestion system 1 according to this embodiment, the weight reference based on the target variable, which is a continuous variable, is used. When the processing in step S534 is completed, the control unit 11 completes the processing for calculating the weight reference shown in step S530 and proceeds to step S540.
[0055] On the other hand, the concrete content of the processing executed in step S530 when the weight reference is calculated based on the statistical quantity indicating the rarity of the explanatory variable included in the measured characteristic data will be described below as steps S536 to S538.
[0056] In step S536, the control unit 11 causes the regression model creation processing unit 112 to determine the number of occurrences of the explanatory variable that takes a specific value or falls within a specified range. Specifically, for example, when the explanatory variable relates to a raw material ratio of various types of raw materials, the number of occurrences of the explanatory variable that takes a value greater than 0 is calculated; in other words, the number of appearances of the explanatory variable indicating the actually used raw material ratio is calculated. When the processing in step S536 is completed, the control unit 11 proceeds to step S538.
[0057] In step S538, the control unit 11 causes the regression model creation processing unit 112 to determine the weight reference using the function of the number of occurrences of the explanatory variable calculated in step S536. A concrete example of this function may be the occurrence probability of the explanatory variable in the entire training data. Furthermore, the regression model creation processing unit 112 may determine the number of occurrences of the explanatory variable as the weight reference instead of determining the occurrence probability of the explanatory variable as the weight reference. Incidentally, the reason why the reference based on the explanatory variable is used is to incorporate the user's, who is the material developer, knowledge of materials science, particularly regarding important parameters not indicated by the target variable.When the processing in step S538 is completed, the control unit 11 completes the processing for calculating the weight reference shown in step S330 and proceeds to step S540.
[0058] Further, in step S540, the regression model creation processing unit 112 performs weighting of the measured feature data or the regression model as described above by performing either the processing for setting the loss function, the oversampling processing for amplifying the rare or highly significant measured feature data and adding the amplified data as learning data, or the undersampling processing for deleting the redundant or less significant measured feature data from the learning data.More specifically, the processing for establishing a regression model further includes one of the following two processings executed based on the calculated weight reference, which are the processing for setting the loss function, the oversampling processing for amplifying the rare or highly significant measured feature data and adding the amplified data as the learning data, and the undersampling processing for deleting the redundant or less significant measured feature data from the learning data.
[0059] Of the processings described above, the concrete content of the processing executed in step S530 when the loss function is set based on the weight reference calculated in step S530 will be explained below as step S542.
[0060] In step S542, the control unit 11 causes the regression model creation processing unit 112 to determine machine learning so that a prediction error of the rare or highly significant measured feature data becomes small based on the weight reference calculated in step S530. This processing is performed by applying a weight to the measured feature data through the loss function during learning. When the processing in step S538 is completed, the control unit 11 completes the processing that causes the weight reference specified in step S540 to be incorporated into the learning data or the regression model and proceeds to step S550.
[0061] Further, the concrete content of the processing called oversampling processing or upsampling processing executed in step S530 when the rare or highly significant measured feature data is amplified and the amplified data is added as learning data based on the weight reference calculated in step S530 will be explained below as step S544.
[0062] In step S544, the control unit 11 causes the regression model creation processing unit 112 to duplicate the rare or highly significant measured feature data, which is the goal of oversampling, based on the weight reference calculated in step S530, and add the learning data to the duplicated data. In this case, the measured feature data is duplicated until its amount is equal to or greater than a specific threshold. Furthermore, the oversampling processing can be performed by various methods such as SMOTE, ADASYN, Borderline-SMOTE, and Safe-level SMOTE, in addition to the method of adding the duplicated measured feature data to the learning data. Accordingly, the rare or highly significant measured feature data is added to the learning data, thereby improving the balance of the learning data.The comparison of the estimation accuracy before and after the weighting processing by oversampling performed in step S544 is shown in . Fig. 6 shown. Fig. Figure 6 shows the measured property data regarding flexural strength, which is one of the properties regarding a ceramic composite made from a variety of specific raw materials, and its predicted values in a scattergram. Fig. 6 shows that, as a result of the weighting of the measured property data relating to certain specific raw materials through the oversampling processing performed in step S544, where the estimation accuracy was not sufficiently achieved due to the small number of uses of these raw materials in the entire measurement data, the errors between the estimated values and the measured values relating to this composite material were successfully reduced; in other words, the estimation accuracy for this composite material was successfully increased. When the processing in step S544 is completed, the control unit 11 completes the processing that causes the weighting reference to be considered in the learning data or the regression model as specified in step S540 and proceeds to step S550.
[0063] Furthermore, the concrete content of the processing called undersampling processing or downsampling processing, which is executed in step S540 when the redundant or less significant measured feature data is deleted from the learning data based on the weight reference calculated in step S530, will be explained below as step S546.
[0064] In step S546, the control unit 11 causes the regression model creation processing unit 112 to delete redundant or less significant measured feature data from the learning data. Accordingly, the redundant or less significant measured feature data is deleted from the learning data, thereby improving the balance of the learning data. When the processing in step S546 is completed, the control unit 11 completes the processing that causes the weight reference to be considered in the learning data or the regression model as specified in step S540 and proceeds to step S550.
[0065] Fig. Figure 7 is a flowchart showing the details of test production condition suggestion processing.
[0066] In step S710, the control unit 11 causes the test production condition suggestion processing unit 113 to start the processing for searching for the test production conditions based on the result obtained in step S560 in Fig. 5. The test production condition suggestion processing unit 113 performs this processing through an optimization process. Incidentally, the test production condition suggestion system 1 according to this embodiment is configured to be capable of applying various types of optimization processing methods, such as mathematical optimization (MO), Bayesian optimization (BO), genetic algorithm (GA), Newton's method (NM), and simplex method (SM). As a result, when a predicted value of the characteristics, i.e., the target variable, becomes the best, the respective explanatory variables are suggested to the user as preliminary test production conditions, indicating the results of the relevant search processing.Further, under these circumstances, the test production condition suggestion processing unit 113 performs a sensitivity analysis of the relevant preliminary test production conditions, evaluates the importance of the respective explanatory variables constituting the relevant preliminary test production conditions, and compiles the evaluation results. Furthermore, the test production condition suggestion system 1 according to this embodiment can select the test production conditions that maximize the detection function when the regression model used is Gaussian process regression. When the processing in step S710 is completed, the control unit 11 proceeds to step S720.
[0067] In step S720, the control unit 11 causes the test production condition suggestion processing unit 113 to accept from the user (a) modification(s) of the preliminary test production condition(s) proposed to the user in step S710. After accepting an input operation regarding the modification(s) of the values of the respective explanatory variables constituting the preliminary test production conditions via the input unit 131 or the communication unit 14, the test production condition suggestion processing unit 113 changes the preliminary test production conditions according to the modification content.Furthermore, under these circumstances, the control unit 11: determines a predicted value for the properties of the material when the material is test-produced under the modified preliminary test production conditions using the regression model; and presents the calculation result to the user. Furthermore, under these circumstances, the test production condition suggestion processing unit 113 also performs the sensitivity analysis of the modified preliminary test production condition in the same manner as in step S710, evaluates the significance of the explanatory variables constituting the relevant modified preliminary test production conditions, and presents the evaluation results together. Furthermore, these evaluation results are updated each time the user modifies the preliminary test production conditions; and the user is always presented with the latest evaluation results.When the processing in step S720 is completed, the control unit 11 proceeds to step S730.
[0068] In step S730, the control unit 11 causes the test production condition suggestion processing unit 113 to judge whether or not the predicted value of the properties obtained in step S720 with respect to the modified preliminary test production conditions is insufficient as the property value of the material subject to test production.For example, if the regression model to be used is Gaussian process regression, this assessment is performed for each test production condition by determining a detection function that indicates an expected value for an improvement in the properties of the material test-produced under the relevant test production condition, with respect to the modified preliminary test production condition, and assessing whether the difference between the value of the respective detection function and a maximum value of the detection function lies within a specific range. Furthermore, the detection function is calculated based on the predicted value µ of the material properties and a standard deviation σ that indicates the deviations of the relevant predicted value when the test production of the material is conducted under any test production condition.If it is judged that the predicted property value is insufficient, the processing returns to step S720 and again accepts an instruction to modify the test production conditions from the user; and if it is judged that the predicted property value is not insufficient, it means that the predicted value of the material's properties is sufficient when performing test production under the relevant modified preliminary test production condition, so the preliminary test production condition(s) is / are determined to be final, and the user is suggested that the test production condition(s) is / are finalized. Specifically, this judgment processing is repeated until it is judged that the predicted value of the material's properties is not insufficient with respect to the preliminary test production condition(s).When the processing in step S730 is completed, the control unit 11 ends the processing shown in the flowchart in FIG. Fig. 7 illustrated test production condition proposal processing.
[0069] Incidentally, the test production condition suggestion system 1 according to this embodiment suggests more optimized test production condition(s) to the user based on the newly input measured property data as described above when the user performs the test production of the material under the condition(s) specified in step S730 in Fig. 7 proposed test production condition(s) and the data indicating the actual measurement result of the product's properties are input as new measured property data. In this case, the control unit 11 for the test production condition suggestion system 1 judges whether or not a missing value is included in the newly inputted measured property data. If, on the one hand, it is judged that a missing value is included, the preprocessing of measured property data is applied to the newly inputted measured property data to supplement this missing value by executing the processing from step S440 in Fig. 4. On the other hand, if it is judged that a missing value is not included, it is not necessary to perform the preprocessing of measured property data on the newly inputted measured property data. Therefore, in such a case, the control unit 11 judges whether it is necessary to update the regression model or not. On the one hand, if it is judged that it is necessary to update the regression model, the control unit 11 executes the regression model creation processing on the newly inputted measured property data by executing the processing from step S510 in Fig. 5 to generate a regression model again. On the other hand, if it is judged that it is not necessary to update the regression model, the control unit 11 skips the regression model creation processing and the regression model evaluation processing and executes the test production condition suggestion processing based on the newly input measured property data by continuing the processing from step S710 in Fig. 7 using the most recently generated regression model. Incidentally, if it is not necessary to update the regression model even if a missing value is included in the newly input measured property data, the control unit 11 similarly omits the regression model creation processing and the regression model evaluation processing.
[0070] According to the above-described embodiment of the present invention, the following operational advantages can be achieved. (1) The test production condition suggestion system 1 is a system for suggesting test production conditions for a material(s) to a material developer and includes the regression model creation processing unit 112 and the test production condition suggestion processing unit 113. The regression model creation processing unit 112 executes the regression model creation processing based on the measured property data indicating the actual measurement result of the material's properties (step S320). The test production condition suggestion processing unit 113 searches for an optimal test production condition for the material using the created regression model and executes test production condition suggestion processing using the search result (step S340). The regression model creation processing ( Fig. 5) includes: processing for calculating the weighting reference, which is a reference for weighting measured property data (step S530); and processing for performing weighting of the measured property data based on the calculated weighting reference (step S540). Accordingly, even when the number of individual learning data used to construct a machine learning model is small or when there is a distortion in the distribution of the learning data, the weighting of the above-described learning data is adjusted accordingly. As a result, it is possible to improve the estimation accuracy and propose the test production condition for a good material with excellent accuracy. (2) The weighting reference is calculated based on the difference between the target variable contained in the measured property data and the target property that indicates a target value of the material's properties (steps S532 to S534). Accordingly, the weighting reference can be determined based on the target variable even for material development for which the weighting reference cannot be determined based on the number of occurrences of the target variable, since regression problems are frequently solved. (3) The weighting reference is calculated based on the statistical quantity indicating the rarity of the explanatory variable in the measured property data (steps S536 to S538). Accordingly, the weighting reference can be calculated based on the explanatory variable. As a result, it is possible to incorporate the user's, who is the material developer, knowledge of materials science, especially with regard to the important parameter(s) not specified by the target variable. (4) The regression model creation processing ( Fig. 5) further includes the processing for setting the loss function based on the calculated weight reference (step S542). Accordingly, it is possible to reduce any prediction errors related to rare or highly significant measured property data. (5) The regression model creation processing ( Fig. 5) further includes oversampling processing (step S544) for amplifying rare or highly significant measured feature data based on the calculated weight reference and adding the amplified data as learning data. Accordingly, the rare or highly significant measured feature data is added to the learning data. As a result, the sensitivity of the rare or highly significant measured feature data can be increased, thereby improving the balance of the learning data. (6) The regression model creation processing ( Fig.5) further includes undersampling processing (step S546) for deleting redundant or less significant measured feature data from the learning data based on the calculated weight reference. Accordingly, redundant or less significant measured feature data are deleted from the learning data. As a result, the sensitivity of the redundant or less significant measured feature data can be reduced, thereby improving the balance of the learning data.
[0071] Incidentally, the present invention is not limited to the above-described embodiment and can be implemented using any components within the scope that does not deviate from the gist of the invention.
[0072] The embodiments and variations described above are merely examples, and the present invention is not limited to their contents unless the characteristics of the invention are impaired. Furthermore, various embodiments and variations are explained above, but the present invention is not limited to their contents. Other aspects conceivable within the scope of the technical concepts of the present invention are also included within the scope of the present invention. LIST OF REFERENCE SYMBOLS 1 test production condition suggestion system 11 Control unit 12 storage units 13 User interface unit 14 Communication unit 111 Preprocessing unit for measured property data 112 Regression model building processing unit 113 Test production condition proposal processing unit 131 Input unit 132 Output unit 400 Internet QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] WO 2021 / 044913
[0003]
Claims
[1] Test production condition suggestion system for suggesting a test production condition for a material to a material developer, where the test production condition suggestion system comprises: a regression model creation processing unit that performs regression model creation processing on measured property data indicating an actual measurement result of properties of the material; and a test production condition suggestion processing unit that searches for an optimal test production condition for the material using the created regression model and executes test production condition suggestion processing based on a search result, where the regression model building processing comprises: processing for calculating a weighting reference, which is a reference for weighting the measured property data; and processing for performing weighting of the measured property data based on the calculated weight reference. [2] The test production condition suggestion system according to claim 1, wherein the weight reference is calculated based on a difference between a target variable included in the measured property data and a target property indicating a target value of the properties of the material. [3] The test production condition suggestion system according to claim 1, wherein the weight reference is calculated based on a statistical quantity indicating the rarity of an explanatory variable included in the measured property data. [4] The test production condition suggestion system according to claim 1, wherein the regression model creation processing further comprises the processing of setting a loss function based on the calculated weight reference. [5] The test production condition suggestion system according to claim 1, wherein the regression model creation processing further comprises oversampling processing for amplifying rare or highly significant measured property data based on the calculated weight reference and adding the amplified measured property data as learning data. [6] The test production condition suggestion system according to claim 1, wherein the regression model creation processing further comprises undersampling processing for deleting redundant or less significant measured feature data from the learning data based on the calculated weight reference. [7] A test production condition proposal method for proposing a test production condition for a material to a material developer using a computer, wherein the computer is caused to execute the following: Regression model creation processing for creating a regression model for measured property data indicating an actual measurement result of the properties of the material; and Test production condition suggestion processing for searching for an optimal test production condition for the material using the created regression model, and suggesting a test production condition for the material based on a search result, where the regression model building processing comprises: Processing for calculating a weighting reference, which is a reference for weighting the measured property data; and Processing to perform weighting of the measured property data based on the calculated weighting reference.
Citation Information
Patent Citations
Preparation and evaluation system, preparation and evaluation method, and program
WO2021044913A1