Data processing device and inference method
The data processing device and method enhance inference result usefulness by varying one explanatory variable while fixing others, addressing the challenge of understanding parameter effects in neural network predictions, thus improving user convenience and efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SHIMADZU SEISAKUSHO LTD
- Filing Date
- 2022-06-23
- Publication Date
- 2026-05-26
AI Technical Summary
Existing methods for estimating the performance of a structural complex using a neural network provide only one estimated value, making it difficult for users to understand how individual parameters affect the prediction, and require multiple dataset iterations for parameter analysis, reducing user convenience and efficiency.
A data processing device and method that predicts target variables from multiple explanatory variables using a trained model, allowing for the variation of a selected explanatory variable while keeping others fixed, and generates data to display the variation of target variables, enhancing the understanding of parameter effects on predictions.
Improves the usefulness of inference results by enabling users to see how changes in individual parameters affect predictions, thereby increasing user convenience and efficiency in the inference process.
Smart Images

Figure 0007865118000001 
Figure 0007865118000002 
Figure 0007865118000003
Abstract
Description
Technical Field
[0001] The present invention relates to a data processing device and an inference method.
Background Art
[0002] Japanese Unexamined Patent Application Publication No. 2018-036131 (Patent Document 1) discloses a method for estimating the state of a target structural complex from a plurality of parameters obtained by measuring the target structural complex using a trained neural network.
[0003] In Patent Document 1, when a data set of a plurality of parameters is input to the input layer of a neural network, an estimated value of the performance of the structural complex is output from the output layer of the neural network.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, in the estimation method described in Patent Document 1, only one estimated value is obtained corresponding to the data set of a plurality of parameters input to the neural network. Therefore, it is considered that it is not easy for the user to know how each of the plurality of parameters affects the predicted performance from this one estimated value. For example, when the values of some of the plurality of parameters fluctuate (increase or decrease), it is considered that it is not easy to predict how the estimated value will fluctuate.
[0006] Furthermore, in order to understand how individual parameters affect the estimated values, it would be necessary to prepare multiple datasets of the multiple parameters input to the neural network and repeat the process of obtaining estimated values for each dataset. This raises concerns about reduced user convenience and decreased efficiency in the inference process. These concerns may become more pronounced as the number of parameters input to the neural network increases.
[0007] This invention was made to solve these problems, and its purpose is to improve the usefulness of the inference results output from a trained model that receives input from multiple explanatory variables. [Means for solving the problem]
[0008] A data processing device according to a first aspect of the present invention comprises an inference unit that predicts at least one target variable from a plurality of explanatory variables using a trained model, and a display data generation unit that generates data for displaying the inference results of the inference unit. The inference unit sets a first explanatory variable selected from the plurality of explanatory variables as a variable value, while setting a second explanatory variable other than the first explanatory variable as a fixed value. The inference unit uses a trained model to predict at least one target variable when the first explanatory variable is continuously varied within a predetermined range of variation. The display data generation unit generates data showing the variation of at least one target variable in response to the variation of the first explanatory variable.
[0009] A second aspect of the present invention relates to an inference method that predicts at least one target variable from a plurality of explanatory variables using a trained model. The inference method includes the steps of: predicting at least one target variable when a first explanatory variable selected from the plurality of explanatory variables is set as a variable value, while a second explanatory variable other than the first explanatory variable is set as a fixed value, and using a trained model, the first explanatory variable is continuously varied within a predetermined range of variation; generating data showing the variation of at least one target variable in response to the variation of the first explanatory variable; and displaying the data generated in the generation step. [Effects of the Invention]
[0010] According to the present invention, the usefulness of the inference results output from a trained model that receives input from multiple explanatory variables can be improved. [Brief explanation of the drawing]
[0011] [Figure 1] This is a schematic diagram illustrating an example of the configuration of the analysis system according to Embodiment 1. [Figure 2] This diagram schematically shows examples of hardware configurations for information processing devices and data processing devices. [Figure 3] This diagram schematically shows the functional configuration of the information processing device and the data processing device. [Figure 4] This is a flowchart illustrating the overview of the processes performed by the data processing device. [Figure 5] This is a flowchart illustrating the process for generating a sample list. [Figure 6] This figure shows an example of the structure of a sample list. [Figure 7] This flowchart illustrates the process of generating training data, machine learning, and storing the trained model. [Figure 8] This figure shows an example of the structure of a training data table. [Figure 9] This figure shows an example of how to construct a list of trained models. [Figure 10] A flowchart for explaining the processing procedure of inference processing in the data processing apparatus according to Embodiment 1. [Figure 11] A diagram schematically showing a first display example of an inference result on a display unit. [Figure 12] A diagram schematically showing a second display example of an inference result on a display unit. [Figure 13] A diagram schematically showing a third display example of an inference result on a display unit. [Figure 14] A flowchart for explaining the processing procedure of inference processing in the data processing apparatus according to Embodiment 2. [Figure 15] A flowchart for explaining the processing procedures of generating teacher data, machine learning, and storing a learned model in the data processing apparatus according to Embodiment 2. [Figure 16] A flowchart for explaining the inference processing in the data processing apparatus according to the first configuration example of Embodiment 3. [Figure 17] A diagram schematically showing a display example of an inference result in the first configuration example. [Figure 18] A flowchart for explaining the inference processing in the data processing apparatus according to the second configuration example of Embodiment 3. [Figure 19] A flowchart for explaining the inference processing in the data processing apparatus according to the third configuration example of Embodiment 3. [Figure 20] A flowchart for explaining the processing procedures of generating teacher data, machine learning, and storing a learned model in the data processing apparatus according to Embodiment 4. [Figure 21] A diagram showing a configuration example of a sample list. [Figure 22] A diagram showing a configuration example of a selected sample extraction table. [Figure 23] A diagram showing a configuration example of a learned model list. [Figure 24] A flowchart for explaining the processing procedure of inference processing in the data processing apparatus according to Embodiment 5. [Figure 25] This diagram schematically shows examples of how multiple inference results are displayed on the display unit. [Figure 26] This is a flowchart illustrating the processing procedure for generating a sample list in the data processing device according to Embodiment 5. [Modes for carrying out the invention]
[0012] Embodiments of the present invention will be described in detail below with reference to the drawings. In the following description, the same or corresponding parts in the drawings will be denoted by the same reference numerals, and their descriptions will not be repeated in principle.
[0013] [Embodiment 1] [Example of an analysis system configuration] Figure 1 is a schematic diagram illustrating an example configuration of the analysis system according to Embodiment 1. The analysis system according to Embodiment 1 can be applied to a system for cross-sectionally analyzing analysis data acquired by multiple analytical devices.
[0014] As shown in Figure 1, the analysis system 100 according to Embodiment 1 comprises a plurality of analysis devices 4 and a data processing device 1.
[0015] Multiple analytical instruments 4 perform measurements of the sample. These analytical instruments 4 include, for example, a liquid chromatograph (LC), a gas chromatograph (GC), a liquid chromatograph-mass spectrometer (LC-MS), a gas chromatograph-mass spectrometer (GC-MS), a pyrolysis gas chromatograph-mass spectrometer (Py-GC / MS), a scanning electron microscope (SEM), a transmission electron microscope (TEM), an energy-dispersive X-ray fluorescence analyzer (EDX), a wavelength-dispersive X-ray fluorescence analyzer (WDX), a nuclear magnetic resonance spectrometer (NMR), and a Fourier transform infrared spectrophotometer (FT-IR). The multiple analytical instruments 4 may further include a photodiode array detector (LC-PDA), liquid chromatograph tandem mass spectrometer (LC / MS / MS), gas chromatograph tandem mass spectrometer (GC / MS / MS), liquid chromatograph ion trap time-of-flight mass spectrometer (LC / MS-IT-TOF), near-infrared spectrometer, tensile tester, compression tester, emission spectrometer (AES), atomic absorption spectrometer (AAS / FL-AAS), plasma mass spectrometer (ICP-MS), organic elemental analyzer, glow discharge mass spectrometer (GDMS), particle composition analyzer, trace total nitrogen automated analyzer (TN), high-sensitivity nitrogen-carbon analyzer (NC), and thermal analyzer. By having multiple analytical instruments 4 of different types, the analysis system 100 makes it possible to analyze a single sample from multiple perspectives using multiple types of analytical data.
[0016] The analyzer 4 includes a main unit 5 and an information processing unit 6. The main unit 5 measures the sample to be analyzed. The information processing unit 6 receives the sample identification information and the measurement conditions for the sample.
[0017] The information processing device 6 controls the measurement in the main unit 5 according to the input measurement conditions. This allows for the acquisition of analytical data based on the measurement results of the sample. The information processing device 6 stores the acquired analytical data, along with the sample identification information and measurement conditions, in a data file and saves this data file in its internal memory.
[0018] The information processing device 6 is connected to the data processing device 1 in a way that allows for mutual communication. The connection between the information processing device 6 and the data processing device 1 may be wired or wireless. For example, the internet can be used as the communication network connecting the information processing device 6 and the data processing device 1. This allows the information processing device 6 of each analysis device 4 to transmit data files for each sample to the data processing device 1.
[0019] The data processing device 1 is primarily a device for managing analytical data acquired by multiple analytical instruments 4. Analytical data from each analytical instrument 4 is input to the data processing device 1. Furthermore, information about the sample (hereinafter also referred to as "sample information") and physical property data of the sample can be input to the data processing device 1.
[0020] Sample information includes identification information for identifying the sample (such as sample ID and sample name) and information regarding the preparation of the sample (hereinafter also referred to as "recipe data"). The recipe data for a sample may include information on the proportions of the raw materials used in the sample and the manufacturing process. For example, if the sample is a three-way catalyst, the recipe data may include the proportions of Pt (platinum) (g), Pd (palladium) (g), stirring time (min), and firing temperature (°C).
[0021] Sample property data refers to data indicating the attributes of the sample, obtained through means other than analysis by the analytical instrument 4. For example, if the sample is a three-way catalyst, the property data would include the NOx (nitrogen oxide) purification rate (%), the CO (carbon monoxide) purification rate (%), the HC (hydrocarbon) purification rate (%), and heat resistance performance.
[0022] The data processing device 1 has a built-in database. The database is a storage unit for saving data exchanged between the data processing device 1 and the multiple analytical devices 4, data input from outside the data processing device 1, and data generated within the data processing device 1. For each sample, the data processing device 1 links the data file with the sample information and the physical property data of the sample and stores them in the database. In the example in Figure 1, the database is built into the data processing device 1, but the database may also be externally connected to the data processing device 1.
[0023] [Example Hardware Configuration for Analysis System] Figure 2 is a schematic diagram showing an example of the hardware configuration of the information processing device 6 and the data processing device 1.
[0024] (Hardware configuration of information processing equipment) As shown in Figure 2, the information processing device 6 includes a CPU (Central Processing Unit) 60 for controlling the entire analysis device 4, and a storage unit for storing programs and data, and is configured to operate according to the program.
[0025] The memory unit includes a ROM (Read Only Memory) 61, a RAM (Random Access Memory) 62, and an HDD (Hard Disk Drive) 65. The ROM 61 stores programs executed by the CPU 60. The RAM 62 temporarily stores data used during program execution by the CPU 60. The RAM 62 functions as a temporary data memory used as a work area. The HDD 65 is a non-volatile storage device that stores information generated by the information processing device 6, such as data files for each sample. In addition to the HDD 65, or instead of the HDD 65, semiconductor storage devices such as flash memory may be used.
[0026] The information processing device 6 further includes a communication interface (I / F) 66, an operation unit 63, and a display unit 64. The communication I / F 66 is an interface for the information processing device 6 to communicate with external devices, including the main unit 5 and the data processing device 1.
[0027] The operation unit 63 receives input from a user (e.g., an analyst) including instructions for the information processing device 6. The operation unit 63 includes a keyboard, mouse, and a touch panel integrated with the display screen of the display unit 64, and receives sample measurement conditions and identification information.
[0028] The display unit 64 can display, for example, the input screen for measurement conditions and sample identification information when setting measurement conditions. During measurement, the display unit 64 can display the measurement data detected by the main unit 5 and the data analysis results from the information processing device 6.
[0029] The processing in the analysis device 4 is realized by each piece of hardware and software executed by the CPU 60. Such software may be pre-stored in ROM 61 or HDD 65. Alternatively, the software may be stored on a storage medium (not shown) and distributed as a program product. The software is then read from HDD 65 by the CPU 60 and stored in RAM 62 in an executable format by the CPU 60. The CPU 60 then executes this program.
[0030] (Hardware configuration of data processing equipment) The data processing device 1 includes a CPU 10 for controlling the entire device and a storage unit for storing programs and data, and is configured to operate according to the program. The storage unit includes a ROM 11, a RAM 12, and a database 15.
[0031] ROM11 stores programs executed by CPU10. RAM12 temporarily stores data used during program execution on CPU10. RAM12 functions as temporary data memory used as a working area.
[0032] The database 15 is a non-volatile storage device that stores data exchanged between the data processing device 1 and the multiple analyzers 4, data input from outside the data processing device 1, and data generated within the data processing device 1.
[0033] The data processing device 1 further includes a communication interface 13 and an input / output interface (I / O) 14. The communication interface 13 is an interface for the data processing device 1 to communicate with external devices, including an information processing device 6.
[0034] I / O14 is an interface for input to or output from the data processing device 1. I / O14 is connected to the display unit 2 and the operation unit 3. As will be described later, the display unit 2 can display information related to the processing when the learning process and inference process are executed in the data processing device 1, and can also display a user interface screen for accepting user input.
[0035] The control unit 3 receives input, including user instructions. The control unit 3 includes a keyboard and mouse, and receives sample information and sample property data. Sample information and sample property data can also be received from external devices via the communication interface 13.
[0036] [Functional Configuration of the Analysis System] Figure 3 is a schematic diagram showing the functional configuration of the information processing device 6 and the data processing device 1.
[0037] (Functional configuration of information processing equipment) As shown in Figure 3, the information processing device 6 has a data acquisition unit 67 and an information acquisition unit 69. These functional configurations are realized in the information processing device 6 shown in Figure 2 by the CPU 60 executing a predetermined program.
[0038] The data acquisition unit 67 acquires analytical data from the main unit 5 based on the measurement results of the sample. For example, if the analyzer 4 is a gas chromatograph-mass spectrometer (GC-MS), the analytical data includes a chromatogram and a mass spectrum. If the analyzer 4 is a scanning electron microscope (SEM) or a transmission electron microscope (TEM), the analytical data includes image data showing a microscopic image of the sample. The data acquisition unit 67 transfers the acquired measurement data to the communication interface 66.
[0039] The information acquisition unit 69 acquires the information received by the operation unit 63. Specifically, the information acquisition unit 69 acquires sample identification information and information indicating the measurement conditions of the sample. Sample identification information includes, for example, the sample name, the name, model number, and serial number of the product to be sampled. The measurement conditions of the sample include instrument parameters, such as the name and model number of the analytical instrument used, and measurement parameters, such as the application conditions of voltage and / or current or temperature conditions.
[0040] Communication I / F66 transmits the acquired analysis data, measurement conditions, and sample identification information as a data file to the data processing device 1.
[0041] (Functional configuration of data processing equipment) The data processing device 1 comprises an analysis data acquisition unit 20, a feature extraction unit 22, a physical property data acquisition unit 24, a sample information acquisition unit 26, a training data generation unit 28, a learning unit 30, an inference unit 32, and a display data generation unit 34. These functional configurations are realized in the data processing device 1 shown in Figure 2 by the CPU 10 executing a predetermined program.
[0042] The analysis data acquisition unit 20 acquires data files transmitted from the information processing device 6 of each analysis device 4 via the communication interface 13. The data files contain the analysis data of the sample.
[0043] The feature extraction unit 22 extracts sample features by analyzing the analysis data acquired by the analysis data acquisition unit 20 using dedicated data analysis software. Sample features include, for example, the composition, concentration, molecular structure, number of molecules, molecular formula, molecular weight, degree of polymerization, particle size, particle area, number of particles, particle dispersion, peak intensity, peak area, peak slope, compound concentration, compound amount, absorbance, reflectance, transmittance, sample test strength, Young's modulus, tensile strength, deformation, strain, fracture time, average interparticle distance, dielectric loss tangent, elongation, spring stiffness, loss coefficient, glass transition temperature, and thermal expansion coefficient.
[0044] The physical property data acquisition unit 24 acquires the physical property data of the sample received by the operation unit 3. The physical property data of the sample is data that indicates the attributes of the sample, and includes, for example, values that indicate the performance of the sample or values that indicate the degree of deterioration of the sample (such as years of use).
[0045] The sample information acquisition unit 26 acquires sample information received by the operation unit 3. The sample information includes sample identification information (sample ID, sample name, etc.) and sample recipe data. The sample recipe data includes information such as the proportions of the sample's raw materials and the manufacturing process.
[0046] Database 15 stores, for each sample, the analysis data acquired by the analysis data acquisition unit 20, the features extracted by the feature extraction unit 22, the physical property data acquired by the physical property data acquisition unit 24, and the sample information acquired by the sample information acquisition unit 26, all linked together. Specifically, a sample list is created in database 15 based on this information. A sample list is a set of datasets created according to the type of project or sample, and its structure is not particularly limited.
[0047] The training data generation unit 28 generates training data (training data) based on the data stored in the database 15 in response to user input operations to the operation unit 3. Training data is data consisting of a set of input (explanatory variables) and output (dependent variable).
[0048] The training data generation unit 28 can generate training data in which, for example, the "analysis data or features" and / or "recipe data" of one sample are used as inputs (explanatory variables) to the prediction model, and the "physical property data" of that sample is used as the output (dependent variable) to the prediction model.
[0049] Alternatively, the training data generation unit 28 can generate training data in which the "recipe data" of one sample is used as the input (explanatory variable) of the prediction model, and the "analysis data or features" or "physical property data" of that sample is used as the output (dependent variable) of the prediction model.
[0050] Alternatively, the training data generation unit 28 can generate training data in which the "physical property data" of one sample is used as the input (explanatory variable) of the prediction model, and the "analysis data or features" or "recipe data" of that sample is used as the output (dependent variable) of the prediction model.
[0051] The generated training data is provided to the learning unit 30. The training data may be stored in the database 15 each time it is generated. As a result, training data is accumulated in the database 15.
[0052] Before storing the training data in the database 15, the training data generation unit 28 uses the display data generation unit 34 to display a confirmation screen on the display unit 2 to confirm whether or not to store the training data in the database 15. If the confirmation screen accepts a user operation to instruct the storage of the training data, the training data generation unit 28 stores the training data in the database 15. On the other hand, if the confirmation screen does not accept the above instruction, the training data generation unit 28 discards the training data.
[0053] The learning unit 30 uses the training data generated by the training data generation unit 28 to perform supervised learning, using the explanatory variables of the training data as inputs to a prediction model and the target variable of the training data as the ground truth output data of the prediction model. In supervised learning, it predicts what kind of output a given input will produce. The machine learning method using the training data in the learning unit 30 is not particularly limited, and known machine learning methods such as neural networks (NN) or support vector machines (SVM) can be used.
[0054] Once training is complete, a trained model is obtained. The generated trained model is stored in database 15. Specifically, the trained model is stored in database 15, linked to identification information for identifying the trained model, the date and time the trained model was created, and identification information for identifying the training data used for training.
[0055] The inference unit 32 uses the trained model stored in the database 15 to predict the output (target variable) from the input data (explanatory variables) newly input from one or two of the analysis data acquisition unit 20, feature extraction unit 22, physical property data acquisition unit 24, and sample information acquisition unit 26. That is, the explanatory variables are one or two of "analysis data or features", "physical property data", and "recipe data", and the target variable is the other one of "analysis data or features", "physical property data", and "recipe data". Alternatively, the explanatory variables are one of "analysis data or features", "physical property data", and "recipe data", and the target variable is the other one or two of "analysis data or features", "physical property data", and "recipe data". As an example, the explanatory variables are "analysis data or features" and / or "recipe data", and the target variable is "physical property data".
[0056] When the display data generation unit 34 obtains the inference result from the inference unit 32, it generates data for displaying the inference result on the display screen of the display unit 2. Furthermore, when the learning process and inference process are executed, the display data generation unit 34 displays information related to the process and provides a user interface (UI) for accepting user input.
[0057] Alternatively, instead of the operation unit 3 and the display unit 2, an information terminal such as a desktop personal computer (PC), notebook PC, or mobile terminal (tablet terminal, smartphone) may be connected to the data processing device 1.
[0058] Furthermore, although the example in Figure 3 shows a configuration in which the data processing device 1 has a feature extraction unit 22, the information processing device 6 may also have a feature extraction unit. In this case, the information processing device 6 will transmit the features along with the sample analysis data to the data processing device 1.
[0059] [Operation of the data processing device] Next, we will explain the processing performed by the data processing device 1.
[0060] Figure 4 is a flowchart illustrating the overview of the processing performed by the data processing device 1. As shown in Figure 4, the processing in the data processing device 1 can be mainly divided into a learning phase and an inference phase.
[0061] <Learning Phase> In the learning phase, training data is generated using the data stored in database 15. Then, supervised learning is performed using the generated training data to produce a trained model.
[0062] As shown in Figure 4, first, step 01 (hereinafter simply referred to as "S") generates a sample list based on the data stored in database 15. The sample list is a set of datasets created depending on the type of project or sample.
[0063] Next, in step S02, training data is generated from the sample list in response to the user's input operation to the control unit 3. The generated training data is stored in the database 15.
[0064] Next, in S03, supervised learning is performed using the training data, with the explanatory variables of the training data as inputs to the prediction model and the target variable of the training data as the ground truth output of the prediction model. Finally, in S04, the trained model generated by supervised learning is stored in the database 15.
[0065] The specific processing steps in the learning phase will be explained below using Figures 5 to 9. (1) Generating a sample list (S01 in Figure 4) Figure 5 is a flowchart illustrating the processing procedure for generating a sample list (S01 in Figure 4). Referring to Figure 5, in S10, sample information is acquired via the operation unit 3. Specifically, the user can input sample information using database operation software (front end) not shown. The sample information includes sample identification information (sample ID, sample name, etc.) and sample recipe data. The sample recipe data includes information such as the proportions of the sample's raw materials and the manufacturing process.
[0066] In S11, physical property data of the sample is acquired via the operation unit 3. The physical property data is data that indicates the attributes of the sample.
[0067] In S12, data files transmitted from the information processing device 6 of each analyzer 4 are acquired via the communication I / F 13. The data files contain the analysis data of the samples.
[0068] In S13, the analysis data acquired in S12 is analyzed using dedicated data analysis software to extract the characteristic features of the sample.
[0069] In S14, the acquired sample information (sample identification information, recipe data), sample physical property data, and sample analysis data and features are entered into the sample list. Figure 6 shows an example of the sample list structure. Figure 6 shows an example of the sample list structure when the sample is a three-way catalyst.
[0070] As shown in Figure 6, the sample list contains the sample name, sample recipe data, physical property data, analytical data, and characteristic values, all linked to the sample ID. The recipe data includes the amount of Pt (g), the amount of Pd (g), stirring time (min), and firing temperature (°C). The physical property data includes the NOx purification rate (%), CO purification rate (%), HC purification rate (%), and heat resistance performance. The analytical data includes analytical data obtained by gas chromatography-mass spectrometry (GC-MS), nuclear magnetic resonance (NMR), scanning electron microscope (SEM), and transmission electron microscope (TEM), etc.
[0071] Feature quantities include peak area for a given mass number obtained by analyzing a chromatogram acquired by a gas chromatograph-mass spectrometer (GC-MS), the abundance ratio of a given substance obtained by analyzing an NMR spectrum acquired by a nuclear magnetic resonance spectrometer (NMR), the particle size and average particle size of particles present in the three-way catalyst obtained by analyzing an SEM image acquired by a scanning electron microscope (SEM), and the particle size of particles present in the three-way catalyst obtained by analyzing an TEM image acquired by a transmission electron microscope (TEM).
[0072] The sample list, which contains sample recipe data, physical property data, analytical data, and feature data, is registered in database 15 with information to identify the sample list (such as the name of the sample list).
[0073] (2) Generation of training data and generation of trained models (S02-S04 in Figure 4) Figure 7 is a flowchart illustrating the processing steps for generating training data (S02 in Figure 4), machine learning (S03 in Figure 4), and storing the trained model (S04 in Figure 4).
[0074] Referring to Figure 7, S20 first selects the samples to be used to generate training data. The user can select samples by operating the UI screen (sample selection screen) displayed on the display unit 2 using the operation unit 3.
[0075] This sample selection screen can be created based on a sample list stored in database 15. For example, the sample selection screen displays a list of all samples stored in database 15, including the sample name, recipe data, etc. In this case, the sample selection screen displays a selection icon corresponding to each sample. The user can select any sample by checking the selection icon using the operation unit 3.
[0076] Alternatively, the sample selection screen may display the names of multiple sample lists stored in database 15. In this case, the sample selection screen displays selection icons corresponding to each sample list. When the user checks a selection icon using the operation unit 3, all samples included in the corresponding sample list are selected.
[0077] Once the sample selection is complete, the data contained in the rows of the selected samples is extracted from the sample list, and a selected sample extraction table is generated. The selected sample extraction table may be configured to be displayed on display unit 2 for confirmation of the selection results.
[0078] Next, S21 and S22 select the explanatory variables and target variable to be used to generate the training data. The display unit 2 shows the UI screens (explanatory variable selection screen and target variable selection screen).
[0079] The explanatory variable selection screen is a UI screen for the user to select the type of data to be used for inputting training data. The explanatory variable selection screen lists the types of recipe data, analysis data, and features of the samples included in the selected sample extraction table. For example, if the selected sample is a three-way catalyst, the explanatory variable selection screen lists the types of recipe data such as "amount of platinum (Pt)" and "amount of palladium (Pd)", the types of analysis data such as "GC-MS" and "NMR", and the types of features such as "peak area" and "particle size". The explanatory variable selection screen displays selection icons corresponding to each type. The user can select any explanatory variable by checking the selection icon using the operation unit 3.
[0080] The target variable selection screen is a user interface (UI) screen for users to select the type of data to be used for outputting training data. The target variable selection screen displays a list of recipe data, analysis data, and feature types for the samples included in the selected sample extraction table. The target variable selection screen displays selection icons corresponding to each type. Users can select any target variable by checking the selection icon using the operation unit 3.
[0081] However, on the dependent variable selection screen, selection icons are not displayed for data types that belong to the same category as the data type selected for the explanatory variable, in order to avoid duplicate selections. Therefore, for example, if a data type belonging to "recipe data" is selected as an explanatory variable, the dependent variable can be selected from either "physical property data" or "analysis data or features". Alternatively, if a data type belonging to "physical property data" is selected as an explanatory variable, the dependent variable can be selected from either "recipe data" or "analysis data or features". Alternatively, if a data type belonging to "analysis data or features" is selected as an explanatory variable, the dependent variable can be selected from either "recipe data" or "physical property data".
[0082] Furthermore, a training data generation application may be prepared for each set of data used as explanatory variables and target variables in supervised learning. In this case, the user can select the type of data to be used for input and output of the training data simply by selecting the training data generation application.
[0083] Once the selection of explanatory variables (types of input data) and target variables (types of output data) is complete, in S23, a training data table is generated by extracting data that matches the explanatory variables and data that matches the target variable from the selected sample extraction table. Figure 8 shows an example of the structure of the training data table. Figure 8 shows an example of the structure of the training data table when the selected sample is a three-way catalyst.
[0084] As shown in Figure 8, the ID and sample name of the sample selected in S20 of Figure 7 are displayed vertically. Furthermore, the explanatory variable (type of input data) selected in S21 of Figure 7 and the target variable (type of output data) selected in S22 are displayed horizontally.
[0085] In the example in Figure 8, recipe data is selected as the explanatory variables. Specifically, the amount of Pt (%), the amount of Pd (%), the stirring time (min), and the firing temperature (°C) are selected as the explanatory variables. Physical property data is selected as the dependent variable. Specifically, the NOx purification rate (%), the CO purification rate (%), and the heat resistance performance are selected as the dependent variables.
[0086] The training data table contains data matching the explanatory variable and data matching the dependent variable for each sample. The training data table is displayed on display unit 2. Users can add or modify samples and data types to the displayed training data table. For example, if a sample is added, the sample is added to the selected sample extraction table, and the data for the added sample is added to the training data table. If a data type for an explanatory variable or dependent variable is added, the data for the added variable is added to each sample in the training data table.
[0087] Once the generation of the training data table is complete, training data is generated based on the generated training data table. In the example in Figure 8, training data is generated with the selected explanatory variables, namely the amount of Pt (g), the amount of Pd (g), the stirring time (min), and the firing temperature (°C), as input, and the selected objective variables, namely the NOx purification rate (%), the CO purification rate (%), and the heat resistance performance, as output.
[0088] Next, in S24, supervised learning is performed using the training data, with the explanatory variables of the training data as input to the training model and the target variable of the training data as the ground truth output of the model. Below, we will explain the case using a Support Vector Machine (SVM) as an example of machine learning.
[0089] The learning unit 30 (Figure 3) inputs the amount of Pt (g), the amount of Pd (g), the stirring time (min), and the firing temperature (°C) into the SVM and obtains the NOx purification rate (%), CO purification rate (%), and heat resistance performance output from the SVM. The learning unit 30 compares the obtained NOx purification rate (%), CO purification rate (%), and heat resistance performance with the NOx purification rate (%), CO purification rate (%), and heat resistance performance included in the training data. The learning unit 30 generates a trained model by updating various parameters within the SVM so that the NOx purification rate (%), CO purification rate (%), and heat resistance performance output from the SVM approach the NOx purification rate (%), CO purification rate (%), and heat resistance performance of the training data, respectively.
[0090] When machine learning is completed in S24, a trained model is obtained (S25 in Figure 7). In S26, the generated trained model is stored in database 15. Specifically, the trained model is registered in the trained model list stored in database 15. Figure 9 shows an example of the structure of the trained model list. In the example in Figure 9, the trained model is associated with identification information to identify the trained model (such as a trained model ID), the date and time the trained model was created, information about the trained model, and identification information to identify the training data used for training.
[0091] Information about the trained model may include the name of the project to which the trained model is applied. For example, it may include information such as "Model for improving the purification performance of a three-way catalyst" or "Model for improving the heat resistance of a three-way catalyst." Identification information for the training data may include information about the samples used to generate the training data (sample ID, sample name, etc.), the types of data selected for the explanatory variables, and the types of data selected for the dependent variable. This information can be obtained from the selected sample data extraction table and the training data table.
[0092] <Inference Phase> In the inference phase, the generated trained model is used to predict the target variable from the given explanatory variables. The inference results are displayed on display unit 2. Returning to Figure 4, in S05, new explanatory variables are first acquired. These explanatory variables are one or two of the following: "analysis data or features," "physical property data," and "recipe data."
[0093] In S06, the trained model receives input from the explanatory variables obtained in S05 and predicts the target variable. The target variable to be predicted is one or two other variables from among "analysis data or features," "physical property data," and "recipe data."
[0094] In S07, the inference results from the trained model are displayed on the display unit 2. This allows the user to see the value of the target variable predicted from the explanatory variables.
[0095] However, with the above configuration, only one inference result is obtained for the newly acquired explanatory variable. Therefore, it is not easy for the user to understand how the explanatory variable affects the predicted dependent variable from that single inference result. For example, it is not easy to predict how the value of the dependent variable will change when the value of one explanatory variable is increased (or decreased) based on the single inference result obtained.
[0096] Therefore, Embodiment 1 describes an inference process that enables the usefulness of the inference results to be enhanced. The specific processing details of the inference process will be explained below using Figures 10 to 12.
[0097] Figure 10 is a flowchart illustrating the processing procedure of the inference process (S05-S07 in Figure 4) in the data processing device according to Embodiment 1.
[0098] Referring to Figure 10, in S30, a pre-trained model to be used for inference processing is selected. The display unit 2 shows a UI screen (pre-trained model selection screen). The pre-trained model selection screen is generated based on the pre-trained model list (Figure 9) stored in the database 15. The pre-trained model selection screen displays identification information (such as the pre-trained model ID) of the pre-trained model included in the pre-trained model list (Figure 9), as well as the date and time the pre-trained model was created, information about the pre-trained model (such as the project name), and identification information to identify the training data used for training.
[0099] The pre-trained model selection screen displays selection icons corresponding to each pre-trained model. Users can select any pre-trained model by checking the selection icon using the operation unit 3.
[0100] The selection of a pre-trained model to be used for inference determines the data types of the explanatory variables input to the pre-trained model, and the data types of the target variable predicted by the pre-trained model. The data for the explanatory variables is one of "recipe data," "physical property data," and "analysis data or features," and the data for the target variable is one or two of the other "recipe data," "physical property data," and "analysis data or features." Alternatively, the data for the explanatory variables is two of the "recipe data," "physical property data," and "analysis data or features," and the data for the target variable is one of the other "recipe data," "physical property data," and "analysis data or features."
[0101] Specifically, the list of trained models (Figure 9) registers each trained model with its associated identification information for the training data used to generate that model. As mentioned above, the identification information for the training data includes information about the samples used to generate the training data (sample ID, sample name, etc.), the types of data selected for the explanatory variables, and the types of data selected for the target variable. Therefore, by selecting a trained model, the types of data for the explanatory variables input to the trained model and the types of data for the target variable predicted by the trained model can be automatically determined.
[0102] Next, in S31, the values of the explanatory variables to be input into the trained model are set. The display unit 2 shows a UI screen (sample selection screen) for selecting the sample to be analyzed. The user can select the sample to be analyzed by operating the sample selection screen using the operation unit 3.
[0103] This sample selection screen can be created based on the sample list (Figure 6) stored in database 15. For example, the sample selection screen displays a list of all samples stored in database 15, including the sample name and recipe data. The sample selection screen also displays selection icons corresponding to each sample. The user can select any sample by checking the selection icon using the operation unit 3.
[0104] Once a sample to be analyzed is selected, the values of the explanatory variables to be input into the trained model are set based on the recipe data, physical property data, analysis data, and features of the selected sample. The user can then adjust each explanatory variable to their desired value by increasing or decreasing these set values from the reference value using the operation unit 3.
[0105] Next, in S32, "explanatory variables that will be variable values" are selected from among multiple explanatory variables. In the inference process according to this embodiment, some of the multiple explanatory variables input to the trained model are set as variable values, and the remainder of these multiple explanatory variables are set as fixed values. The number of such partial explanatory variables may be one or two or more. As will be described later, the user can select the explanatory variables that will be variable values by operating the user interface screen displayed on the display unit 2 using the operation unit 3.
[0106] A "variable explanatory variable" is an explanatory variable whose value fluctuates within a predetermined range during the inference process. In contrast, a "fixed explanatory variable" is an explanatory variable whose value remains constant during the inference process.
[0107] In S33, the range of variation for the explanatory variables, which are the variable values, is set. As will be described later, the user can set the range of variation for the explanatory variables by using the operation unit 3 to input the upper and lower limits of the range of variation on the user interface screen displayed on the display unit 2. The range of variation for the explanatory variables can also be set automatically by the data processing device 1 based on the sample list (Figure 6) stored in the database 15. For example, data corresponding to the explanatory variables that are the variable values can be extracted from the training data used to generate the trained model, the minimum value of the extracted data can be set as the lower limit of the recommended range, and the maximum value can be set as the upper limit of the recommended range.
[0108] In S34, the target variables to be displayed are selected from among the target variables predicted by the trained model. The user can select the target variables to be displayed by operating the user interface screen (target variable selection screen) displayed on the display unit 2 using the operation unit 3. The number of target variables to be displayed may be one or two or more.
[0109] This target variable selection screen displays a list of data types for the target variable predicted by the trained model selected in S30. The target variable selection screen also displays selection icons corresponding to each target variable. The user can select any target variable to display by checking the selection icon using the operation unit 3.
[0110] Next, in S35, the target variable is predicted by inputting the explanatory variables set in S31-S33 to the trained model selected in S30. In this inference process, the target variable is predicted in accordance with the values of some of the continuously fluctuating explanatory variables among the multiple explanatory variables. In other words, how the target variable will fluctuate based on the fluctuations of these some explanatory variables is predicted.
[0111] In S36, the inference results obtained from the inference processing in S35 are displayed on the display unit 2. The display unit 2 displays a graph showing the variation of the dependent variable selected for display in S34 in relation to the variation of some of the explanatory variables.
[0112] Next, we will explain an example of how the inference results are displayed in the display unit 2 using Figures 11 and 12. Figure 11 schematically shows a first example of the display of inference results in the display unit 2. Figure 11 illustrates the inference results displayed in the display unit 2 when "recipe data" is selected as the explanatory variable input to the trained model and "physical property data" is selected as the target variable to be predicted. The sample is a three-way catalyst.
[0113] The example display in Figure 11 is generated by selecting a pre-trained model to be used for inference processing, selecting the samples to be analyzed, and selecting the target variable to be displayed.
[0114] As shown in Figure 11, the display unit 2 displays a GUI (Graphical User Interface) 70. GUI 70 is a GUI for selecting explanatory variables that will be variable values from among multiple explanatory variables input to the trained model. Specifically, GUI 70 includes GUI 80 for selecting explanatory variables that will be variable values, GUI 84 for setting the range of variation of the explanatory variables, and GUI 86 for setting the step size of the explanatory variables within that range of variation. The user can select explanatory variables that will be variable values, as well as set the range of variation and step size of those explanatory variables, by operating GUI 80, 84, and 86 using the operation unit 3.
[0115] In the right corner of GUI80, an icon 82 is shown for selecting an explanatory variable that will be a variable value. When the user clicks icon 82 using the control unit 3, a GUI (not shown) for displaying candidate explanatory variables that will be variable values is displayed at the bottom of GUI80. This GUI lists the data types of multiple explanatory variables associated with the selected trained model. When the user selects an explanatory variable that will be a variable value from among the multiple explanatory variables in this GUI, the data type of the selected explanatory variable is written to GUI80. In the example in Figure 11, "amount of Pt (g)" is selected as the explanatory variable that will be a variable value.
[0116] GUI84 is configured to allow input of the lower and upper limits of the variation range for the explanatory variable that will be a variable value. GUI86 is configured to allow input of the step size when the explanatory variable that will be a variable value is continuously varied. The user can set the variation range of the explanatory variable that will be a variable value in GUI84, and set the step size of the explanatory variable that will be a variable value in GUI86. In the example in Figure 11, the lower limit X1_a and upper limit X1_b of the variation range for "Pt amount (g)" and the step size dx1 are set.
[0117] GUI70 can include GUI88 to show recommended ranges for the explanatory variables that represent variable values. These recommended ranges can be set based on the training data used to generate the trained model. For example, data corresponding to the explanatory variables that represent variable values (e.g., the amount of Pt in the mixture (g)) can be extracted from the training data, and the minimum value X1min among the extracted data can be set as the lower limit of the recommended range, while the maximum value X1max can be set as the upper limit of the recommended range. In this case, the user can set the range of variation in GUI84 while referring to the recommended range shown in GUI88.
[0118] The display unit 2 also displays a GUI 74 for setting the values of fixed explanatory variables among the multiple explanatory variables input into the trained model. GUI 74 displays the data type and the value of each data for the fixed explanatory variables in a table format. In the example in Figure 11, GUI 74 shows the values X2, X3, and X4 of the explanatory variables other than "amount of Pt (g)" (such as "amount of Pd (g)", "stirring time (min)", and "firing temperature (°C)"). The values X2, X3, and X4 are set based on the data of the sample to be analyzed.
[0119] After the values of each of the multiple explanatory variables have been set using the procedure described above, clicking GUI72 to instruct the execution of inference will start the inference process. During the inference process, the explanatory variables, which will be the variable values, are continuously varied in predetermined increments, and the target variable corresponding to each value of the explanatory variables is predicted using the trained model.
[0120] The display area 76 of the display unit 2 displays the inference results. As shown in Figure 11, the display area 76 displays graphs 90 and 92 showing the relationship between the explanatory variables, which are the fluctuating values, and the target variable to be displayed. In the example in Figure 11, "NOx purification rate (%)" and "heat resistance performance" are selected as the target variables (physical property data) to be displayed.
[0121] Graph 90 is a two-dimensional graph with the explanatory variable (amount of Pt in the mixture (g)) as the horizontal axis and the target variable (NOx purification rate (%)) as the vertical axis. Graph 92 is a two-dimensional graph with the explanatory variable (amount of Pt in the mixture (g)) as the horizontal axis and the target variable (heat resistance performance) as the vertical axis.
[0122] Graphs 90 and 92 show how the dependent variable changes when the explanatory variable, which is a variable value, is continuously varied within a predetermined range of variation using a predetermined step size. According to Graph 90, the NOx purification rate increases as the amount of Pt increases, but beyond a certain value, the NOx purification rate decreases as the amount of Pt increases. According to Graph 92, the heat resistance performance increases as the amount of Pt increases, but beyond a certain value, the heat resistance performance decreases. Furthermore, comparing Graphs 90 and 92, it can be seen that the amount of Pt at which the NOx purification rate peaks is different from the amount of Pt at which the heat resistance performance peaks.
[0123] In this way, by referring to the inference results displayed on the display unit 2, the user can easily predict how the value of the target variable will change when one explanatory variable is continuously varied. For example, according to graphs 90 and 92, it is possible to predict the amount of Pt suitable for realizing a three-way catalyst with desired physical properties.
[0124] Furthermore, users can obtain graphs 90 and 92 corresponding to various explanatory variables by changing the values of the explanatory variables, which are fixed values in GUI74, and running the inference process again. In addition, users can obtain graphs showing the change in the dependent variable in relation to the continuous change in the explanatory variables for other explanatory variables by changing the type, range, and step size of the explanatory variables, which are variable values in GUI70, and running the inference process again.
[0125] The data representing the inference results (raw data obtained from the inference process and data from graphs 90 and 92) are stored in database 15, along with identification information for identifying the trained model used in the inference process, and information about the explanatory variables input to the trained model. Furthermore, the data representing the inference results stored in database 15 can be output (exported) from the data processing device 1 to an external device via the communication interface 13. The output format can be, for example, CSV (Comma-Separated Values) format, or a format that can be displayed by other AI software, statistical analysis software, or other related software.
[0126] Figure 12 schematically shows a second example of the display of inference results in the display unit 2. Similar to Figure 11, Figure 12 illustrates the inference results displayed in the display unit 2 when "recipe data" is selected as the explanatory variable input to the trained model and "physical property data" is selected as the target variable to be predicted. The sample is a three-way catalyst. The display example in Figure 12 is generated by selecting the trained model used for the inference process, selecting the sample to be analyzed, and selecting the target variable to be displayed.
[0127] In the example shown in Figure 12, there are two explanatory variables that represent the variable values. Display unit 2 displays GUIs 70 and 71 for selecting the explanatory variables that represent the variable values. The configuration of GUIs 70 and 71 is the same as that of GUI 70 shown in Figure 11. Therefore, the user can select the explanatory variables that represent the variable values in GUIs 70 and 71, and set the variation range and step size for the selected explanatory variables.
[0128] In the example in Figure 12, "amount of Pt (g)" and "stirring time (min)" are selected as explanatory variables that represent the fluctuating values. For "amount of Pt (g)", the lower limit X1_a and upper limit X1_b of the fluctuation range and the step size dx1 are set, and for "stirring time (min)", the lower limit X3_a and upper limit X3_b of the fluctuation range and the step size dx3 are set.
[0129] GUI74 shows the values X2 and X4 for explanatory variables other than "amount of Pt (g)" and "stirring time (min)" among the multiple explanatory variables input to the trained model (such as "amount of Pd (g)" and "firing temperature (°C)").
[0130] After the values of each of the multiple explanatory variables are set using the procedure described above, the inference process is executed when GUI72, which instructs the execution of inference, is clicked. In the inference process, one of the two explanatory variables that will be variable is made to vary continuously, while the other explanatory variable is set to a fixed value (for example, the median of the range of variation). Then, the target variable corresponding to each value of the one explanatory variable is predicted using the trained model. In the example in Figure 12, the inference process is executed when "amount of Pt (g)" is set to a variable value and the remaining explanatory variables are set to fixed values, and when "stirring time (min)" is set to a variable value and the remaining explanatory variables are set to fixed values.
[0131] The display area 76 of the display unit 2 displays the inference results. The display area 76 also displays graphs 90 and 94, which show the relationship between the explanatory variables that are variable values and the target variable that is to be displayed. In the example in Figure 12, "NOx purification rate (%)" is selected as the target variable (physical property data) to be displayed.
[0132] Graph 90 is a two-dimensional graph with the explanatory variable (amount of Pt added (g)) as the horizontal axis and the target variable (NOx purification rate (%)) as the vertical axis. Graph 94 is a two-dimensional graph with the explanatory variable (stirring time (min)) as the horizontal axis and the target variable (NOx purification rate (%)) as the vertical axis.
[0133] According to Graph 90, the NOx purification rate increases as the amount of Pt increases, but beyond a certain value, the NOx purification rate decreases as the amount of Pt increases. According to Graph 94, the NOx purification rate increases as the stirring time increases, but beyond a certain value, the NOx purification rate decreases. Furthermore, by comparing Graphs 90 and 94, we can find out the degree of influence of multiple explanatory variables input into the trained model on a single target variable.
[0134] As described above, with the data processing device according to Embodiment 1, the user can easily find out how the target variable changes when the first explanatory variable, which is a variable value, is continuously varied based on the displayed data. Therefore, the usefulness of the inference results can be improved.
[0135] Furthermore, in Embodiment 1, the relationship between the first explanatory variable and the dependent variable is represented by a two-dimensional graph. Based on the displayed two-dimensional graph, the user can easily and visually predict the change in the dependent variable in response to the change in the first explanatory variable.
[0136] In Figures 11 and 12, a configuration is shown in which the number of explanatory variables that fluctuate in a single inference process is 1. However, a configuration in which the number of explanatory variables that fluctuate in a single inference process is 2 or more is also possible. For example, if the number of explanatory variables that fluctuate is 2, each of the two explanatory variables is continuously varied by a predetermined step size, and the target variable corresponding to each value of the two explanatory variables is predicted using the trained model. In this case, as an inference result, a 3D graph can be displayed in which the first explanatory variable of the two explanatory variables is the X-axis, the second explanatory variable is the Y-axis, and the target variable to be displayed is the Z-axis.
[0137] Furthermore, while Figure 12 describes a configuration in which two graphs 90 and 94 are displayed on the display unit 2, corresponding to the two explanatory variables, when there are two explanatory variables that represent the fluctuating values, these two graphs may also be displayed superimposed on each other. Figure 13 schematically shows a third example of the display of the inference results on the display unit 2. The display example in Figure 13 differs from the display example in Figure 12 in how the inference results are displayed. In the display example in Figure 13, a graph 99 showing the relationship between the two explanatory variables that represent the fluctuating values and the target variable to be displayed is displayed in the display area 76. This graph 99 is equivalent to graph 90 in Figure 12 with graph 94 superimposed on it. In this way, the user can relatively evaluate the degree of influence that the two explanatory variables have on one target variable based on a single graph 99.
[0138] Figure 13 shows an example of superimposing two-dimensional graphs 90 and 94, but three-dimensional graphs can also be superimposed. In this way, users can relatively evaluate the degree of influence that four explanatory variables have on one dependent variable based on a single graph.
[0139] [Embodiment 2] As shown in the second display example in Figure 12, when there are multiple explanatory variables that represent variable values, the degree to which each explanatory variable influences the dependent variable differs depending on the type of explanatory variable. Embodiment 2 describes a configuration that compares the degree to which multiple explanatory variables that represent variable values influence the dependent variable and saves the type of explanatory variable that has the greatest influence.
[0140] Figure 14 is a flowchart illustrating the processing procedure for the inference process (S05-S07 in Figure 4) in the data processing device according to Embodiment 2. The flowchart shown in Figure 14 is the same as the flowchart shown in Figure 10, with S37 and S38 added.
[0141] Referring to Figure 14, the inference process is performed according to the same S30-S36 steps as in Figure 10. When the inference results obtained from the inference process are displayed on the display unit 2, S37 calculates the amount of change in the target variable for each of the multiple explanatory variables that represent the variable values, in relation to the change in the explanatory variables. The amount of change in the target variable corresponds to the absolute value of the difference between the maximum and minimum values of the target variable when the explanatory variables are continuously varied within the range of variation set in S32.
[0142] In S38, based on the amount of change in the dependent variable calculated in S37, explanatory variables that have a large influence on the dependent variable are identified.
[0143] In S38, the system may be configured so that the user identifies explanatory variables that have a large influence on the target variable by comparing multiple graphs displayed in the display area 76 of the display unit 2. Alternatively, the data processing device 1 may be configured to identify the explanatory variable that has the largest fluctuation in the dependent variable among multiple explanatory variables that represent fluctuating values, and this variable may have a large influence on the dependent variable.
[0144] The types of explanatory variables that have a significant impact on the dependent variable, as identified in S38, are stored in database 15, linked to the corresponding dependent variable type and information about the trained model used in the inference process. The information about the trained model includes the name of the project to which the trained model is applied. If the sample is a three-way catalyst, the project name might be, for example, "Improving the purification performance of three-way catalysts" or "Improving the heat resistance of three-way catalysts."
[0145] The information stored in database 15 at S38 in Figure 14 can be used in the learning phase. Figure 15 is a flowchart illustrating the processing procedures for generating training data (S02 in Figure 4), machine learning (S03 in Figure 4), and storing the trained model (S04 in Figure 4) in the data processing device according to Embodiment 2. The flowchart shown in Figure 15 is the same as the flowchart shown in Figure 7, with S200 added.
[0146] Referring to Figure 15, in the same S20 as in Figure 7, the UI screen (sample selection screen) is displayed on the display unit 2. The user can select samples to be used to generate training data by operating the sample selection screen using the operation unit 3. Once the sample selection is complete, the selected sample extraction data is generated.
[0147] In S200, the display unit 2 displays information stored in the database 15 in S38 of Figure 14 during past inference processing. Specifically, the display unit 2 displays information about the trained model used in past inference processing (project name), as well as information about the target variable predicted from the trained model and the types of explanatory variables that have a large influence on the target variable.
[0148] Next, the explanatory and target variables to be used to generate the training data are selected using the same S21 and S22 steps as in Figure 7. The display unit 2 shows the UI screens (explanatory variable selection screen and target variable selection screen). As described above, the display unit 2 shows information regarding the name of the project to which the trained model is applied, the type of target variable predicted from the trained model, and the types of explanatory variables that have a large influence on the target variable. Therefore, the user can refer to this information and select the explanatory and target variables that make up the training data according to the project to which the newly generated trained model is applied. For example, the user can select the target and explanatory variables so that they include the target variable and explanatory variables that have a large influence on the target variable, which are associated with the trained model in the same project.
[0149] As described above, the data processing device according to Embodiment 2 can generate a trained model using training data that includes explanatory variables that have a significant influence on the target variable. This makes it possible to improve the usefulness of the trained model for the project.
[0150] [Embodiment 3] In the embodiment 1 described above, a configuration was described in which, during the inference phase, the explanatory variables that will be the variable values among the multiple explanatory variables given to the trained model are selected based on user input to GUIs 70 and 71 (see Figures 11 and 12).
[0151] With the above configuration, as a result of the inference, a graph showing the relationship between the explanatory variable and the dependent variable can be displayed on display unit 2 for the explanatory variable selected by the user. On the other hand, the selection of which explanatory variable to use as a variable depends on the user's experience and skill level. Therefore, even if an explanatory variable has a large impact on the dependent variable, a graph showing the relationship between that explanatory variable and the dependent variable cannot be displayed unless the user selects it as a variable. As a result, there is a concern that the user may overlook important explanatory variables.
[0152] Therefore, Embodiment 3 describes a configuration for displaying inference results for explanatory variables that have a large influence on the target variable. Note that the operation of the data processing device according to Embodiment 3 is basically the same as that of the data processing device according to Embodiment 1 described above, except for the inference process described below.
[0153] (1) First Configuration Example Figure 16 is a flowchart illustrating the inference process in the data processing device according to the first configuration example of Embodiment 3. The flowchart shown in Figure 16 is the same as the flowchart shown in Figure 10, but with S32 replaced by S320 and S350 to S352 added.
[0154] Referring to Figure 16, the same S30 and S31 steps as in Figure 10 select a pre-trained model to be used for inference processing, and once the values of the multiple explanatory variables to be input to the pre-trained model are set, in S320, the range of variation for each of the multiple explanatory variables is set. In other words, the first configuration example differs from the embodiment described above in that all of the multiple explanatory variables to be input to the pre-trained model are variable values.
[0155] The range of variation for each explanatory variable can be set based on the training data used to generate the trained model. For example, data corresponding to each explanatory variable can be extracted from the training data, the minimum value of the extracted data can be set as the lower limit of the range of variation, and the maximum value of that data can be set as the upper limit of the range of variation.
[0156] As shown in Figure 10, step S34 selects the target variable to be displayed from among the target variables predicted by the trained model. The user can select the target variable to be displayed by operating the user interface screen displayed on the display unit 2 using the operation unit 3.
[0157] In S35, the same as in Figure 10, the target variable is predicted by inputting multiple explanatory variables into the trained model selected in S30. In this inference process, one of the multiple explanatory variables is continuously varied in predetermined increments, and the target variable corresponding to each value of that explanatory variable is predicted.
[0158] When one explanatory variable is varied, the values of the other explanatory variables are kept constant. The values of the other explanatory variables are fixed to the values set in S31. These values are based on the recipe data, physical property data, analytical data, and features of the sample being analyzed.
[0159] When the change in the dependent variable is predicted for a change in one explanatory variable, the change in the dependent variable is predicted by changing another explanatory variable. Once the change in the dependent variable is predicted for all of the explanatory variables, the inference process in S35 is completed.
[0160] When the inference process in S35 is completed, a graph to be displayed as the inference result is selected from among the inference results of multiple dependent variables corresponding to multiple explanatory variables. Specifically, first, in S350, the amount of change in the dependent variable in relation to the change in the explanatory variable is calculated for each of the multiple explanatory variables. The amount of change in the dependent variable corresponds to the absolute value of the difference between the maximum and minimum values of the dependent variable when the explanatory variable is continuously varied within the range of variation set in S320.
[0161] Next, in S351, the graphs to be displayed are selected based on the amount of change in the dependent variable calculated in S350. In S351, graphs showing the relationship between the explanatory variables and the dependent variable are selected in order from the largest amount of change in the dependent variable among the multiple inference results. The number of graphs to be selected for display can be set in advance by the user. For example, a predetermined number of graphs can be selected to be displayed, counting from the one with the largest amount of change in the dependent variable. Alternatively, graphs in which the amount of change in the dependent variable is greater than or equal to a predetermined value can be selected to be displayed.
[0162] In S352, the display order of the graphs selected in S351 is set. Specifically, the graph with the largest variation in the dependent variable is placed first, and the graphs are displayed in descending order of the variation in the dependent variable.
[0163] In S36, the inference results obtained from the inference processing in S35 are displayed on the display unit 2. The graphs selected as display targets are displayed on the display unit 2 according to the display order set in S352.
[0164] Figure 17 is a schematic diagram showing an example of how the inference results are displayed in the first configuration example. In Figure 17, the display area 76 for displaying the inference results is extracted from the display unit 2 shown in Figure 11 and schematically shown.
[0165] In the example shown in Figure 17, the display area 76 of the display unit 2 displays several graphs 94, 96, and 98 that show the relationship between the explanatory variables, which are the fluctuating values, and the dependent variable, which is the target variable to be displayed. Graph 94 is a two-dimensional graph with the explanatory variable X1 on the horizontal axis and the dependent variable Y1 on the vertical axis. Graph 96 is a two-dimensional graph with the explanatory variable X3 on the horizontal axis and the dependent variable Y1 on the vertical axis. Graph 98 is a two-dimensional graph with the explanatory variable X7 on the horizontal axis and the dependent variable Y1 on the vertical axis.
[0166] In graphs 94, 96, and 98, ΔY1 represents the amount of change in the dependent variable Y1 when the corresponding explanatory variable is continuously varied within its range of variation. Graph 94 has the largest amount of change ΔY1, graph 96 has the second largest amount of change ΔY1, and graph 98 has the smallest amount of change ΔY1. In other words, the display area 76 shows multiple graphs 94, 96, and 98 side by side, in descending order of the amount of change ΔY1 of the dependent variable Y1.
[0167] As mentioned above, the number of graphs to be displayed in the display area 76 can be set in advance by the user. For example, if the number of graphs to be displayed in the display area 76 is set to N (N≧1), then a total of N graphs will be displayed in the display area 76 in order from those with the largest variation ΔY1 of the dependent variable Y1.
[0168] Alternatively, the system may be configured to display graphs in the display area 76 where the variation amount ΔY1 of the dependent variable Y1 is greater than or equal to a predetermined value. In this case, graphs where the variation amount ΔY1 of the dependent variable Y1 is greater than or equal to the predetermined value are displayed in the display area 76 in order from those with the largest variation amount ΔY1.
[0169] As explained above, in the first configuration example, among the multiple explanatory variables given to the trained model, those in which the amount of change in the target variable is large in relation to the change in the explanatory variable are preferentially selected, and a graph showing the relationship between the selected explanatory variables and the target variable is displayed on the display unit 2. With this, regardless of the user's experience level or skill level, explanatory variables with a large influence on the target variable are automatically selected, and the inference result of the change in the target variable in relation to the change in that explanatory variable is displayed. Therefore, the possibility of the user overlooking important explanatory variables can be reduced.
[0170] Furthermore, the display unit 2 displays graphs in order from those showing the largest variation in the dependent variable in response to the variation in the explanatory variables, thus effectively displaying graphs for explanatory variables that have a significant impact on the dependent variable. This improves the usefulness of the inference results.
[0171] (2) Second Configuration Example In the first configuration example described above, all of the multiple explanatory variables given to the trained model are treated as variable values, raising concerns that the computational load required for inference processing will increase as the number of explanatory variables increases. Therefore, in the second configuration example and the third configuration example described later, a configuration is described in which the data processing device 1 automatically selects the explanatory variables that will be treated as variable values before executing the inference processing.
[0172] Figure 18 is a flowchart illustrating the inference process in the data processing device according to the second configuration example of Embodiment 3. The flowchart shown in Figure 18 is the same as the flowchart shown in Figure 10, but with S32 replaced by S321 and S353 added.
[0173] Referring to Figure 18, in the same S30 and S31 steps as in Figure 10, a trained model to be used for inference processing is selected, and the values of the multiple explanatory variables to be input to the trained model are set. Then, in S321, the inference unit 32 of the data processing device 1 obtains the importance of each explanatory variable input to the trained model.
[0174] The importance of each explanatory variable quantifies how much that particular explanatory variable contributes to the model's performance. Specifically, the importance of each explanatory variable can be calculated by applying a decision tree algorithm to multiple explanatory variables. Any known algorithm can be used as the decision tree algorithm; for example, a random forest can be used.
[0175] The inference unit 32 selects explanatory variables to be variable values based on the importance of each explanatory variable. Specifically, the inference unit 32 prioritizes selecting explanatory variables with high importance as explanatory variables to be variable values. The number of explanatory variables to be variable values can be set in advance by the user. For example, a predetermined number of explanatory variables can be selected as variable values, starting from the one with the highest importance. Alternatively, explanatory variables whose importance is equal to or greater than a predetermined value can be selected as variable values.
[0176] In S321, when explanatory variables that will have variable values are selected, in S33, the inference unit 32 sets the variable range for each explanatory variable. The variable range for each explanatory variable can be set based on the training data used to generate the trained model. For example, data corresponding to each explanatory variable can be extracted from the training data, the minimum value of the extracted data can be set as the lower limit of the variable range, and the maximum value of the extracted data can be set as the upper limit of the variable range.
[0177] As shown in Figure 10, step S34 selects the target variable to be displayed from among the target variables predicted by the trained model. The user can select the target variable to be displayed by operating the user interface screen displayed on the display unit 2 using the operation unit 3.
[0178] In S35, the same as in Figure 10, the trained model selected in S30 is input with the explanatory variables set in S31 and S321, and the target variable is predicted. In this inference process, the inference unit 32 continuously varies one of the explanatory variables that will be variable values in predetermined increments, and predicts the target variable corresponding to each value of that explanatory variable.
[0179] When one explanatory variable is varied, the values of the other explanatory variables are kept fixed. For example, the values of the other explanatory variables are fixed to the values set in S31. These values are based on the recipe data, physical property data, analysis data, and features of the sample being analyzed. When the inference unit 32 predicts a change in the target variable in response to a change in one explanatory variable, it varies another explanatory variable to predict a change in the target variable. Once a change in the target variable is predicted for all explanatory variables that will have varying values, the inference process in S35 is completed.
[0180] When the inference process in S35 is completed, in S353, the display data generation unit 34 sets the display order of the graphs to be displayed as inference results based on the importance of each explanatory variable. Specifically, the graph showing the inference result for the explanatory variable with the highest importance is set as the first graph, and the display order of the graphs is set so that the explanatory variables are arranged in order from the most important to the least important.
[0181] In S36, the display data generation unit 34 displays the inference results obtained from the inference processing in S35 on the display unit 2. The display unit 2 displays the graphs selected as display targets according to the display order set in S352.
[0182] As explained above, in the second configuration example, among the multiple explanatory variables given to the trained model, explanatory variables with high importance in the trained model are selected as fluctuation values, and a graph showing the relationship between the selected explanatory variables and the target variable is displayed on the display unit 2. This means that the inference results of the fluctuation of the target variable in relation to the fluctuation of explanatory variables that have a large influence on the target variable are displayed, regardless of the user's experience level or skill level. Therefore, the possibility of the user overlooking important explanatory variables can be reduced.
[0183] Furthermore, the display unit 2 displays graphs of explanatory variables in order of their influence on the dependent variable, from most to least. This allows for the effective display of graphs related to explanatory variables that have a significant impact on the dependent variable. Therefore, the usefulness of the inference results can be improved.
[0184] (3) Third Configuration Example Figure 19 is a flowchart illustrating the inference process in the data processing device according to the third configuration example of Embodiment 3. The flowchart shown in Figure 19 is the same as the flowchart shown in Figure 10, but with S32 replaced by S322 and S323, and S354 added.
[0185] Referring to Figure 19, the same S30 and S31 steps as in Figure 10 select a pre-trained model to be used for inference processing. Once the values of the multiple explanatory variables to be input to the pre-trained model are set, the inference unit 32 selects the explanatory variable that will be the variable from among the multiple explanatory variables. In the third configuration example, the explanatory variable that will be the variable is selected using principal component analysis performed on the multiple explanatory variables.
[0186] Principal component analysis (PCA) is generally performed as a data preprocessing step to reduce the dimensionality of large datasets. By performing PCA, multiple explanatory variables are aggregated into a smaller number of composite variables (principal components). The results of PCA are obtained as principal component scores, which are the transformed values corresponding to the original explanatory variables, and principal component loadings, which are the weights of the explanatory variables for each principal component score.
[0187] In S322, one principal component is selected from a predetermined number of principal components obtained by principal component analysis. For example, the inference unit 32 can select one principal component according to user input. In this case, the user can select one principal component based on the contribution rate of each principal component. The contribution rate of a principal component is obtained by dividing the eigenvalues of each principal component by their sum, and indicates what proportion of the overall variation each principal component accounts for. Alternatively, the inference unit 32 may be configured to select the first principal component with the highest contribution rate, regardless of user input.
[0188] In S323, the inference unit 32 selects the explanatory variables that will be the variable values based on the weights (principal component loadings) of each explanatory variable in the one principal component selected in S322.
[0189] Specifically, the i-th principal component z is a composite obtained by multiplying the original p variables X1, X2, ..., Xp by a weight w (principal component loading), and can be expressed by the following equation. Note that the sum of the squares of the p wj (j=1,2,...p) is 1.
[0190] z = w1X1 + w2X2 + ... + wpXp In the above equation, the larger the absolute value of the weight w (principal component loading), the greater the contribution of the corresponding explanatory variable X to the principal component z; in other words, it is an explanatory variable that characterizes the principal component. Therefore, in S323, the inference unit 32 preferentially selects explanatory variables with large weights (principal component loadings) as explanatory variables that will be the variable values.
[0191] The number of explanatory variables that will be used as variable values can be pre-set by the user. For example, a predetermined number of explanatory variables can be used as variable values, starting from the one with the highest weight (principal component loading). Alternatively, explanatory variables whose weight (principal component loading) is equal to or greater than a predetermined value can be used as variable values.
[0192] In S323, when explanatory variables that will have variable values are selected, the inference unit 32 sets the range of variation for each explanatory variable in S33. The range of variation for each explanatory variable can be set based on the training data used to generate the trained model. For example, data corresponding to each explanatory variable can be extracted from the training data, the minimum value of the extracted data can be set as the lower limit of the range of variation, and the maximum value of the extracted data can be set as the upper limit of the range of variation.
[0193] As shown in Figure 10, step S34 selects the target variable to be displayed from among the target variables predicted by the trained model. The user can select the target variable to be displayed by operating the user interface screen displayed on the display unit 2 using the operation unit 3.
[0194] In S35, the same as in Figure 10, the inference unit 32 predicts the target variable by inputting the explanatory variables set in S31 and S321 to the trained model selected in S30. In this inference process, one of the explanatory variables that will fluctuate is continuously varied in predetermined increments, and the target variable corresponding to each value of that explanatory variable is predicted. When that one explanatory variable is varied, the values of the other explanatory variables are fixed. For example, the values of the other explanatory variables are fixed to the values set in S31. When the inference unit 32 predicts a change in the target variable for a change in one explanatory variable, it varies another explanatory variable and predicts a change in the target variable. When the change in the target variable is predicted for all of the explanatory variables that will fluctuate, the inference process in S35 ends.
[0195] Once the inference process in S35 is complete, in S353, the display data generation unit 34 sets the display order of the graphs to be displayed as inference results based on the weights (principal component loadings) of each explanatory variable. Specifically, the graph showing the inference result for the explanatory variable with the highest weight (principal component loading) is set as the first graph, and the display order of the graphs is set so that the explanatory variables are arranged in descending order of their weights (principal component loadings).
[0196] In S36, the display data generation unit 34 displays the inference results obtained from the inference processing in S35 on the display unit 2. The display unit 2 displays the graphs selected as display targets according to the display order set in S354.
[0197] As explained above, in the third configuration example, among the multiple explanatory variables given to the trained model, explanatory variables with large weights (principal component loadings) for specific principal components are selected as fluctuation values, and a graph showing the relationship between the selected explanatory variables and the target variable is displayed on the display unit 2. This allows the inference results of the fluctuation of the target variable in relation to the fluctuation of explanatory variables that have a high contribution to the principal components to be displayed, regardless of the user's experience level or skill level. Therefore, the possibility of the user overlooking important explanatory variables can be reduced.
[0198] Furthermore, since the display unit 2 displays graphs of explanatory variables in descending order of their principal component loadings, it is possible to effectively display graphs of explanatory variables that contribute significantly to the principal components. Therefore, the usefulness of the inference results can be improved.
[0199] [Embodiment 4] In the learning phase, supervised learning is performed, using the explanatory variables of the generated training data as input to the learning model and the target variable of the training data as the ground truth data for the output of the learning model. Embodiment 4 describes a configuration in which the user can select the learning model. The operation of the data processing device according to Embodiment 4 is basically the same as the operation of the data processing device according to Embodiment 1 described above, except for the learning process described below.
[0200] Figure 20 is a flowchart illustrating the processing procedures for generating training data (S02 in Figure 4), machine learning (S03 in Figure 4), and storing the trained model (S04 in Figure 4) in the data processing device according to Embodiment 4. The flowchart shown in Figure 20 is the same as the flowchart shown in Figure 7, with the addition of S230.
[0201] Referring to Figure 20, once the training data is generated by steps S20-S23, the same as in Figure 7, a learning model to be used for machine learning is selected in step S230. The display unit 2 shows a UI screen (model selection screen). The model selection screen is a UI screen for the user to select a learning model. Multiple learning models are listed on the model selection screen. These multiple learning models are, for example, polynomial regression models, with polynomial degrees differing from each other. In addition, the user can add interaction terms between terms, as well as logarithmic and exponential terms.
[0202] In machine learning, as you increase the degree of the polynomial or make the model more complex by including interaction terms, logarithms, and exponents, the accuracy on the training data improves, but the accuracy on unknown data can decrease, leading to "overfitting."
[0203] In Embodiment 4, by making the learning model more complex, it is possible to represent the complex relationship between the explanatory variables and the target variable that make up the training data. On the other hand, by simplifying the learning model, it is possible to avoid the overfitting mentioned above. In S230 of Figure 20, the user can select the order of the learning model after weighing these advantages and disadvantages, thereby enabling optimal machine learning.
[0204] [Embodiment 5] In the embodiment 1 described above, a sample list (Figure 6) is generated based on the data stored in the database 15 (S01 in Figure 4 and Figure 5). At this time, the characteristic quantities of the sample are extracted by analyzing the sample analysis data using dedicated data analysis software (S13 in Figure 5). The characteristic quantities include the peak area for a predetermined mass number obtained by analyzing the chromatogram obtained by GC-MS, the abundance ratio of a predetermined substance obtained by analyzing the NMR spectrum obtained by NMR, the particle diameter and average particle diameter of particles present in the three-way catalyst obtained by analyzing the SEM image obtained by SEM, and the particle diameter of particles present in the three-way catalyst obtained by analyzing the TEM image obtained by TEM.
[0205] In the process of extracting these features, changing the conditions for processing the analysis data or changing the conditions for calculating the features can result in different values for the extracted features, even for the same sample of analysis data. In this case, the training data generated from the sample list will differ depending on the processing conditions and feature calculation conditions of the analysis data, and therefore the trained model generated from the training data will also differ depending on the processing conditions and feature calculation conditions of the analysis data. If the trained model is different, the predicted target variable may differ in the inference process, even if the explanatory variables given to the trained model are the same. Therefore, the degree to which the explanatory variables influence the target variable, as derived from the inference results, may also differ depending on the differences in the trained model.
[0206] Embodiment 5 describes a configuration for obtaining appropriate data processing conditions for considering the degree of influence of explanatory variables on the target variable. Below, the processing performed by the data processing device according to Embodiment 5 in the learning phase and the inference phase will be described.
[0207] <Learning Phase> Figure 21 shows an example of a sample list structure. Figure 21 shows an example of a sample list structure when the sample is a three-way catalyst. The sample list shown in Figure 21 differs from the sample list shown in Figure 6 in that it includes multiple features extracted by processing the sample analysis data under multiple processing conditions.
[0208] In the example shown in Figure 21, the features include peak area 1 for the first mass number and peak area 2 for the second mass number, which were obtained by analyzing the chromatogram acquired by GC-MS.
[0209] Peak area 1 consists of three values Pa, Pb, and Pc, each with different data processing conditions (methods for calculating peak area). Pa is peak area 1 calculated using processing condition A, Pb is peak area 1 calculated using processing condition B, and Pc is peak area 1 calculated using processing condition C. Peak area 2 consists of three values Qa, Qb, and Qc, each with different data processing conditions (methods for calculating peak area). Qa is peak area 2 calculated using processing condition A, Qb is peak area 2 calculated using processing condition B, and Qc is peak area 2 calculated using processing condition C.
[0210] In the process of generating training data (S02 in Figure 4), when a sample to be used to generate training data is selected (S20 in Figure 7), data contained in the row of the selected sample is extracted from the sample list shown in Figure 21, and a selected sample extraction table is generated. Figure 22 shows an example of the configuration of the selected sample extraction table. Three types of selected sample extraction tables are shown in Figure 22.
[0211] Selective sample extraction table A is composed of peak area 1 and peak area 2, which are features obtained using processing condition A. Selective sample extraction table B is composed of peak area 1 and peak area 2, which are features obtained using processing condition B. Selective sample extraction table C is composed of peak area 1 and peak area 2, which are features obtained using processing condition C.
[0212] In other words, although selection sample extraction tables A to C were generated from the same sample analysis data, the data processing conditions for extracting features from that analysis data differ from one another. As a result, although the data types in selection sample extraction tables A to C are the same, the data values are different from one another.
[0213] Once the explanatory and dependent variables to be used to generate training data are selected (S21, S22 in Figure 7), data matching the explanatory variables and data matching the dependent variable are extracted from each of the selected sample extraction tables A to C, generating three types of training data tables. Each of the three types of training data tables is then populated with data matching the explanatory variables and data matching the dependent variable for each sample. Based on these three generated training data tables, training data A to C are then generated.
[0214] Three types of trained models are generated by supervised learning using training data A to C. Trained model MODEL1a is a trained model generated by machine learning using training data A. Trained model MODEL1b is a trained model generated by machine learning using training data B. Trained model MODEL1c is a trained model generated by machine learning using training data C.
[0215] The generated trained models MODEL1a, MODEL1b, and MODEL1c are registered in the trained model list stored in database 15. Figure 23 shows an example of the structure of the trained model list. The trained model list shown in Figure 23 differs from the trained model list shown in Figure 9 in that the identification information for identifying the training data includes the data processing conditions used to extract features from the analysis data.
[0216] The trained models MODEL1a, MODEL1b, and MODEL1c share the same project name to which the trained model is applied, the same sample information used to generate the training data, and the same types of data selected for the explanatory and dependent variables. However, they differ in the processing conditions of the analytical data used to generate the data (for example, the method for calculating the peak area of the chromatogram).
[0217] <Inference Phase> Figure 24 is a flowchart illustrating the processing procedure for the inference process (S05 to S07 in Figure 4) in the data processing device according to Embodiment 5. The flowchart shown in Figure 24 is obtained by replacing S30 with S300 and S36 with S360 to S362 in the flowchart shown in Figure 10.
[0218] Referring to Figure 24, in S300, multiple pre-trained models to be used for inference processing are selected. The display unit 2 displays a generated UI screen (pre-trained model selection screen) based on the list of pre-trained models (Figure 23) stored in the database 15. The user can select multiple pre-trained models that differ only in their processing conditions for the analysis data by operating the UI screen using the operation unit 3. In the following, we will assume that pre-trained models MODEL1a, MODEL1b, and MODEL1c are selected.
[0219] In S31, the same as in Figure 10, the values of the explanatory variables to be input to each trained model are set. In S32, the "explanatory variables that will have variable values" are selected from among the multiple explanatory variables. In S33, the range of variation is set for the explanatory variables that will have variable values. In S34, the target variable to be displayed is selected from among the target variables predicted by each trained model.
[0220] In S35, the same as in Figure 10, the target variable is predicted by inputting the explanatory variables set in S31 to S33 for each of the multiple trained models selected in S300. In this inference process, for each trained model, how the target variable will change is predicted in relation to the values of some of the continuously fluctuating explanatory variables among the multiple explanatory variables.
[0221] In S360, multiple inference results obtained from the inference processing in S35 are displayed on the display unit 2. The display unit 2 displays graphs corresponding to each of the multiple trained models, showing the variation in some explanatory variables for the target variable selected as the display target.
[0222] Figure 25 is a schematic diagram showing an example of the display of multiple inference results in the display unit 2. In Figure 25, the display area 76 of the display unit 2 that displays the inference results is extracted and shown.
[0223] The display area 76 includes a display area 76A that displays the inference results of the inference process using the trained model MODEL1a, a display area 76B that displays the inference results of the inference process using the trained model MODEL1b, and a display area 76C that displays the inference results of the inference process using the trained model MODEL1c.
[0224] Each of the display areas 76A, 76B, and 76C displays graphs 90 and 92, which show the relationship between the explanatory variable that represents the variable and the target variable that is displayed. In the example in Figure 25, "Peak Area 1" is selected as the explanatory variable that represents the variable, and "NOx Purification Rate (%)" and "Heat Resistance Performance" are selected as the target variables that are displayed. Graphs 90 and 92 show how the NOx Purification Rate (%) and Heat Resistance Performance change when Peak Area 1 is continuously varied within a predetermined range of variation.
[0225] Comparing graph 90 across display areas 76A, 76B, and 76C reveals that, due to differences in the trained models, even with the same sample being analyzed, there are differences in the magnitude of the influence of the explanatory variables on the dependent variable. The same can be said for graph 92.
[0226] By comparing graphs 90 and 92 displayed in these three display areas 76A, 76B, and 76C, the user can select a pre-trained model that they deem appropriate for examining the relationship between the explanatory and dependent variables.
[0227] Figure 25 shows an example where three display areas 76A, 76B, and 76C are displayed side by side. However, it is also possible to configure the system so that only one display area is displayed initially, and the display areas 76A, 76B, and 76C are switched according to the user's operation.
[0228] Returning to Figure 24, in S361, a suitable pre-trained model is selected by the user. For example, the display unit 2 displays a UI screen for selecting a suitable pre-trained model. The user can select a suitable pre-trained model by operating the UI screen using the operation unit 3. In S362, information about the suitable pre-trained model selected by the user is stored in the database 15. The information about the suitable pre-trained model includes the name of the project to which the pre-trained model is applied, the sample information used to generate the training data, the types of data selected for the explanatory and dependent variables, and the processing conditions for the analysis data used to extract the data.
[0229] By storing information about the appropriate pre-trained model selected based on the inference results in database 15, in the future, when generating a pre-trained model using samples similar to those analyzed in the current inference process, it will be possible to retrieve the data processing conditions used to generate the appropriate pre-trained model from database 15 and present them to the user. Similar samples mean that at least one of the sample's recipe data, physical property data, and analysis data is the same or similar.
[0230] Figure 26 is a flowchart illustrating the processing procedure for generating the sample list (S01 in Figure 4). The flowchart shown in Figure 26 is the same as the flowchart shown in Figure 5, with S120 and S121 added.
[0231] Once sample information, physical property data, and analysis data for the sample are acquired through steps S10-S12 as shown in Figure 5, step S120 retrieves the processing conditions for analysis data of samples similar to the current sample from database 15 by referencing this data, and displays them on display unit 2. The processing conditions for analysis data retrieved from database 15 include information about trained models that the user has previously deemed suitable for examining the relationship between explanatory variables and the dependent variable.
[0232] In S121, the processing conditions for the analysis data acquired in S12 are set via the operation unit 3. In S13, the same as in Figure 5, the analysis data acquired in S12 is processed using the processing conditions set in S121 to extract the sample features. In S14, the acquired sample information, sample physical property data, and sample analysis data and features are entered into the sample list (Figure 6). In S15, the sample list is registered in the database 15 with the sample list identification information attached.
[0233] In the learning phase, features are extracted from sample analysis data using data processing conditions deemed appropriate for considering the relationship between explanatory variables and the target variable, and a sample list is generated. Then, a trained model is generated using the training data generated based on this sample list. In this way, in the inference phase, the relationship between the explanatory variables given to the trained model and the target variable predicted by the trained model will follow what the user deems appropriate. Therefore, it is possible to improve the usefulness of the inference results.
[0234] [Other configuration examples] (1) In the above-described embodiment, an example configuration was described in which a UI screen for accepting user input when inference processing is performed and the inference results (see Figures 11 and 12) are displayed on a display unit 2 connected to the data processing device 1. However, instead of the display unit 2, an information terminal such as a desktop personal computer (PC), a notebook PC, or a mobile terminal (tablet terminal, smartphone) may be connected to the data processing device 1, and the UI screen and inference results may be displayed on the information terminal.
[0235] (2) In the above-described embodiment, an example configuration was explained in which the data types of the explanatory variables input to the pre-trained model and the data types of the target variable predicted by the pre-trained model are automatically determined by selecting the pre-trained model to be used for inference processing. However, it is also possible to configure the system so that the pre-trained model to be used for inference processing is automatically determined by selecting the data types of the explanatory variables input to the pre-trained model and the data types of the target variable to be predicted. In this case, the user can select the data types of the explanatory variables and the data types of the target variable by checking the selection icons on the UI screen (target variable selection screen and explanatory variable selection screen) displayed on the display unit 2 using the operation unit 3. The inference unit 32 can refer to the pre-trained list (Figure 9) stored in the database 15 and determine the pre-trained model associated with the training data, which includes the selected data types of the explanatory variables and the data types of the target variable, as the pre-trained model to be used for estimation processing.
[0236] (3) In the above-described embodiment, the data processing device 1 is shown as having a learning unit 30 and an inference unit 32 (see Figure 3), but the learning unit 30 and the inference unit 32 may be provided as separate units.
[0237] [Pattern] Those skilled in the art will understand that the above-described exemplary embodiments are specific examples of the following embodiments.
[0238] (Section 1) A data processing device according to one embodiment comprises an inference unit that predicts a target variable from a plurality of explanatory variables using a trained model, and a display data generation unit that generates data for displaying the inference results by the inference unit. The inference unit sets a first explanatory variable selected from the plurality of explanatory variables as a variable value, while setting a second explanatory variable other than the first explanatory variable as a fixed value. The inference unit predicts the target variable when the first explanatory variable is continuously varied within a predetermined range of variation using a trained model. The display data generation unit generates data showing the variation of the target variable in response to the variation of the first explanatory variable.
[0239] According to the data processing device described in paragraph 1, the user can easily predict how the target variable will change when the first explanatory variable, which is a variable value, is continuously varied based on the displayed data. Therefore, the usefulness of the inference results can be increased.
[0240] (Section 2) In the data processing device described in Section 1, the display data generation unit generates a two-dimensional graph with the first explanatory variable as the first axis and the target variable as the second axis.
[0241] According to the data processing device described in paragraph 2, the user can easily and visually predict the variation of the dependent variable in relation to the variation of the first explanatory variable based on the displayed two-dimensional graph.
[0242] (Section 3) In the data processing device described in Section 1 or 2, the inference unit selects two or more first explanatory variables from among a plurality of explanatory variables. For each of the two or more selected first explanatory variables, the inference unit uses a trained model to predict the target variable when the first explanatory variables are continuously changed within a range of variation. A display unit is connected to the data processing device. The display data generation unit generates two or more two-dimensional graphs corresponding to each of the two or more first explanatory variables. The display data generation unit displays the two or more generated two-dimensional graphs on the display unit so that they overlap each other.
[0243] According to the data processing device described in paragraph 3, it becomes possible to relatively evaluate the degree of influence that each of two or more first explanatory variables has on the dependent variable based on two or more two-dimensional graphs superimposed on the display unit.
[0244] (Section 4) In the data processing device described in Sections 1 to 3, the display data generation unit is configured to provide a first user interface for selecting a first explanatory variable and setting a range of variation. The first user interface includes information regarding a recommended range for the range of variation.
[0245] The data processing device described in paragraph 4 can improve user convenience in inference processing.
[0246] (Section 5) In the data processing device described in Section 4, the trained model is a model generated by machine learning using training data that takes multiple explanatory variables as input and the target variable as the correct output. The recommended range is set based on the value of the first explanatory variable included in the training data.
[0247] According to the data processing device described in Section 5, it is possible to provide the user with a recommended range in which the accuracy of the inference results is guaranteed.
[0248] (Clause 6) In the data processing device described in paragraph 4 or 5, the display data generation unit further provides a second user interface for setting the value of a second explanatory variable. The data processing device described in paragraph 6 can improve user convenience in the inference process.
[0249] (Section 7) In the data processing device described in Sections 1 to 6, the trained model is a model generated by machine learning using training data that takes multiple explanatory variables as input and the target variable as the correct output. The inference unit selects at least one first explanatory variable from among the multiple explanatory variables based on the importance of each explanatory variable in the trained model. For each of the selected at least one first explanatory variable, the inference unit uses the trained model to predict the target variable when the first explanatory variable is continuously varied within a range of variation.
[0250] According to the data processing device described in Section 7, among the multiple explanatory variables provided to the trained model, explanatory variables with high importance in the trained model are selected as fluctuation values, and a graph showing the relationship between the selected explanatory variables and the target variable is generated. This allows for the inference of the change in the target variable in relation to the change in explanatory variables that have a large impact on the target variable, regardless of the user's experience level or skill level. Therefore, the possibility of the user overlooking important explanatory variables can be reduced.
[0251] (Clause 8) In the data processing device described in Clause 7, the range of variation is set based on the value of the first explanatory variable included in the training data. According to the data processing device described in Clause 8, a range of variation can be set in which the accuracy of the inference result is guaranteed.
[0252] (Clause 9) In the data processing device described in paragraph 7 or 8, a display unit is connected to the data processing device. The display data generation unit generates at least one data corresponding to at least one first explanatory variable. The display data generation unit displays the generated at least one data on the display unit in order from the most important of the corresponding first explanatory variables.
[0253] According to the data processing device described in paragraph 9, the display unit shows graphs of explanatory variables in descending order of their influence on the dependent variable. Therefore, graphs of explanatory variables that have a large influence on the dependent variable can be effectively displayed. This improves the usefulness of the inference results.
[0254] (Clause 10) In the data processing device described in paragraphs 1 to 6, the inference unit selects at least one first explanatory variable from among the multiple explanatory variables based on the absolute values of the principal component loadings of each explanatory variable for a specific principal component obtained by principal component analysis of the multiple explanatory variables. For each of the selected at least one first explanatory variable, the inference unit uses a trained model to predict the target variable when the first explanatory variable is continuously varied within the range of variation.
[0255] According to the data processing device described in Section 10, among the multiple explanatory variables provided to the trained model, explanatory variables with large weights (principal component loadings) for specific principal components are selected as fluctuation values, and a graph showing the relationship between the selected explanatory variables and the target variable is generated. This allows for inference results of the fluctuation of the target variable in relation to the fluctuation of explanatory variables that contribute highly to the principal components, regardless of the user's experience level or skill level. Therefore, the possibility of the user overlooking important explanatory variables can be reduced.
[0256] (Section 11) In the data processing device described in Section 10, the trained model is a model generated by machine learning using training data that takes multiple explanatory variables as input and the target variable as the correct output. The range of variation is set based on the value of the first explanatory variable included in the training data.
[0257] According to the data processing device described in paragraph 11, it is possible to set a variation range in which the accuracy of the inference results is guaranteed.
[0258] (Clause 12) In the data processing device described in paragraph 10 or 11, a display unit is connected to the data processing device. The display data generation unit generates at least one data corresponding to at least one first explanatory variable. The display data generation unit displays the generated at least one data on the display unit in order from the highest principal component loading of the corresponding first explanatory variable.
[0259] According to the data processing device described in paragraph 12, the display unit shows graphs of explanatory variables in descending order of their principal component loadings, thus effectively displaying graphs of explanatory variables that contribute significantly to the principal components. Therefore, the usefulness of the inference results can be improved.
[0260] (Section 13) In the data processing device described in Sections 1 to 6, a display unit is connected to the data processing device. The inference unit selects each of the multiple explanatory variables in order as the first explanatory variable. For each selected explanatory variable, the inference unit uses a trained model to predict the target variable when the first explanatory variable is continuously varied within the range of variation. The display data generation unit generates multiple data points corresponding to each of the multiple explanatory variables. The display data generation unit displays the generated multiple data points on the display unit in order from those with the largest variation in the target variable.
[0261] According to the data processing device described in Section 13, among the multiple explanatory variables provided to the trained model, those in which the amount of variation in the target variable is large in relation to the variation in the explanatory variables are preferentially selected, and a graph showing the relationship between the selected explanatory variables and the target variable is generated. As a result, explanatory variables with a large influence on the target variable are automatically selected, regardless of the user's experience level or skill level, and the inference results of the change in the target variable in relation to the change in those explanatory variables are displayed. Therefore, the possibility of the user overlooking important explanatory variables can be reduced. Furthermore, since the display unit shows graphs in order from the amount of variation in the target variable in relation to the change in the explanatory variables, graphs for explanatory variables with a large influence on the target variable can be effectively displayed. Thus, the usefulness of the inference results can be improved.
[0262] (Section 14) In the data processing device described in Sections 1 to 13, the inference unit selects two or more first explanatory variables from among multiple explanatory variables. For each of the two or more selected first explanatory variables, the inference unit uses a trained model to predict the target variable when the first explanatory variable is continuously changed within a range of variation. The display data generation unit generates two or more data points corresponding to each of the two or more first explanatory variables. The data processing device further includes a database for storing the type of first explanatory variable that has the greatest influence on the target variable among the two or more first explanatory variables, linked to information about the project to which the trained model is applied.
[0263] According to the data processing apparatus described in claim 14, when the learning model is to be learned next time, depending on the project to which the learning model for inputting teacher data while referring to the information stored in the database is applied, explanatory variables and objective variables can be selected.
[0264] (Claim 15) The data processing apparatus described in claim 14 further includes a teacher data generation unit that generates teacher data having a plurality of explanatory variables as inputs and an objective variable as a correct output, and a learning unit that generates a learned model by machine learning using the teacher data. The teacher data generation unit presents the project and the type of the first explanatory variable associated with the project to the user.
[0265] According to the data processing apparatus described in claim 15, the user can select explanatory variables and objective variables depending on the project to which the learning model for inputting teacher data is applied. For example, the user can select the objective variable of the learned model having the same project and the explanatory variable having a large influence degree on the objective variable as the objective variable and the explanatory variable, respectively. According to this, since the learned model is generated using the explanatory variable having a large influence degree on the objective variable as the teacher data, the usefulness of the learned model for the project can be enhanced.
[0266] (Claim 16) The data processing apparatus described in claims 1 to 13 further includes a teacher data generation unit that generates teacher data having a plurality of explanatory variables as inputs and an objective variable as a correct output, a learning unit that generates a learned model by machine learning using the teacher data, and a database for storing the learned model in association with the teacher data.
[0267] According to the data processing apparatus described in claim 16, machine learning of the learning model and inference using the learned model can be executed on one device.
[0268] (Item 17) In the data processing apparatus according to Item 16, the teacher data generation unit generates a plurality of teacher data so as to each include a plurality of feature amounts extracted using a plurality of different data processing conditions from one data group. The learning unit generates a plurality of learned models respectively corresponding to the plurality of teacher data by machine learning. The learning unit stores each of the generated plurality of learned models in a database in association with the corresponding data processing condition.
[0269] According to the data processing apparatus described in Item 17, a plurality of teacher data with different data processing conditions are generated from one data group, and a plurality of learned models are respectively generated using this plurality of teacher data. By performing inference by giving a common explanatory variable to this plurality of learned models, the relationship between the data processing condition and the influence degree that the explanatory variable gives to the objective variable can be known.
[0270] (Item 18) In the data processing apparatus according to Item 17, the inference unit predicts the objective variable when the first explanatory variable is continuously varied within the variable range using each of the plurality of learned models. The display data generation unit generates a plurality of data indicating the variation of the objective variable with respect to the variation of the first explanatory variable corresponding to each of the plurality of learned models.
[0271] According to the data processing apparatus described in Item 18, the user can select a learned model (that is, an appropriate data processing condition) that seems appropriate in considering the relationship between the first explanatory variable and the objective variable by comparing the plurality of generated data.
[0272] (Item 19) In the data processing apparatus according to Item 18, when one data is selected by the user from the plurality of data, the display data generation unit stores the learned model corresponding to the selected data in the database as an appropriate learned model. When a feature amount is extracted from a data group similar to the above one data group, the teacher data generation unit presents the data processing condition associated with the appropriate learned model to the user.
[0273] According to the data processing device described in Section 19, in the learning phase, features can be extracted from sample analysis data and a sample list can be generated using data processing conditions deemed appropriate for considering the relationship between the first explanatory variable and the target variable. A trained model is then generated using the training data generated based on this sample list. Therefore, in the inference phase, the relationship between the first explanatory variable given to the trained model and the target variable predicted by the trained model will follow the user's appropriate criteria. Consequently, the usefulness of the inference results can be improved.
[0274] (Clause 20) An inference method according to one embodiment predicts a target variable from multiple explanatory variables using a trained model. The inference method includes the steps of: predicting the target variable when a first explanatory variable selected from the multiple explanatory variables is set as a variable value, while a second explanatory variable other than the first explanatory variable is set as a fixed value, and using a trained model, the first explanatory variable is continuously varied within a predetermined range of variation; generating data showing the variation of the target variable in response to the variation of the first explanatory variable; and displaying the data generated in the generation step.
[0275] According to the inference method described in Section 20, users can easily predict how the dependent variable will change when the first explanatory variable, which represents the variable value, is continuously varied based on the displayed data. Therefore, the usefulness of the inference results can be enhanced.
[0276] The embodiments disclosed herein should be considered in all respects to be illustrative and not restrictive. The scope of the present invention is indicated by the claims rather than by the foregoing description, and all modifications within the meaning and scope equivalent to the claims are intended to be included. [Explanation of Symbols]
[0277] 1 Data processing unit, 2,64 Display unit, 3,63 Operation unit, 4 Analysis device, 5 Main unit, 6 Information processing unit, 11,61 ROM, 12,62 RAM, 13,66 Communication I / F, 15 Database, 20 Analysis data acquisition unit, 22 Feature extraction unit, 24 Physical property data acquisition unit, 26 Sample information acquisition unit, 28 Training data generation unit, 30 Learning unit, 32 Inference unit, 34 Display data still image unit, 67 Data acquisition unit, 69 Information acquisition unit, 76,76A~76C Display area, 82 Icons, 90,92,94,96,98,99 Graphs, 100 Analysis system.
Claims
1. An inference unit that uses a pre-trained model to predict the target variable from multiple explanatory variables, The system includes a display data generation unit that generates data for displaying the inference results from the inference unit, The inference unit, A first explanatory variable selected from the aforementioned plurality of explanatory variables is set as a variable value, while a second explanatory variable other than the first explanatory variable is set as a fixed value, and Using the trained model, predict the target variable when the first explanatory variable is continuously varied within a predetermined range of variation. The display data generation unit is configured to generate data showing the change in the target variable in response to the change in the first explanatory variable. The inference unit, Based on the absolute values of the principal component loadings of each explanatory variable for a specific principal component obtained by principal component analysis of the aforementioned plurality of explanatory variables, at least one of the aforementioned first explanatory variables is selected from the plurality of explanatory variables, and, A data processing device that predicts the target variable when the first explanatory variable is continuously varied within the range of variation for each of the selected at least one first explanatory variable, using the trained model.
2. The data processing apparatus according to claim 1, wherein the display data generation unit generates a two-dimensional graph with the first explanatory variable as the first axis and the objective variable as the second axis.
3. The inference unit, Select two or more of the above-mentioned first explanatory variables from the above-mentioned plurality of explanatory variables, and For each of the two or more selected first explanatory variables, the trained model is used to predict the target variable when the first explanatory variable is continuously changed within the range of variation. A display unit is connected to the aforementioned data processing device. The aforementioned display data generation unit, Two or more two-dimensional graphs are generated corresponding to each of the two or more first explanatory variables. The data processing apparatus according to claim 2, wherein the two or more generated two-dimensional graphs are displayed on the display unit such that they are superimposed on each other.
4. The display data generation unit is configured to provide a first user interface for selecting the first explanatory variable and setting the range of variation. The data processing device according to claim 1, wherein the first user interface includes information regarding a recommended range for the variation range.
5. The aforementioned trained model is a model generated by machine learning using training data that takes the multiple explanatory variables as input and the objective variable as the correct output. The data processing device according to claim 4, wherein the recommended range is set based on the value of the first explanatory variable included in the training data.
6. The data processing apparatus according to claim 4 or 5, wherein the display data generation unit further provides a second user interface for setting the value of the second explanatory variable.
7. The aforementioned trained model is a model generated by machine learning using training data that takes the multiple explanatory variables as input and the objective variable as the correct output. The data processing device according to claim 1, wherein the range of variation is set based on the value of the first explanatory variable included in the training data.
8. A display unit is connected to the aforementioned data processing device. The aforementioned display data generation unit, At least one of the aforementioned data is generated corresponding to each of the at least one first explanatory variable. The data processing apparatus according to claim 1 or 7, wherein the generated at least one data is displayed on the display unit in order from the highest principal component loading of the corresponding first explanatory variable.
9. The inference unit, Select two or more of the first explanatory variables from the aforementioned plurality of explanatory variables, For each of the two or more selected first explanatory variables, the trained model is used to predict the target variable when the first explanatory variable is continuously changed within the range of variation. The display data generation unit generates two or more data corresponding to the two or more first explanatory variables, The data processing device according to claim 1, further comprising a database for storing, in association with information about the project to which the trained model is applied, the type of the first explanatory variable that has the greatest influence on the target variable among the two or more first explanatory variables.
10. A training data generation unit that takes the aforementioned multiple explanatory variables as input and generates training data with the aforementioned target variable as the correct output, The system further comprises a learning unit that generates the trained model by machine learning using the aforementioned training data, The data processing device according to claim 9, wherein the training data generation unit presents to the user information about the project and the type of the first explanatory variable associated with the project.
11. An inference unit that predicts a target variable from multiple explanatory variables using a trained model, The system includes a display data generation unit that generates data for displaying the inference results from the inference unit, The inference unit, A first explanatory variable selected from the aforementioned plurality of explanatory variables is set as a variable value, while a second explanatory variable other than the first explanatory variable is set as a fixed value, and Using the trained model, predict the target variable when the first explanatory variable is continuously varied within a predetermined range of variation. The display data generation unit is configured to generate data showing the change in the target variable in response to the change in the first explanatory variable. A training data generation unit that takes the aforementioned multiple explanatory variables as input and generates training data with the aforementioned target variable as the correct output, A learning unit that generates the trained model by machine learning using the aforementioned training data, The system further comprises a database for storing the trained model in association with the training data, The training data generation unit generates multiple training data sets, each containing multiple feature quantities extracted from a single data set using multiple different data processing conditions. The aforementioned learning unit, The aforementioned machine learning generates multiple pre-trained models from the multiple training data sets, A data processing device that stores each of the generated trained models in the database, linked to the corresponding data processing conditions.
12. The inference unit predicts the target variable when the first explanatory variable is continuously varied within the range of variation, using each of the plurality of trained models. The data processing apparatus according to claim 11, wherein the display data generation unit generates a plurality of data showing the variation of the target variable in response to the variation of the first explanatory variable, corresponding to each of the plurality of trained models.
13. When one of the aforementioned data is selected by the user from among the multiple data, the display data generation unit stores the trained model corresponding to the selected data in the database as an appropriate trained model. The data processing apparatus according to claim 12, wherein the training data generation unit presents the user with the data processing conditions associated with the appropriate trained model when features are extracted from a data set similar to the one data set.
14. An inference unit that predicts a target variable from multiple explanatory variables using a trained model, The system includes a display data generation unit that generates data for displaying the inference results from the inference unit, The inference unit, A first explanatory variable selected from the aforementioned plurality of explanatory variables is set as a variable value, while a second explanatory variable other than the first explanatory variable is set as a fixed value, and Using the trained model, predict the target variable when the first explanatory variable is continuously varied within a predetermined range of variation. The display data generation unit is configured to generate data showing the change in the target variable in response to the change in the first explanatory variable. The aforementioned trained model is a model generated by machine learning using training data that takes the multiple explanatory variables as input and the objective variable as the correct output. The inference unit, By applying the decision tree algorithm to the multiple explanatory variables input to the aforementioned trained model, the importance of each explanatory variable is calculated. Based on the calculated importance of each explanatory variable, the explanatory variable with the highest importance is preferentially selected as the first explanatory variable, and, A data processing device that predicts the target variable when the selected first explanatory variable is continuously varied within the range of variation using the trained model.
15. The data processing device according to claim 14, wherein the range of variation is set based on the value of the first explanatory variable included in the training data.
16. A display unit is connected to the aforementioned data processing device. The aforementioned display data generation unit, At least one of the aforementioned data is generated corresponding to each of the at least one first explanatory variable, and The data processing apparatus according to claim 14, wherein the generated at least one data is displayed on the display unit in order from the highest importance of the corresponding first explanatory variable.
17. A computer-based inference method that uses a pre-trained model to predict a target variable from multiple explanatory variables, The steps include: using a first explanatory variable selected from the plurality of explanatory variables as a variable value, while setting a second explanatory variable other than the first explanatory variable as a fixed value, and predicting the target variable when the first explanatory variable is continuously varied within a predetermined range of variation using the trained model stored in the computer's memory; A step of generating data showing the variation of the dependent variable in response to the variation of the first explanatory variable, The process includes a step of displaying the data generated by the above-mentioned generation step, The aforementioned prediction step is, The steps include selecting at least one of the multiple explanatory variables from among the multiple explanatory variables based on the absolute values of the principal component loadings of each explanatory variable for a specific principal component obtained by principal component analysis of the multiple explanatory variables, An inference method comprising the step of predicting the target variable when the first explanatory variable is continuously varied within the range of variation for each of the selected at least one first explanatory variable, using the trained model.