Estimation device, estimation method, and program

The estimation device and method address inefficiencies in predictive model performance estimation by using explanatory variable data within the prediction model, enabling real-time predictive performance assessment without dependent variables.

JP7790562B2Active Publication Date: 2025-12-23NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024521508
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-12-23
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

Existing prediction models face challenges in calculating performance until dependent data is obtained, as the performance index for the predictive model cannot be determined, as the value of the performance index for the predictive model cannot be calculated until the dependent variable is obtained, which results in inefficiencies in performance estimation.

Method used

An estimation device and method that estimates predictive performance without using data on dependent variables by selecting and comparing explanatory variable data within a prediction model, utilizing data selection and performance estimation means to determine predictive performance.

Benefits of technology

Enables predictive performance estimation during operation without waiting for dependent variable data, allowing for real-time performance assessment of predictive models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007790562000001
    Figure 0007790562000001
  • Figure 0007790562000002
    Figure 0007790562000002
  • Figure 0007790562000003
    Figure 0007790562000003
Patent Text Reader

Abstract

An estimation device 100 according to the present invention comprises: a data selection means 121 for selecting, from first explanatory variable data prepared according to a preset prediction model and first objective variable data corresponding to the first explanatory variable data, first explanatory variable data and first objective variable data corresponding to the first explanatory variable data, on the basis of second explanatory variable data that is not associated with an objective variable; and a performance estimation means 122 for estimating the performance of a prediction model's prediction about the second explanatory variable data, on the basis of the comparison between prediction data obtained by inputting the first explanatory variable data selected according to the prediction model and the first objective variable data corresponding to the first explanatory variable data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an estimation device, an estimation method, and a program. [Background technology]

[0002] A method has been studied for creating a prediction model that predicts dependent variable data from explanatory variable data based on explanatory variable data and dependent variable data corresponding to the explanatory variable data, and then putting the created prediction model into practical use. For example, Patent Document 1 discloses a technology for creating multiple prediction model candidates from explanatory variable data and new data generated from the explanatory variable data, and selecting a prediction model by evaluating the prediction accuracy of the multiple prediction model candidates. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2021 / 229648 Summary of the Invention [Problem to be solved by the invention]

[0004] However, depending on the operating conditions of a prediction model, it may take time to obtain data on the dependent variable that is the target of prediction by the prediction model. In such cases, a problem occurs in that the value of the performance index for the prediction of the prediction model cannot be calculated until the dependent variable is obtained.

[0005] In view of the above-mentioned problems, an object of the present disclosure is to provide an estimation device, an estimation method, and a program that can estimate the predictive performance of a predictive model without using data on dependent variables. [Means for solving the problem]

[0006] An estimation device according to one aspect of the present invention includes: a data selection means for selecting, from first explanatory variable data prepared according to a preset prediction model and first dependent variable data corresponding to the first explanatory variable data, the first explanatory variable data and the first dependent variable data corresponding to the first explanatory variable data, based on second explanatory variable data to which no dependent variable is associated; a performance estimation means for estimating the prediction performance of the prediction model for the second explanatory variable data based on a comparison between prediction data obtained by inputting the first explanatory variable data selected for the prediction model and the first dependent variable data corresponding to the first explanatory variable data; having The structure is as follows.

[0007] Furthermore, an estimation method according to one aspect of the present invention includes: selecting, from first explanatory variable data prepared according to a preset prediction model and first dependent variable data corresponding to the first explanatory variable data, the first explanatory variable data and the first dependent variable data corresponding to the first explanatory variable data, based on second explanatory variable data to which no dependent variable is associated; estimating the prediction performance of the prediction model for the second explanatory variable data based on a comparison between prediction data obtained by inputting the selected first explanatory variable data into the prediction model and the first dependent variable data corresponding to the selected first explanatory variable data; The structure is as follows.

[0008] Furthermore, a program according to one aspect of the present invention includes: selecting, from first explanatory variable data prepared according to a preset prediction model and first dependent variable data corresponding to the first explanatory variable data, the first explanatory variable data and the first dependent variable data corresponding to the first explanatory variable data, based on second explanatory variable data to which no dependent variable is associated; estimating the prediction performance of the prediction model for the second explanatory variable data based on a comparison between prediction data obtained by inputting the selected first explanatory variable data into the prediction model and the first dependent variable data corresponding to the selected first explanatory variable data; Have the computer perform the process, The structure is as follows. [Effects of the Invention]

[0009] According to the present disclosure, it is possible to estimate the predictive performance of a predictive model during operation without using data on dependent variables. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a block diagram illustrating an example of a hardware configuration of an estimation device according to a first embodiment of the present disclosure. [Figure 2] 1 is a block diagram illustrating an example of a configuration of an estimation device according to a first embodiment of the present disclosure. [Figure 3] 4 is a flowchart illustrating an example of an operation of the estimation device according to the first embodiment of the present disclosure. [Figure 4] FIG. 10 is a block diagram illustrating an example of the configuration of an estimation device according to a second embodiment of the present disclosure. [Figure 5] FIG. 10 is a block diagram illustrating an example of the operation of an estimation device according to a second embodiment of the present disclosure. [Figure 6] FIG. 10 is a diagram illustrating an example of an estimation process performed by an estimation device according to a second embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] Preferred embodiments of the present disclosure will be described in detail with reference to the drawings.

[0012] First Embodiment First, an example of a first embodiment of the present disclosure will be described. Fig. 1 is a block diagram showing an example of the hardware configuration of an estimation device according to this embodiment, and Fig. 2 is a block diagram showing an example of the configuration of the estimation device. Fig. 3 is a flowchart showing an example of the operation of the estimation device. Note that this embodiment shows an outline of the configuration of an estimation device and an estimation method described in a second embodiment, which will be described later.

[0013] First, the hardware configuration of an estimating device 100 according to this embodiment will be described with reference to Fig. 1. The estimating device 100 is configured as a general information processing device, and is equipped with the following hardware configuration, for example. ·CPU(Central Processing Unit)101(Arithmetic unit) ROM (Read Only Memory) 102 (storage device) RAM (Random Access Memory) 103 (storage device) Programs 104 loaded into RAM 103 A storage device 105 for storing a group of programs 104 A drive device 106 that reads and writes from a storage medium 110 external to the information processing device A communication interface 107 that connects to a communication network 111 outside the information processing device Input / output interface 108 for inputting and outputting data Bus 109 connecting each component

[0014] 1 shows an example of the hardware configuration of the information processing device that is the estimation device 100, and the hardware configuration of the information processing device is not limited to the above-described case. For example, the information processing device may be configured with a part of the above-described configuration, such as not including the drive device 106. Furthermore, the information processing device may use a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an MPU (Micro Processing Unit), an FPU (Floating point number Processing Unit), a PPU (Physics Processing Unit), a TPU (Tensor Processing Unit), a quantum processor, a microcontroller, or a combination thereof, instead of the above-described CPU.

[0015] The estimation device 100 can be equipped with the data selection means 121 and performance estimation means 122 shown in Fig. 2 by having the CPU 101 acquire and execute the program group 104. The program group 104 is stored in advance in, for example, the storage device 105 or the ROM 102, and is loaded into the RAM 103 and executed by the CPU 101 as needed. The program group 104 may be supplied to the CPU 101 via the communication network 111, or may be stored in advance in the storage medium 110, and the drive device 106 may read out the programs and supply them to the CPU 101. However, the data selection means 121 and performance estimation means 122 described above may be constructed using electronic circuits dedicated to realizing such means.

[0016] First, the estimation device 100 has a function of acquiring a prediction model, first explanatory variable data, first objective variable data corresponding to the first explanatory variable data, and second explanatory variable data. The prediction model, when executed within the estimation device 100, has a function of outputting a prediction result for the corresponding objective variable data when explanatory variable data is input. In this case, the output data is referred to as prediction data. The first explanatory variable data and first objective variable data are data prepared according to the prediction model, and include, for example, first explanatory variable data and first objective variable data as training data for generating the prediction model by machine learning, and first explanatory variable data and first objective variable data for verifying the prediction model. Furthermore, the second explanatory variable data is explanatory variable data acquired when the prediction model is used in another task, and there is no objective variable data corresponding to the second explanatory variable data, and the second explanatory variable data is not associated with the second explanatory variable data.

[0017] The estimation device 100 may acquire the prediction model and each piece of data from outside the estimation device 100. For example, if the estimation device 100 is connected to other devices via a network, the prediction model and each piece of data may be acquired from another device that exists on the same network as the estimation device 100. In this case, the estimation device 100 may include a communication device for communicating with other devices on the network. Furthermore, if the estimation device 100 includes a storage device (not shown), the prediction model and each piece of data may be acquired from the storage device.

[0018] The data selection means 121 selects, based on the second explanatory variable data, some or all of the first explanatory variable data and the corresponding first objective variable data from the first explanatory variable data and the first objective variable data corresponding to the first explanatory variable data acquired by the estimation device 100. At this time, the data selection means 121 selects, based on the content of the second explanatory variable data, all or part of the first explanatory variable data that is more suitable for estimating the prediction performance of the prediction model for the second explanatory variable data. More specifically, the data selection means 121 selects all or part of the first explanatory variable data based on the content of the second explanatory variable data so that the estimated value of the prediction performance index for the second explanatory variable data of the prediction model by the performance estimation means 122 (described later) is closer to the actual value. For example, in estimating the prediction performance of a prediction model, it is generally desirable to use data of explanatory variables that the prediction model did not use during training. Therefore, the data selection means 121 may select all or part of the first explanatory variable data by selecting data of explanatory variables that were not used during training of the prediction model. As another example, when estimating the predictive performance of a prediction model for data on a second explanatory variable, it is preferable to use data on an explanatory variable corresponding to data on an explanatory variable that is close to the data on the second explanatory variable in the explanatory variable space. Therefore, as a method for selecting all or part of the data on the first explanatory variable by the data selection means 121, a selection method may be used in which the distribution of the selected explanatory variable data is close to the spatial distribution of the explanatory variables (hereinafter simply referred to as "distribution") of the data on the second explanatory variable. Note that the method for selecting the data on the reference explanatory variable may be more appropriately determined depending on the performance estimation method used by the performance estimation means 122, which will be described later. The data on the explanatory variables selected by the data selection means 121 is referred to as data on the reference explanatory variable. The data selection means 121 also selects data on the objective variable corresponding to the selected reference explanatory variable data, which is referred to as data on the reference objective variable. The reference explanatory variable data and the data on the objective variable corresponding to it are collectively referred to as reference data.

[0019] The performance estimation means 122 estimates the prediction performance of the prediction model for the second explanatory variable data based on a comparison between the prediction data obtained by inputting the reference explanatory variable data into the prediction model and the reference objective variable data corresponding to the reference explanatory variable data. More specifically, for example, the performance estimation means 122 may use the value of a performance index calculated using the prediction data obtained by inputting the reference explanatory variable data into the prediction model and the reference objective variable data as the estimated value of the performance index for the second explanatory variable data of the prediction model. In this case, the performance estimation means 122 may further consider the difference in trend between the second explanatory variable data and the reference explanatory variable data, and adjust the estimated value of the performance index so that the larger the difference in trend, the worse the performance index. Here, the difference in trend is information based on the results of comparing the contents of two data, and is, for example, a numerical non-negative value that represents the difference in content between the two data, i.e., the degree to which the contents of the two data differ. As an example of the difference in trend, a distance or index defined between the distributions of two data may be used.

[0020] The performance index used by the performance estimation means 122 quantitatively evaluates the quality or poorness of the prediction performance of the prediction model. The actual value of the performance index is calculated based on a comparison of the predicted data obtained by inputting data of the second explanatory variable into the prediction model with the data of the dependent variable corresponding to the data of the second explanatory variable. Therefore, if data of the dependent variable is not available, the actual value of the performance index cannot be calculated. Specific examples of performance indexes include mean absolute error, mean squared error, and coefficient of determination in the case of regression, and accuracy, F1 score, and cross entropy in the case of discrimination. Of course, the performance indexes that can be used by the performance estimation means 122 are not limited to these, and any evaluation index can be used as the estimation target.

[0021] The estimation device 100 having the above-described configuration executes the estimation method shown in the flowchart of FIG. 3 by using the functions of the data selection means 121 and performance estimation means 122 described above.

[0022] As shown in FIG. 3, the estimation device 100 From first explanatory variable data and first dependent variable data corresponding to the first explanatory variable data prepared according to a preset prediction model, the first explanatory variable data and the first dependent variable data corresponding to the first explanatory variable data are selected based on second explanatory variable data to which no dependent variable is associated (step S101); The prediction performance of the prediction model for the second explanatory variable data is estimated based on a comparison between the prediction data obtained by inputting the first explanatory variable data selected for the prediction model and the first dependent variable data corresponding to the first explanatory variable data (step S102).

[0023] As described above, in the present disclosure, for example, by using a prediction model currently in operation, explanatory variable data used in creating the prediction model as data for the first explanatory variable (i.e., explanatory variable data used in training the prediction model and explanatory variable data used in validating the prediction model), objective variable data used in creating the prediction model as data for the objective variable corresponding to the data for the first explanatory variable (i.e., objective variable data used in training the prediction model and objective variable data used in validating the prediction model), and explanatory variable data used in running the prediction model as data for the second explanatory variable, it is possible to estimate the predictive performance for the explanatory variable data used in running the prediction model. In other words, when running the prediction model, the estimation device 100 can estimate the predictive performance of the prediction model during operation using the explanatory variable data during operation without waiting for the objective variable data during operation to be obtained.

[0024] The configuration shown in FIG. 2 is merely an example and does not necessarily limit the configuration of the estimation device according to this embodiment. For example, some functions of the estimation device 100 may be realized by multiple devices working together. As a specific example, the estimation device 100 may be configured with a first estimation device that acquires data on a first explanatory variable, data on a target variable corresponding to the data on the first explanatory variable, and data on a second explanatory variable, and a second estimation device that acquires a prediction model. Furthermore, in order to estimate the values ​​of performance indicators from multiple angles, the estimation device 100 may include multiple data selection means and multiple performance estimation means. Furthermore, the estimation device 100 may include a device that calculates a final estimate by the estimation device 100 based on estimates by the multiple performance estimation means. If the estimation device 100 also has functions other than estimating the values ​​of performance indicators of a prediction model, it may include devices not shown in FIG. 2.

[0025] <Second embodiment> Next, an example of the second embodiment of the present disclosure will be described with reference to Fig. 4 to Fig. 6. Fig. 4 is a block diagram for explaining an example of the configuration of an estimation device according to this embodiment, and Fig. 5 is a flowchart for explaining an example of the operation of the estimation device. Fig. 6 is a diagram for explaining the state when the estimation device is actually operated.

[0026] In this embodiment, the data of the explanatory variables used to create the prediction model is used as the data of the first explanatory variables. Furthermore, the data of the first objective variable corresponding to the data of the first explanatory variables is used as the data of the objective variable corresponding to the data of the explanatory variables used to create the prediction model. Furthermore, the data of the explanatory variables at the time of operation of the prediction model is used as the data of the second explanatory variables. However, these are merely examples of the present embodiment as an example of an embodiment of the present disclosure and do not limit the embodiment of the present disclosure itself. As an example that is not limited to each data of the present disclosure, for example, a prediction model created using data of explanatory variables and data of objective variables obtained in a certain task may be applied to another task to estimate the predictive performance of the prediction model for the data of explanatory variables of the other task. In this case, the data of the explanatory variables of the other task can be used as the data of the second explanatory variables, rather than the data of explanatory variables at the time of operation of the prediction model, and the estimation device according to the present disclosure can be applied. Alternatively, the data of the first explanatory variables and the data of the first objective variable corresponding to the data of the first explanatory variables may be data other than that used to create the prediction model. For example, the data of the explanatory variables at the time of operation of a prediction model for which data of the corresponding objective variable has already been obtained and the data of the corresponding objective variable may be used.

[0027] The estimation device 1 in this embodiment is configured with one or more information processing devices each including a calculation device and a storage device. As shown in FIG. 4 , the estimation device 1 includes a second data acquisition unit 10, an output unit 20, a control unit 30, and a storage unit 40. The control unit 30 further includes a reference data selection unit 3 and a performance estimation unit 4. The functions of the second data acquisition unit 10, the output unit 20, and the reference data selection unit 3 and the performance estimation unit 4 included in the control unit 30 can be realized by the calculation device executing a program for realizing each function stored in the storage device. The estimation device 1 also includes a storage unit 40 in the storage device that stores first data 41, a prediction model 42, and parameter information 43. The control unit 30 communicates data with the second data acquisition unit 10, the output unit 20, and the storage unit 40 via a communication network or the like.

[0028] The estimation device 1 has the function of estimating the prediction performance of the prediction model 42 stored in the storage unit 40 with respect to the data of the second explanatory variables acquired by the second data acquisition unit 10, i.e., the data of the explanatory variables during operation of the prediction model 42, by using the above-mentioned components. The components will be described in detail below.

[0029] The second data acquisition unit 10 acquires explanatory variable data during operation of the prediction model as data on the second explanatory variables. The explanatory variable data during operation (i.e., data on the second explanatory variables) is data obtained during operation, and the prediction model predicts data on the corresponding objective variable. For example, if a user of the estimation device 1 inputs explanatory variable data during operation, the second data acquisition unit 10 may include an interface for accepting user input, specifically, a touch panel, buttons, a voice input device, or the like. Furthermore, if the estimation device 1 itself has the function of collecting data on the second explanatory variables, examples of the second data acquisition unit 10 include a sensor device such as a camera, a device for processing information acquired from the sensor device, and a device for storing the processed information as data on the second explanatory variables. If the second explanatory variable data is acquired from a device other than the estimation device 1, a communication device with the other device, or the like, may be included.

[0030] The output unit 20 outputs the estimated value of the performance index calculated by the control unit 30. When the estimated value is output to the user of the estimation device 1, the output unit 20 corresponds to a display for displaying the estimated value, a speaker for outputting sound, or the like. When the estimated value is used in a device other than the estimation device 1, the output unit 20 corresponds to a device for communicating with the other device, or the like. The estimation device 1 may also have an alert function for notifying the user of a deterioration in the prediction performance when the estimated value falls below a certain value, for example. In this case, the output unit 20 corresponds to a device for notifying the user of the deterioration in the prediction performance.

[0031] The following describes the configuration of the storage unit 40. As shown in Fig. 4, the storage unit 40 stores in advance first data 41, a prediction model 42, and parameter information 43. The storage unit 40 transmits the first data 41, the prediction model 42, and the parameter information 43 to the control unit 30 as necessary.

[0032] The first data 41 includes data on explanatory variables used in creating the prediction model 42 (first explanatory variable data) and data on objective variables corresponding to the explanatory variable data used in creating the prediction model 42 (first objective variable data). The first data 41 includes training data used in training the prediction model 42 (i.e., data on explanatory variables used in training the prediction model 42 and data on objective variables corresponding to the explanatory variable data), and verification data used in verifying the generalization performance of the prediction model 42 (i.e., data on explanatory variables used in verifying the prediction model 42 and data on objective variables corresponding to the explanatory variable data).

[0033] The prediction model 42 is a prediction model created using the first data 41. Specifically, the prediction model 42 may be created by so-called supervised machine learning so that, when data of explanatory variables in training data among the first data 41 is input, the prediction model 42 predicts data of the objective variable in the training data. Furthermore, the generalization performance of the prediction model 42 may be verified using validation data among the first data 41. Examples of algorithms that realize the prediction model 42 include linear regression, decision trees, random forests, and neural networks. These are merely examples, and the algorithm that realizes the prediction model 42 is not limited as long as it is possible to predict data of the objective variable based on data of the explanatory variables.

[0034] The parameter information 43 is information on parameters necessary for estimating the value of the performance index, and may include, for example, information on the performance index to be estimated, a method for acquiring reference data (described later), and a method for estimating the performance index.

[0035] In this embodiment, the prediction model, the data of the first explanatory variables, and the data of the first objective variable corresponding to the data of the first explanatory variables do not change for each estimation, and therefore the processing required for acquisition can be reduced by storing them in the storage unit 40. On the other hand, the data of the second explanatory variables, i.e., the data of the explanatory variables used when operating the prediction model, changes for each estimation, and therefore by acquiring the data using the second data acquisition unit 10, it becomes possible to estimate the value of the performance index for the data of the explanatory variables during the most recent operation.

[0036] Next, the configuration of the control unit 30 will be described. As shown in Fig. 4, the control unit 30 has a reference data selection unit 3 and a performance estimation unit 4. The control unit 30 also has a CPU, ROM, RAM, etc. (not shown), and performs various controls and calculations on the reference data selection unit 3 and the performance estimation unit 4.

[0037] The reference data selection unit 3 (data selection means) selects all or part of the explanatory variable data of the first data 41 based on the second explanatory variable data acquired by the second data acquisition unit 10, the prediction model 42 acquired from the storage unit 40, the first data 41, and the parameter information 43. At this time, the reference data selection unit 3 also selects the objective variable data corresponding to the explanatory variable data of the selected first data 41. Here, the selected explanatory variable data is referred to as the reference explanatory variable data. Furthermore, the objective variable data of the first data 41 corresponding to the reference explanatory variable data is referred to as the reference objective variable data. The reference explanatory variable data and the reference objective variable data are collectively referred to as the reference data.

[0038] As a specific example of a method for selecting reference data by the reference data selection unit 3, a method may be used in which only the validation data of the first data 41 is used as the reference data. The performance index estimation formula (1) described below uses the value of the performance index of the prediction model 42 for the reference data as a basis for the estimated value of the performance index of the prediction model 42 for the data of the second explanatory variable in the calculation. By not including the training data of the first data 41 and using only the validation data in the reference data, it is possible to avoid overestimation of the value of the performance index due to overfitting of the training data of the prediction model 42, and more appropriately calculate the estimated value of the performance index of the prediction model 42. Of course, all or part of the training data may be included in the reference data, and increasing the number of samples included in the reference data, including the training data, may be useful for appropriately estimating the value of the performance index. For example, when it is not possible to select data of the first explanatory variable that is related to the data of the second explanatory variable, as described below, because the attributes or characteristics of the data cannot be classified based on the content of the data, the reference data selection unit 3 may select only the validation data or all data including the training data.

[0039] As another example of a method for selecting reference data by the reference data selector 3, all or part of the first data 41 may be selected as reference data, which is highly relevant to the prediction performance of the prediction model 42 for the data of the second explanatory variable. Here, "highly relevant" refers to a positive correlation between the prediction performance of the prediction model 42 for the data of the second explanatory variable and the prediction performance of the prediction model 42 for the data of the reference explanatory variable, such that the two prediction performances are considered to be comparable. In other words, in this case, a comparison of the data of the second explanatory variable with the data of the first explanatory variable in the first data 41 indicates that the data content is highly relevant according to a predetermined standard for the second explanatory variable data. The performance index estimation formula (1) described below uses the value of the performance index of the prediction model 42 for the reference data as a standard for the calculation of the estimated value of the performance index of the prediction model 42 for the data of the second explanatory variable. Therefore, by using a portion of the data highly relevant to the data of the second explanatory variable as reference data, the estimated value of the performance index of the prediction model 42 can be more appropriately calculated. When selecting data highly relevant to the data of the second explanatory variable from the first data 41, empirical knowledge about the data (generally referred to as domain knowledge) or the characteristics of the prediction model 42 may be used. A specific method for acquiring highly relevant data is, for example, when the data of the explanatory variable includes date and time information, to acquire data of the same month and day, the same day of the week, or the same season as any sample included in the data of the second explanatory variable from the first data 41. As another example, for an explanatory variable that the prediction model 42 particularly places importance on during prediction, data whose values ​​are close to those of the data of the second explanatory variable may be acquired from the first data 41.

[0040] As another example of a method for selecting reference data by the reference data selector 3, a portion of the first data 41 may be selected as reference data, which has a smaller difference in trend from the data of the second explanatory variable, i.e., a smaller difference in data content based on a preset standard. The performance index estimation formula (1) described below calculates an estimated value such that the larger the difference in trend between the data of the second explanatory variable and the data of the reference explanatory variable, the worse the value of the performance index of the prediction model 42 for the data of the second explanatory variable. The data of the second explanatory variable often contains only a limited amount of data and a limited range in the explanatory variable space compared to the first data 41. Under such circumstances, if all the explanatory variable data of the first data 41 were used as the data of the reference explanatory variable, the difference in trend between the data of the second explanatory variable and the data of the reference explanatory variable would be large, resulting in a poorly estimated value of the performance index of the prediction model 42 for the data of the second explanatory variable. However, there are cases in which the explanatory variable data of the first data 41 contains samples that are close to each sample of the second explanatory variable in the explanatory variable space. In other words, the prediction model 42 may have already been trained or validated using data (first data 41) used to create the prediction model 42 that is close to the explanatory variable data (second explanatory variable data) during operation. In such cases, the actual value of the performance index for prediction of the prediction model 42 for the second explanatory variable data will be better. Therefore, by selecting some of the explanatory variable data in the first data 41 that has a small difference in trend from the second explanatory variable data and using this as the reference explanatory variable data, it is possible to estimate a value close to the actual value of the performance index for prediction of the explanatory variable data of the prediction model 42. Furthermore, the trend difference index may be calculated using the same index as the trend difference index used in the performance index estimation formula (1) described below, or may be calculated using a different index. Furthermore, it may be calculated using a combination of multiple different indexes.

[0041] Here, as an example of a specific method for the reference data selection unit 3 to select data with a small difference in trend from the explanatory variable data of the first data 41, one or more samples closest to the second explanatory variable data are obtained from the explanatory variable data of the first data 41 and used as temporary reference explanatory variable data. Next, from each sample of the explanatory variable data of the first data 41 that is not included in the temporary reference explanatory variable data, one or more samples that, when added to the temporary reference explanatory variable data, will better reduce the difference in trend between the temporary reference explanatory variable data and the second explanatory variable data are obtained and added to the temporary reference explanatory variable data. The above addition process is repeated, for example, until the number of samples included in the temporary reference explanatory variable data exceeds a certain number, and finally, the temporary reference explanatory variable data is used as the reference explanatory variable data, thereby enabling data with a small difference in trend from the explanatory variable data of the first data 41 to be selected.

[0042] The reference data selection unit 3 may use a selection method that combines two or more examples of the above-described reference data selection methods. For example, the reference data selection unit 3 may first select validation data from the first data 41, then select a portion of data that is highly relevant to the data of the second explanatory variable, and further select a portion of data that has a small difference in trend from the data of the second explanatory variable. Furthermore, the reference data selection method may be a more suitable selection method that takes into account the specific estimation method used by the performance estimation unit 4, which will be described later. In other words, if it is empirically or theoretically considered that the estimation method used by the performance estimation unit 4 will enable more accurate calculation of performance estimates using a specific reference data selection method, the above-described selection method may be used.

[0043] Note that if the difference in trend between the explanatory variable data of the first data 41 and the second explanatory variable data is large, and any sample of the second explanatory variable data has low relevance to samples included in the explanatory variable data of the first data 41, the difference in trend with the second explanatory variable data will never be zero, no matter how all or part of the explanatory variable data of the first data 41 is selected. Therefore, if the trend between the explanatory variable data of the first data 41 and the second explanatory variable data is large, even if all or part of the explanatory variable data of the first data 41 is selected as described above, it is possible to estimate a deterioration in the prediction performance of the prediction model 42 for the second explanatory variable data.

[0044] The performance estimation unit 4 (performance estimation means) calculates an estimated value P of the performance index of the prediction model 42 for the data of the second explanatory variable, for example, using equation (1) described later, based on the reference data, the data of the second explanatory variable, and the prediction model 42 and parameter information 43 acquired from the storage unit 40. The value of the performance index estimated by the calculation is output to the output unit 20. <Equation for estimating performance index> P = B + D × A…(1) however, P: Estimated value of the performance index of the prediction model 42 for the data of the second explanatory variable, B: The performance index values ​​of the predictive model for the standard explanatory variable data and the standard target variable data. D: A non-negative value indicating the difference in trend between the data of the second explanatory variable and the data of the base explanatory variable. A: The rate of change in the value of the performance index according to D, calculated based on the data of the standard explanatory variables and a comparison between the predicted data obtained by inputting the data of the standard explanatory variables into the prediction model 42 and the data of the standard objective variables, is.

[0045] In equation (1), B is the value of the performance index of the predictive model for the reference explanatory variable data and the reference dependent variable data. In particular, when there is absolutely no difference in trend between the reference explanatory variable data and the second explanatory variable data (i.e., when D in equation (1) is 0), the estimated performance index P is the value of B. In fact, when there is absolutely no difference in trend between the reference explanatory variable data and the second explanatory variable data, the second explanatory variable data and the reference explanatory variable data can be considered identical. Therefore, unless there is a change in the relationship between the explanatory variable data and the dependent variable data (i.e., concept drift), the predictive performance of the predictive model 42 for the second explanatory variable data will be almost identical to the predictive performance of the predictive model 42 for the reference explanatory variable data. Equation (1) enables the estimation of the performance index value based on this fact. In other words, when concept drift does not occur and the reference explanatory variable data can be selected so that it has the same trend as the second explanatory variable data, equation (1) can accurately estimate the performance index value from the reference explanatory variable data and the reference dependent variable data.

[0046] Furthermore, when there is a difference in trend between the data for the reference explanatory variable and the data for the second explanatory variable (i.e., when the value of D in formula (1) is greater than 0), formula (1) estimates the deterioration in performance based on the values ​​of D and A in formula (1), with the value of B in formula (1) set as the best value. Generally, the predictive performance of a predictive model will deteriorate for explanatory variable data that has a trend different from the explanatory variable data of the training data for the predictive model, and formula (1) can estimate the deterioration in the predictive performance of the predictive model 42 for the data for the second explanatory variable based on this.

[0047] D in equation (1) is a nonnegative value that quantitatively indicates the difference in trend between the data for the second explanatory variable and the data for the reference explanatory variable. D is 0 when there is no difference in trend between the data for the second explanatory variable and the data for the reference explanatory variable, and a large value when the difference is large. As a specific example of D, the difference between the distribution of the data for the second explanatory variable and the distribution of the data for the reference explanatory variable can be calculated and used as D. Examples of indices for measuring the difference in data distribution include Kullback-Leibler divergence, Jensen-Shannon divergence, Wasserstein distance, and Maximum Mean Discrepancy (MMD). These indices are 0 when there is no difference in data distribution and take larger values ​​as the difference in data distribution increases. As another example, D can be calculated by calculating statistical quantities such as the mean and variance of the explanatory variables for the data for the second explanatory variable and the data for the reference explanatory variable, and then using the difference between these values ​​to calculate D. Another example is to calculate the proportion of samples in the data for the second explanatory variable that do not contain any samples of the reference explanatory variable within a certain distance from each sample, and use this value as D.

[0048] In the above example, the value of D may be calculated using all of the explanatory variables in the data for the second explanatory variable and the data for the reference explanatory variable, or only some of the explanatory variables. For example, only some of the explanatory variables that have a particularly strong influence on the prediction by the prediction model 42 may be used. Furthermore, in the above example, the value of D may be calculated using all of the data for the second explanatory variable and the data for the reference explanatory variable, or only some of the data. For example, when the data for the second explanatory variable consists of a very large number of samples, the data for the second explanatory variable may be sampled to calculate the value of D. This reduces the time required to calculate D in equation (1), enabling high-speed estimation.

[0049] Additionally, in calculating the value of D in equation (1), in order to adjust the influence of each explanatory variable on D, the values ​​of each explanatory variable in the data for the second explanatory variable and the data for the reference explanatory variable may be transformed before calculating D. For example, in order to make the influence of each explanatory variable on D the same, each explanatory variable may be transformed using min-max normalization or z-score normalization. As another example, in order to increase the influence of some explanatory variables that have a particularly strong influence on the prediction results of the prediction model 42 on D, the values ​​of the explanatory variables in the data for the second explanatory variable and the data for the reference explanatory variable may be doubled, for example.

[0050] A in formula (1) is the rate of change in the value of the performance index according to D, calculated based on the data of the standard explanatory variables and a comparison of the predicted data obtained by inputting the data of the standard explanatory variables into the prediction model 42 with the data of the standard objective variable. A is a negative value when the higher the value of the performance index, the better the performance of the prediction model (for example, when the coefficient of determination, accuracy, F1 score, etc. are used as performance indexes), and a positive value when the lower the value of the performance index, the better the performance of the prediction model (for example, when the mean absolute error, mean squared error, cross entropy, etc. are used as performance indexes). This makes it possible to formulate the deterioration in the predictive performance of the prediction model 42 for the data of the second explanatory variable as the value of D in formula (1) increases.

[0051] The value of A in formula (1) may be, for example, a constant. More specifically, for example, when data on explanatory variables and data on dependent variables corresponding to the data on explanatory variables are available in addition to the first data, the value of A may be set so that the estimated value of P according to formula (1) when the data on the explanatory variables is used as data on the second explanatory variable becomes the actual value of the performance index of the prediction model 42 for the data on the explanatory variables and the data on the dependent variable. Furthermore, the value of A may be suitably set based on, for example, theoretical analysis or empirical experimental results.

[0052] Furthermore, the value of A in formula (1) may be calculated based on the data of the reference explanatory variables and a comparison between the predicted data obtained by inputting the data of the reference explanatory variables into the prediction model 42 and the data of the reference objective variable. As an example of a specific calculation method, the calculation formula for the rate of deterioration of prediction performance based on distribution robust optimization shown in the following formula (2) may be used. This calculation formula for the rate of deterioration of prediction performance based on distribution robust optimization is a method of mathematically calculating the value of A from the value of the prediction performance index for the data of the reference explanatory variables of the prediction model 42 and the distribution of the data of the reference explanatory variables. Note that when the calculation formula according to formula (2) is used, the value of A is positive, and therefore, the prediction performance index used is one in which the lower the value of the index, the better the performance of the prediction model.

[0053] <Formula for predicting performance degradation rate based on distribution robust optimization> A=(Σij Kij×Li×Lj)^(1 / 2)…(2) however, A: the value of A in equation (1), Σij: The sum symbol when i and j are changed from 1 to the number of data points of the standard explanatory variables. Kij: The Laplace kernel calculation result of the i-th and j-th explanatory variables of the data of the reference explanatory variable, Li: The value of the index of predictive performance calculated based on the i-th value of the data of the reference objective variable and the output value when the i-th sample of the data of the reference explanatory variable is input into the predictive model. Lj: The value of the predictive performance index calculated based on the j-th value of the data of the reference objective variable and the output value when the j-th sample of the data of the reference explanatory variable is input into the predictive model. is.

[0054] By using the MMD value based on the Laplace kernel for D in equation (1) and then using the value calculated by equation (2) as A in equation (1), it is possible to estimate the value of the performance index when the prediction performance is the worst, based on the MMD distance between the data of the reference explanatory variable and the data of the second explanatory variable.

[0055] <Explanation of operation> Next, an example of the overall operation according to this embodiment will be described in detail with reference to the block diagram of Fig. 4 and the flowchart of Fig. 5. Fig. 5 is an example of a flowchart showing the processing procedure of the estimation device 1 in the first embodiment.

[0056] First, when the processing of the estimation device 1 is started, the second data acquisition unit 10 acquires data of the second explanatory variables (step S1). Next, the control unit 30 acquires the first data 41, the prediction model 42, and the parameter information 43 from the storage unit 40 (step S2). The reference data selection unit 3 selects reference data based on the data of the second explanatory variable, the first data 41, and the parameter information 43 (step S3). The performance estimation unit 4 calculates an estimated value of the performance index using equation (1) based on the reference data, the data of the second explanatory variable, the prediction model 42, and the parameter information 43 (step S4). Finally, the estimation device 1 outputs the estimated value of the performance index obtained by the performance estimation unit 4 through the output unit 20 (step S5).

[0057] 5 is merely an example of this embodiment and does not necessarily limit the operation flow according to this embodiment. As a specific example, the process S1 for acquiring the second explanatory variable may be executed after the process (step S2) for acquiring the first data 41, the prediction model 42, and the parameter information 43. Furthermore, step S2 may be divided into multiple steps, and for example, the process for acquiring the prediction model 42 in step S2 may be executed after the process (step S3) for selecting the reference data.

[0058] <Example> Next, as an example of an estimation device according to an embodiment of the present disclosure, an example of operation of the system assuming a specific use case will be described.

[0059] (Example 1: Forecasting demand for ice cream in a supermarket one month from now) First, in Example 1, an example will be described in which the demand for ice cream one month from now is predicted as the objective variable in a physical supermarket, using the current month, date, day of the week, season, weather, temperature, humidity, number of customers, and number of related products sold as explanatory variables. In this case, the actual sales volume cannot be known until one month has passed and the ice cream is actually sold. In other words, the demand for ice cream one month from now, which is the value of the objective variable, cannot be known until one month has passed since the prediction was made, and the prediction performance of the current prediction model, more specifically, for example, the mean squared error of the predicted value of the prediction model for the most recent week, cannot be known. Therefore, the present disclosure estimates the mean squared error of the prediction model's prediction for the most recent week.

[0060] The prediction model is created using data acquired over a certain period of time in the past, for example, three years' worth of data (i.e., including daily explanatory variable data over the three years and daily objective variable data, which is the number of ice creams sold one month after each day) as data for creating the prediction model. Here, the three years' worth of data is referred to as data for the first explanatory variable and data for the objective variable corresponding to the data for the first explanatory variable, and is collectively referred to as the first data. The created prediction model and the first data are stored in a predetermined storage area.

[0061] Next, during actual operation, the mean square error of the prediction model for the most recent week is estimated. That is, the data for the most recent week's explanatory variables is used as the data for the second explanatory variables. Estimating prediction performance requires a method for selecting reference data, a method for calculating the value of D in equation (1), and a method for calculating the value of A in equation (1).

[0062] First, as the reference data, data from the first data with dates that match the dates included in the data for the second explanatory variable are selected. This is based on the idea that the predictive performance of the prediction model for the same month in the same year will generally be similar, i.e., the predictive performance of the prediction model for the data for the second explanatory variable is highly relevant. This idea can be obtained, for example, from experience with supermarket sales or from a detailed analysis of the first data (i.e., domain knowledge). This makes it possible to more appropriately estimate the mean squared error value for the most recent week, which is the target of estimation, based on data that is highly relevant to the data for the second explanatory variable.

[0063] Next, the MMD of the data for the reference explanatory variable and the data for the second explanatory variable is used as the value of D in equation (1). More specifically, MMD is calculated using a Laplace kernel with sigma (parameter) set to 1. Furthermore, Z-Score normalization is performed before calculating D. The method for calculating the value of D using MMD, combined with the method for calculating the value of A in equation (1) using equation (3) described below, is a method for estimating mean squared error based on distribution robust optimization, and is an estimation method based on theoretical analysis.

[0064] The value of A in equation (1) is calculated by the following equation (3) by specifying the above equation (2) based on the reference data and the prediction model. The calculation of the value of A using equation (3) is a calculation method derived from analysis based on distribution robust optimization.

[0065] <Calculation formula for rate of change A> A=(Σij Kij×(Yi−Pi)^2×(Yj−Pj)^2)^(1 / 2)…(3) however, A: the value of A in equation (1), Σij: The sum symbol when i and j are changed from 1 to the number of data points of the standard explanatory variables. Kij: The value of the Laplace kernel calculation result with sigma (parameter) set to 1 for the i-th and j-th explanatory variables of the Z-Score normalized explanatory variable data. Yi: the i-th value of the data of the reference variable, Yj: the jth value of the data of the reference variable, Pi: The output value when the i-th sample of the standard explanatory variable data is input to the prediction model Pj: The output value when the jth sample of the data of the standard explanatory variable is input to the prediction model. is.

[0066] By using the method for selecting reference data specified above, the method for calculating the value of D in equation (1), and the method for calculating the value of A in equation (1) using equation (3), it is possible to estimate the mean square error of the forecast model for the most recent week during actual operation.

[0067] FIG. 6 is a schematic diagram of a user interface that displays the mean square error estimation results according to this embodiment to the user in the form of a graph. The graph in FIG. 6 shows the actual measured mean square error values ​​and the predicted values ​​according to this embodiment in a line graph, with the horizontal axis representing the date and the vertical axis representing the mean square error. FIG. 6 also shows a bar graph showing the breakdown of P in equation (1) for the predicted values ​​of the mean square error according to this embodiment, broken down into the value of B in the first term on the right-hand side of equation (1) and the value of D×A in the second term. Because it takes one month to determine the value of the dependent variable, i.e., the number of ice creams sold, the actual measured mean square error values ​​can only be calculated and plotted for predictions up to one month in advance. On the other hand, the prediction performance estimation according to this disclosure can be calculated to include predictions up to today. Simultaneous display of the breakdown also makes it easier for users to understand the estimated values ​​of prediction performance and changes in the estimated values.

[0068] Example 2: Disease diagnosis prediction based on early symptoms Next, as Example 2, an example of disease diagnosis prediction based on early symptoms will be described. More specifically, a prediction model is used to predict the patient's illness, such as a cold, influenza, allergic rhinitis, streptococcal infection, acute bronchitis, or measles, using the patient's today's body temperature, yesterday's body temperature, the day before yesterday's body temperature, and symptoms such as sore throat, stuffy nose, cough, and fatigue as explanatory variables. The objective variable is a predictive variable, and the patient's illness, such as a cold, influenza, allergic rhinitis, streptococcal infection, acute bronchitis, or measles, is predicted. Actual illness diagnosis requires a doctor's diagnosis, and it may take time for the diagnosis to be made, the symptoms required for diagnosis to develop, and the test results required for diagnosis to be obtained. Therefore, the prediction accuracy of the prediction model for, for example, the most recent 30 patients cannot be known until the diagnosis results for all patients are available. Therefore, the estimation device according to the present disclosure estimates the prediction accuracy of predictions for the most recent 30 patients.

[0069] First, a prediction model is created by using, as data for creating the prediction model, symptoms, which are explanatory variable data, and disease assessment results, which are objective variable data, for a certain number of people in the past, for example, 1,000 people. Here, the data for the past 1,000 people is referred to as first explanatory variable data and objective variable data corresponding to the first explanatory variable data, and is collectively referred to as first data. The created prediction model and first data are stored in a predetermined storage area.

[0070] Next, during actual operation, the prediction accuracy of the prediction model for the most recent 30 people is estimated. That is, the data of the explanatory variables for the most recent 30 people is used as the data of the second explanatory variable. Estimating the prediction accuracy requires a method for selecting reference data, a method for calculating the value of D in equation (1), and a method for calculating the value of A in equation (1).

[0071] In this example, all of the first data is used as the reference data. In other words, all of the first data is selected as the reference data. In this embodiment, the value of D in formula (1) is calculated based on whether the symptom is unknown compared to each symptom in the reference data, as described below. In reality, all of the first data are known symptoms. Therefore, by selecting all of the first data as the reference data, all of the first data can be considered as known symptoms, enabling more suitable estimation of prediction accuracy. Note that, depending on the calculation method for the value of D, the number of first data, and the presence or absence of domain knowledge related to the symptoms, the reference data may be selected in a more detailed manner; however, in this embodiment, all of the data is selected for the reasons described above.

[0072] The value of D in equation (1) is the proportion of samples in the Min-Max normalized second explanatory variable data where no sample of the reference explanatory variable data exists within a certain distance, for example, a Euclidean distance of 1.0, from each sample. This calculates the proportion of samples in the second explanatory variable data that are not included in the first explanatory variable data, i.e., samples with explanatory variables unknown to the predictive model that the predictive model has not trained or validated. If all samples in the second explanatory variable data are included in the first explanatory variable data, the value of D will be 0.

[0073] Finally, the value of A in equation (1) is a constant of -1. This, along with the calculation method for D, is an estimation method based on the assumption that predictions made by a prediction model for samples that the model has not been trained or validated on will generally be incorrect.

[0074] By selecting the reference data as described above and setting the values ​​of D and A in equation (1), we can estimate the prediction accuracy of the prediction model for the most recent 30 people, assuming that the accuracy of the prediction model for the data used to create the model is the upper limit, and that predictions by the prediction model for samples that the prediction model has not trained or verified will be incorrect.

[0075] Although the present disclosure has been described above with reference to the above-described embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each of the above-described means and functions may be executed by an information processing device installed and connected anywhere on a network, i.e., may be executed by so-called cloud computing.

[0076] The above-described program can be stored and supplied to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). The program may also be supplied to a computer by various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable media can supply the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.

[0077] <Additional Notes> A part or all of the above-described embodiments can be described as follows: The following provides an outline of the configurations of the estimation device, estimation method, and program according to the present invention. However, the present invention is not limited to the following configurations. (Appendix 1) a data selection means for selecting, from first explanatory variable data prepared according to a preset prediction model and first dependent variable data corresponding to the first explanatory variable data, the first explanatory variable data and the first dependent variable data corresponding to the first explanatory variable data, based on second explanatory variable data to which no dependent variable is associated; a performance estimation means for estimating the prediction performance of the prediction model for the second explanatory variable data based on a comparison between prediction data obtained by inputting the first explanatory variable data selected for the prediction model and the first dependent variable data corresponding to the first explanatory variable data; An estimation device having: (Appendix 2) 10. The estimation device of claim 1, the performance estimation means estimates the prediction performance of the prediction model for the second explanatory variable data based on a comparison between the prediction data and the first dependent variable data and a comparison between the second explanatory variable data and the selected first explanatory variable data. Estimation device. (Appendix 3) 3. The estimation device according to claim 2, the performance estimation means estimates the prediction performance of the prediction model for the selected first explanatory variable data based on a comparison between the prediction data and the first dependent variable data, and based on a difference in data content between the second explanatory variable data and the selected first explanatory variable data based on a preset criterion. Estimation device. (Appendix 4) 3. The estimation device according to claim 2, the performance estimation means estimates that the performance of the prediction model for the second explanatory variable data will deteriorate as a difference based on a preset criterion in data content between the second explanatory variable data and the selected first explanatory variable data increases. Estimation device. (Appendix 5) 5. The estimation device according to claim 4, the performance estimation means estimates that a value of prediction performance of the prediction model for the selected first explanatory variable data based on a comparison between the prediction data and the first dependent variable data will deteriorate as a difference based on a preset criterion in data content between the second explanatory variable data and the selected first explanatory variable data increases. Estimation device. (Appendix 6) 5. The estimation device according to claim 4, the performance estimation means calculates a deterioration rate of performance of the prediction model for the second explanatory variable data according to a difference between a distribution of the second explanatory variable data and a distribution of the selected first explanatory variable data, using an equation for a deterioration rate of prediction performance based on distribution robust optimization calculated from the selected first explanatory variable data and a comparison between the prediction data and the selected first dependent variable data, and estimates the performance of the prediction model for the second explanatory variable data based on the deterioration rate; Estimation device. (Appendix 7) 10. The estimation device of claim 1, the data selection means selects, from the first explanatory variable data, data that is different from data used at the time of training the prediction model. Estimation device. (Appendix 8) 10. The estimation device of claim 1, the data selection means selects, from the first explanatory variable data, data that has a predetermined correlation with the second explanatory variable data with respect to the prediction performance of the prediction model. Estimation device. (Appendix 9) 10. The estimation device of claim 1, the data selection means selects, from the first explanatory variable data, the first explanatory variable data whose difference from the second explanatory variable data based on a preset criterion of data content is smaller than that of other data. Estimation device. (Appendix 10) selecting, from first explanatory variable data prepared according to a preset prediction model and first dependent variable data corresponding to the first explanatory variable data, the first explanatory variable data and the first dependent variable data corresponding to the first explanatory variable data, based on second explanatory variable data to which no dependent variable is associated; estimating the prediction performance of the prediction model for the second explanatory variable data based on a comparison between prediction data obtained by inputting the selected first explanatory variable data into the prediction model and the first dependent variable data corresponding to the selected first explanatory variable data; Estimation method. (Appendix 11) selecting, from first explanatory variable data prepared according to a preset prediction model and first dependent variable data corresponding to the first explanatory variable data, the first explanatory variable data and the first dependent variable data corresponding to the first explanatory variable data, based on second explanatory variable data to which no dependent variable is associated; estimating the prediction performance of the prediction model for the second explanatory variable data based on a comparison between prediction data obtained by inputting the selected first explanatory variable data into the prediction model and the first dependent variable data corresponding to the selected first explanatory variable data; A computer-readable storage medium that stores a program for causing a computer to execute a process. [Explanation of symbols]

[0078] 1 Estimation device 3. Reference data selection section 4 Performance estimation part 10 Second data acquisition unit 20 Output section 30 Control Unit 40 Storage section 41 First Data 42 Predictive Models 43 Parameter Information 100 Estimator 101 CPU 102 ROM 103 RAM 104 Programs 105 Storage device 106 Drive device 107 Communication Interface 108 Input / Output Interface 109 Bus 110 Storage medium 111 Communication Network 121 Data Selection Methods 122 Performance estimation means

Claims

1. a data selection means for selecting, from first explanatory variable data prepared according to a preset prediction model and first dependent variable data corresponding to the first explanatory variable data, the first explanatory variable data and the first dependent variable data corresponding to the first explanatory variable data, based on second explanatory variable data to which no dependent variable is associated; a performance estimation means for estimating the prediction performance of the prediction model with respect to the second explanatory variable data based on a comparison between prediction data obtained by inputting the first explanatory variable data selected for the prediction model and the first dependent variable data corresponding to the first explanatory variable data; An estimation device having:

2. The estimation device according to claim 1, the performance estimation means estimates prediction performance of the prediction model with respect to the second explanatory variable data based on a comparison between the prediction data and the first dependent variable data and a comparison between the second explanatory variable data and the selected first explanatory variable data. Estimation device.

3. 3. The estimation device according to claim 2, the performance estimation means estimates the prediction performance of the prediction model with respect to the selected first explanatory variable data based on a comparison between the prediction data and the first dependent variable data, and based on a difference in data content between the second explanatory variable data and the selected first explanatory variable data based on a preset criterion. Estimation device.

4. 3. The estimation device according to claim 2, the performance estimation means estimates that the performance of the prediction model for the second explanatory variable data will deteriorate as a difference based on a preset criterion in data content between the second explanatory variable data and the selected first explanatory variable data increases. Estimation device.

5. The estimation device according to claim 4, the performance estimation means estimates that a value of prediction performance of the prediction model for the selected first explanatory variable data based on a comparison between the prediction data and the first dependent variable data will deteriorate as a difference based on a preset criterion in data content between the second explanatory variable data and the selected first explanatory variable data increases. Estimation device.

6. The estimation device according to claim 4, the performance estimation means calculates a deterioration rate of performance of the prediction model with respect to the second explanatory variable data according to a difference between a distribution of the second explanatory variable data and a distribution of the selected first explanatory variable data, using an equation for a deterioration rate of prediction performance based on distribution robust optimization calculated from the selected first explanatory variable data and a comparison between the prediction data and the selected first dependent variable data, and estimates the performance of the prediction model with respect to the second explanatory variable data based on the deterioration rate; Estimation device.

7. The estimation device according to claim 1, the data selection means selects, from the first explanatory variable data, data that is different from data used at the time of training the prediction model. Estimation device.

8. An estimation device according to claim 1, The predictive model is created by machine learning. Estimation device.

9. An information processing device, selecting, from first explanatory variable data prepared according to a preset prediction model and first dependent variable data corresponding to the first explanatory variable data, the first explanatory variable data and the first dependent variable data corresponding to the first explanatory variable data, based on second explanatory variable data to which no dependent variable is associated; estimating the prediction performance of the prediction model for the second explanatory variable data based on a comparison between prediction data obtained by inputting the selected first explanatory variable data into the prediction model and the first dependent variable data corresponding to the selected first explanatory variable data; Estimation method.

10. selecting, from first explanatory variable data prepared according to a preset prediction model and first dependent variable data corresponding to the first explanatory variable data, the first explanatory variable data and the first dependent variable data corresponding to the first explanatory variable data, based on second explanatory variable data to which no dependent variable is associated; estimating the prediction performance of the prediction model for the second explanatory variable data based on a comparison between prediction data obtained by inputting the selected first explanatory variable data into the prediction model and the first dependent variable data corresponding to the selected first explanatory variable data; A program that causes a computer to execute a process.

Citation Information

Patent Citations

  • Work support system, and work support method

    JP2022034697A

  • Method to continuously diagnose and model changes of real-valued streaming variables

    US20070260563A1

  • Estimation system, estimation method, and estimation program

    WO2019229977A1

  • Mathematical model generation system, mathematical model generation method, and mathematical model generation program

    WO2021229648A1