Information processing apparatus, information processing method, program

The information processing apparatus uses multiple methods and normalization techniques to select important explanatory variables with high accuracy by averaging and normalizing importance scores, addressing the limitations of single-model approaches.

JP7703834B2Active Publication Date: 2025-07-08FUJI ELECTRIC CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2020149044
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-18
Filing Date
2020-09-04
Publication Date
2025-07-08
Estimated Expiration
2040-09-04

AI Technical Summary

Technical Problem

Existing methods for selecting important explanatory variables for a target variable with continuous values, such as multiple regression, support vector regression, and decision trees, often fail to achieve high accuracy due to their dependency on a single model or method.

Method used

An information processing apparatus that calculates importance values using multiple methods (e.g., multiple regression, support vector regression, and decision trees) and selects variables based on a combination of these values, incorporating normalization and weight coefficients to account for variability across methods.

Benefits of technology

This approach allows for the selection of important explanatory variables with high accuracy by averaging and normalizing the importance scores across different methods, reducing reliance on a single model and enhancing the robustness of the selection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007703834000001
    Figure 0007703834000001
  • Figure 0007703834000002
    Figure 0007703834000002
  • Figure 0007703834000003
    Figure 0007703834000003
Patent Text Reader

Abstract

To provide an information processing device capable of selecting highly-accurate and important explanatory variables.SOLUTION: An information processing device includes: a calculation section which calculates a first value indicating an important level for each of a plurality of explanatory variables against prescribed objective variables, using each of a plurality of m different methods; and a first selection section which, based on a plurality of the first values corresponding to each of the m different methods, selects n explanatory variables.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing system, and a program.

Background Art

[0002] When using data with a large number of explanatory variables, it may be necessary to select important explanatory variables for the target variable (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] By the way, when the target variable is a continuous value, as a method for selecting important explanatory variables for the target variable, for example, there are methods using models based on multiple regression, support vector regression, decision trees, etc., and methods using correlation coefficients. However, when selecting explanatory variables using a specific single model or method, depending on the data, it may not always be possible to select important explanatory variables with high accuracy.

[0005] The present invention has been made in view of the above-described conventional problems, and an object thereof is to provide an information processing apparatus capable of selecting important explanatory variables with high accuracy.

Means for Solving the Problems

[0006] The main invention of the present invention for solving the above-described problems is a computing unit that calculates a first value indicating the importance of each of a plurality of explanatory variables with respect to a predetermined target variable using each of a plurality of m different methods, and a first selection unit that selects n explanatory variables based on the plurality of the first values corresponding to each of the m different methods. The information processing apparatus includes:

Effect of the Invention

[0007] According to the present invention, it is possible to provide an information processing apparatus capable of selecting important explanatory variables with high accuracy.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Mode for Carrying Out the Invention

[0009] At least the following matters become clear from the description in this specification and the accompanying drawings.

[0010] =====First Embodiment===== FIG. 1 is a diagram showing the relationship between a data processing apparatus 10 and an information processing apparatus 20 which is an embodiment of the present invention.

[0011] The data processing apparatus 10 is an apparatus that outputs data x1 to x10 indicating measurement values and various information acquired in a predetermined power system (hereinafter referred to as "power system A"). Specifically, the data processing apparatus 10 outputs ten pieces of data x1 to x10 including data indicating, for example, the temperature, humidity, sunshine amount, etc. of the power system A and data indicating the attributes of the measurement date (for example, whether it is a "weekday" or a "holiday") at predetermined intervals.

[0012] The information processing device 20 is a device that selects data important for the power demand of customers in power system A among data x1 to x10. Specifically, the information processing device 20 is a device that selects important explanatory variables when using the power demand as the target variable and data x1 to x10 as the explanatory variables. Note that the "important explanatory variable" is, for example, a variable that has a large influence on the power demand. Also, hereinafter, in this embodiment, data x is appropriately referred to as an "explanatory variable", and data y (described later) indicating the power demand is referred to as a "target variable".

[0013] The information processing device 20 is a computer including a CPU (Central Processing Unit) 30, a memory 31, a storage device 32, an input device 33, a display device 34, and a communication device 35.

[0014] The CPU 30 realizes various functions in the information processing device 20 by executing programs stored in the memory 31 and the storage device 32.

[0015] The memory 31 is, for example, a RAM (Random-Access Memory) or the like, and is used as a temporary storage area for programs, data, and the like.

[0016] The storage device 32 is a non-volatile storage device that stores various information such as programs and data sets executed by the CPU 30.

[0017] The input device 33 is a device that accepts input of commands and data by the user, and includes an input interface such as a keyboard and a touch sensor that detects a touch position on a touch panel display.

[0018] The display device 34 is, for example, a display device, and the communication device 35 exchanges various programs and data with other computers via a network (not shown).

[0019] FIG. 2 is a diagram showing an example of information stored in the storage device 32. The storage device 32 stores a control program 40, a data set 41, importance data 42, normalization data 43, and difference data 44.

[0020] The control program 40 is a program for realizing various functions of the information processing apparatus 20, and includes, for example, an OS (Operating System) or the like.

[0021] As shown in FIG. 3, the data set 41 includes past data x1 to x10 and y output from the data processing apparatus 10.

[0022] Here, "data x1" is data indicating, for example, the temperature of power system A, and "data x2" is data indicating the humidity of power system A. Also, "data x9" is data indicating the attribute of the day when data such as x1 was acquired (for example, whether it is a "weekday" or a "holiday"), and "data x10" is data indicating, for example, the sunshine amount of power system A. Further, "data y" is the power consumption of power system A corresponding to data x1 to x10. In this embodiment, in data x9, "1" indicates a weekday and "0" indicates a holiday.

[0023] Also, data x3 to x8 are data related to a predetermined physical quantity (for example, wind speed) of power system A similar to x1, x2, etc., and thus detailed description is omitted here.

[0024] The data set 41 is a set of numerical data including data x1 to x10 at different times, i (for example, 10,000) in number. The first data of the data set 41 is acquired at time t1, for example. The data x1 indicating temperature is "27°C", the data x2 indicating humidity is "62%", the data x9 indicating the attribute of the acquisition date is "1 (weekday)", and the data x10 indicating the sunshine amount is "0.6MJ / m 2 ". Also, the data y indicating the power consumption at this timing is "2443kW". In this embodiment, the "numerical data" means data in which data y indicates a numerical value, and is also referred to as continuous value data or continuous data.

[0025] FIG. 4 is a diagram showing an example of importance data 42. The importance data 42 is data including values indicating the importance of each of data x1 to x10, which are explanatory variables, with respect to data y, which is an objective variable. Although details will be described later, the importance data 42 of the present embodiment is obtained by using three different variable selection methods M1 to M3 (for example, a method using a multiple regression, a kernel method, and a model based on a decision tree). Further, in the importance data 42 obtained by the method M1, the value of data x1 is “9.1” and the value of data x2 is “5.1”. In the present embodiment, it is shown that the larger the value of the importance data 42, the more important it is with respect to the objective variable. Note that the values indicating the importance of each of data x1 to x10 correspond to “first values”.

[0026] FIG. 5 is a diagram showing an example of normalized data 43. The normalized data 43 is data obtained by normalizing the importance data 42 of each of the methods M1 to M3. In the present embodiment, the importance data 42 is subjected to a predetermined normalization process (described later) so that the values S(j, k) of the normalized data 43 of each of the methods M1 to M3 fall within the range of 0 to 1. Here, j is a value indicating any one of 1 to 10, k is a value indicating any one of 1 to 3, and the value S(j, k) is the value of the normalized data 43 of data xj obtained based on the method Mk. Thereby, the importance of data x1 to x10 obtained by different methods M1 to M3 can be objectively grasped. Note that the respective values of the normalized data 43 correspond to “second values”.

[0027] FIG. 6 is a diagram showing an example of difference data 44. The difference data 44 is data indicating how much the normalized values S(j, k) obtained by each of the methods M1 to M3 deviate from the average value in each of data x1 to x10, and includes a value D(j, k).

[0028] Here, in data xj, the average value Ave(xj) of the normalized data 43 obtained by the methods M1 to M3 can be calculated by the following formula (1).

[0029] Ave(xj) = (S(j,1) + S(j,2) + S(j,3)) / 3 ··· (1) And from Equation (1), when j = 1, the average value Ave(x1) is Ave(x1) = (0.91 + 0.85 + 0.93) / 3. Also, the value of the square of the deviation between the value of the normalized data 43 of the data xj obtained based on the method Mk and the average value Ave(xj), D(j,k), can be calculated by Equation (2).

[0030] D(j,k) = (Ave(xj) ― S(j,k)) 2 ··· (2) Here, when j = 1 and k = 1, the value D(j,k) = (-0.013) 2 is obtained. And in this embodiment, by calculating such a value D(j,k), it is possible to grasp how far the value S(j,k) indicating the importance obtained by the methods M1 to M3 in the data xj is from the average value.

[0031] ==Functional Blocks of Information Processing Apparatus 20== FIG. 7 is a diagram showing an example of functional blocks implemented in the information processing apparatus 20. When the CPU 30 of the information processing apparatus 20 executes the control program 40, an importance calculation unit 50, a normalization unit 51, selection units 52, 56, a weight coefficient calculation unit 53, a multiplication unit 54, and an addition unit 55 are implemented in the information processing apparatus 20.

[0032] The importance calculation unit 50 calculates the importance of the explanatory variables x1 to x10 based on three different methods M1 to M3. Specifically, as the method M1, the importance calculation unit 50 calculates the importance of the explanatory variables using a model based on multiple regression. Also, as the method M2, the importance calculation unit 50 calculates the importance of the explanatory variables using a model based on support vector regression, and as the method M3, the importance calculation unit 50 calculates the importance of the explanatory variables using a model based on a decision tree. As a result, for example, the importance data 42 shown in FIG. 4 can be obtained.

[0033] Hereinafter, in the present embodiment, the "model" refers to a function represented by y = f(x1, x2, ···, x10). Also, the model is created, for example, by using the dataset 41 as teaching data and stored in the storage device 32.

[0034] In the present embodiment, for example, the models used in each of the methods M1 to M3 are assumed to be multiple regression, support vector regression, and decision tree-based models, but it is not limited to this. For example, models based on ridge regression, k-nearest neighbor method, nearest neighbors, etc. may be used. Further, the importance calculation unit 50 may calculate the importance by a method using the correlation coefficient without using a model.

[0035] Also, as an index indicating the importance of the explanatory variables, the importance calculation unit 50 may use, for example, the importance (Permutation Importance) obtained by rearranging the explanatory variables, disclosed in Japanese Patent Application Laid-Open No. 2019-121162. Further, as the importance of the explanatory variables (that is, features), a method of recursive feature elimination may be used, and a general index when the features are reduced may be used.

[0036] In addition, for example, the importance of the explanatory variables may be obtained by approximating a model constructed using a neural network with a linear regression method such as LIME (Local Interpretable Model-agnostic Explanations). Also, the Shapley values (SHAP (SHapley Additive exPlanation) Values) obtained using cooperative game theory may be used as a value indicating the importance of the explanatory variables.

[0037] Also, the "method for calculating importance" executed by the importance calculation unit 50 is a process excluding the process of selecting variables among general variable selection (Feature Selection) methods. In the present embodiment, three (m) different methods are used, but it is not limited to this, and a plurality of methods may be used.

[0038] The normalization unit 51 performs normalization processing for each of the methods M1 to M3 on the value of the importance data 42 to calculate the normalized data 43. As the normalization method, for example, the maximum value and the minimum value may be set as predetermined values respectively, or the average value and the variance may be used.

[0039] The selection unit 52 selects explanatory variables for which all values of the normalized data 43 with respect to the explanatory variables are equal to or greater than a predetermined value τ (for example, 0.01). Thereby, explanatory variables with low importance can be removed. For example, in the data x3 which is an explanatory variable in FIG. 5, the value based on the method M2 is "0.008". In such a case, even if the values based on other methods of the data x3 are equal to or greater than the predetermined value τ (for example, 0.01), the data x3 will not be selected by the selection unit 52. The selection unit 52 corresponds to the "second selection unit".

[0040] The weight coefficient calculation unit 53 calculates weight coefficients Wk (k = 1 to 3) corresponding to each of the methods M1 to M3. As shown in FIG. 8, the weight coefficient calculation unit 53 includes an average value calculation unit 60, a difference calculation unit 61, a total value calculation unit 62, and a coefficient calculation unit 63.

[0041] The average value calculation unit 60 calculates the average value Ave(xj) of the normalized data 43 for each data xj using the above-described formula (1).

[0042] The difference calculation unit 61 calculates the squared deviation values D(j,1) to D(j,3) between each of the three values of the normalized data 43 of the data xj and the average value Ave(xj) using the above-described formula (2).

[0043] The total value calculation unit 62 calculates the total value Tk (k = 1 to 3) of the values D(1,k) to (10,k) in the method Mk. That is, the total value calculation unit 62 calculates the sum of the ten squared deviation values of D(1,k) to (10,k) based on the following formula (3) in the method Mk.

[0044] Tk = D(1,k) + D(2,k) + ··· + D(10,k) ··· (3) Accordingly, it is possible to objectively grasp how much each of the methods M1 to M3 deviates from the average of the methods M1 to M3.

[0045] The coefficient calculation unit 63 calculates the weight coefficient Wk for the method Mk based on the total value Tk of the method Mk. Specifically, the coefficient calculation unit 63 calculates the weight coefficient Wk using Equation (4).

[0046] Wk = 1 / Tk ··· (4) Therefore, for the method Mk that deviates significantly from the average of the methods M1 to M3, the weight coefficient Wk becomes small.

[0047] The multiplication unit 54 in FIG. 7 multiplies the value S(j,k) of the normalized data 43 of the method Mk and the weight coefficient Wk, and calculates Wk × S(j,k).

[0048] The addition unit 55 adds the multiplication results of the multiplication unit 54 corresponding to each of the methods Mk (k = 1 to 3). Specifically, the addition unit 55 calculates the total value Ej of the multiplication results for each explanatory variable (data xj) using Equation (5).

[0049] Ej = W1 × S(j,1) + W2 × S(j,2) + W3 × S(j,3) ··· (5) The selection unit 56 selects a predetermined n (for example, 5) explanatory variables with large total values Ej among the total values Ej. Note that the selection unit 56 corresponds to the "first selection unit".

[0050] <<<Explanatory Variable Selection Process S10>>> FIG. 9 is a flowchart showing an example of the processing executed by each functional block of the information processing apparatus 20. First, the importance calculation unit 50 calculates the importance of all explanatory variables (data x1 to x10) using a predetermined model (not shown) corresponding to the method Mk stored in the storage device 32 (S20: calculation process).

[0051] FIG. 10 is a diagram showing an example of the process S20 executed by the importance calculation unit 50. The importance calculation unit 50 first calculates the importance of the explanatory variables using a model based on multiple regression corresponding to method M1 (S30). Also, the importance calculation unit 50 calculates the importance of the explanatory variables using a model based on support vector regression corresponding to method M2 (S31). Further, the importance calculation unit 50 calculates the importance of the explanatory variables using a model based on a decision tree corresponding to method M3 (S32).

[0052] As a result, for example, importance data 42 shown in FIG. 4 is obtained. Also, the normalization unit 51 normalizes the values of the importance data 42 for method Mk (S21). Thereby, for example, normalized data 43 shown in FIG. 5 is obtained.

[0053] The selection unit 52 selects explanatory variables for which all the values of the normalized data 43 for the explanatory variables are equal to or greater than a predetermined value τ (for example, 0.01) (S22). As a result, in the following processes S23 to S26, only the values of the normalized data 43 for the selected explanatory variables are used. When explaining processes S23 to S26, the explanatory variables are the explanatory variables selected in process S22, but for the sake of convenience, they may simply be referred to as explanatory variables.

[0054] The weight coefficient calculation unit 53 calculates the weight coefficient Wk for method Mk (S23). FIG. 11 is a flowchart showing an example of the calculation process S23 of the weight coefficient Wk. The average value calculation unit 60 calculates the average value Ave(xj) of the normalized data 43 for each selected explanatory variable (data xj) using Equation (1) (S40).

[0055] The difference calculation unit 61 calculates the squared deviation values D(j,1) to D(j,3) between each of the three values of the normalized data 43 of the data xj and the average value Ave(xj) using Equation (2) (S41).

[0056] In method Mk, the total value calculation unit 62 calculates the total value Tk (k = 1 to 3) of the values D(j, k) using Equation (3) (S42). In process S42, only the j selected by the selection unit 52 out of j = 1 to 10 is used. For example, when j = 1, 2, 4 to 10 are selected, the total value Tk (k = 1 to 3) is the sum of the nine selected values D(j, k).

[0057] When the total value Tk is calculated, the coefficient calculation unit 63 calculates the weight coefficient Wk for method Mk using Equation (4) (S43). As a result, process S23 in FIG. 9 ends. When the weight coefficient Wk is calculated, the multiplication unit 54 calculates Wk × S(j, k), which is the product of the weight coefficient Wk and the value S(j, k) of the normalized data 43 (S24).

[0058] Then, the addition unit 55 adds the multiplication results of the multiplication unit 54 corresponding to each of methods Mk (k = 1 to 3) using the above-described Equation (5) to calculate the total value Ej (S25). When the total value Ej is calculated, the selection unit 56 selects a predetermined n (for example, 5) explanatory variables with large total values Ej from the total values Ej (S26: first selection process). In processes S24 to 26 as well, similar to process S23, only the j selected by the selection unit 52 out of j = 1 to 10 is used.

[0059] As a result, in the information processing apparatus 20, n explanatory variables with high importance are selected from among the explanatory variables (data x1 to x10).

[0060] =====Second Embodiment (When Using a Learning Model)===== Here, a case will be described in which three different learning models are used as each of methods M1 to M3, and a weight coefficient W corresponding to the accuracy of the learning model is used. The information processing apparatus 21 in FIG. 1 is the same apparatus as the information processing apparatus 20, and selects data important for the power demand of the customers in power system A from among data x1 to x10.

[0061] The information processing device 21 is a computer including a CPU 30, a memory 31, an input device 33, a display device 34, a communication device 35, and a storage device 100. Since the blocks other than the storage device 100 of the information processing device 21 are the same as the respective blocks of the information processing device 20, detailed description thereof is omitted here.

[0062] FIG. 12 is a diagram showing an example of information stored in the storage device 100. The storage device 100 stores a control program 110, a learning model 111, a data set 112, importance data 113, and normalization data 114.

[0063] The control program 110 is a program for realizing various functions of the information processing device 21, and includes, for example, an OS or the like.

[0064] The learning model 111 is a learning model that has learned the relationship between power demand and data x1 to x10 when the power demand is the target variable and the data x1 to x10 are the explanatory variables. Specifically, the learning model 111 is a model obtained by using the data x1 to x10 and the values of the power demand as teacher data and executing predetermined machine learning.

[0065] In the present embodiment, as the learning model 111, three learning models M L , M k , M T are included. Here, the "learning model M L " is a model obtained by using a regression method (for example, linear regression), and the "learning model M k " is a model obtained by using a kernel method (for example, support vector machine). Further, the "learning model M T " is a model obtained by using a tree structure method (for example, decision tree). Note that the three learning models M L , M k , M T of the present embodiment are pre-calculated by a predetermined server (not shown).

[0066] The dataset 112 is data including past data x1 to x10 and y output from the data processing device 10, and is, for example, the same as the dataset 41 shown in FIG. 3.

[0067] The importance data 113 is data including values indicating the importance of each of the data x1 to x10 as explanatory variables with respect to the data y as the target variable, and is, for example, the same as the importance data 42 shown in FIG. 4. Specifically, the importance data 113 in the present embodiment is for three learning models M L ,M k ,M T and is a value indicating the importance of the data x1 to x10 as explanatory variables when used. Each value of the importance data 113 corresponds to the "first value".

[0068] The value indicating importance can be calculated, for example, by using the above-described neighborhood method (e.g., k-nearest neighbor method, nearest neighbors method), Permutation Importance method, cooperative game theory method, wrapper method (e.g., Forward Selection, Backward Elimination), etc., to determine which of the 10 data x1 to x10 as explanatory variables is important. When calculating the importance data 113 when using each of the learning models M L ,M k ,M T , the same method (e.g., k-nearest neighbor method) may be used for the three learning models, or all different methods may be used. Also, when calculating the value indicating the importance of the explanatory variables when using, for example, the learning model M T , "Gini impurity" may be used.

[0069] In the above-described embodiment, it is assumed that three different variable selection methods M1 to M3 (e.g., methods using multiple regression, kernel method, and decision tree-based models) are used, but each of the methods M1 to M3 corresponds to the method of using the learning models M L ,M k ,M T .

[0070] The normalized data 114 is data obtained by normalizing the importance data 113 corresponding to each of the three learning models, and is the same as the normalized data 43 in FIG. 5. In this embodiment, for the learning models M L , M k , M T each, the importance data 113 is subjected to a predetermined normalization process (described later) so that the value S(j,k) of the normalized data 114 falls within the range of 0 to 1. Note that each value of the normalized data 114 corresponds to the "second value".

[0071] FIG. 13 is a diagram showing an example of functional blocks implemented in the information processing apparatus 21. When the CPU 30 of the information processing apparatus 21 executes the control program 110, an importance calculation unit 120, a normalization unit 121, a weight coefficient calculation unit 122, a multiplication unit 123, an addition unit 124, and a selection unit 125 are implemented in the information processing apparatus 20.

[0072] The importance calculation unit 120 calculates values indicating the importance of the data x1 to x10, which are explanatory variables, when using the three learning models M L , M k , M T each. As described above, since the three different methods M1 to M3 correspond to the methods of using the three learning models M L , M k , M T each, the importance calculation unit 120 and the importance calculation unit 50 are substantially the same. As a result, for example, the importance data 113 shown in FIG. 4 can be obtained.

[0073] The normalization unit 51 performs a normalization process on the values of the importance data 113 for each of the three learning models M L , M k , M T each to calculate the normalized data 114. As the normalization method, for example, the maximum value and the minimum value may be set as predetermined values, or the average and variance may be used.

[0074] The weight coefficient calculation unit 122 is for the three learning models ML , M k , M T Values indicating the accuracy of each of them are calculated as weight coefficients Wa (a = L, K, T). Specifically, the weight coefficient calculation unit 122 uses the past data x1 to x10 and y output from the data processing device 10 as test data, and three learning models M L , M k , M T to calculate values indicating the accuracy of each of them. Here, if the values indicating the accuracy of the three learning models M L , M k , M T are denoted as P L , P k , P T , then the weight coefficient Wa (a = L, K, T) is represented by Equation (6). Wa = Pa ··· (6) Note that the values P L , P k , P T indicating accuracy are values in the range of 0 to 1 that increase as the accuracy increases. Also, here, "a" is a variable indicating any one of L, K, and T. Therefore, the higher the accuracy of the learning model M L , M k , M T , the larger the weight coefficient Wa becomes.

[0075] The multiplication unit 123 multiplies the value S(j, a) of the normalized data 114 of the learning model Ma by the weight coefficient Wa, and calculates Wa × S(j, a) = Pa × S(j, a).

[0076] The addition unit 124 calculates the total value Fj of the multiplication results (Pa × S(j, a)) for each explanatory variable (data xj) using Equation (7).

[0077] Fj = P L × S(j, L) + P K × S(j, K) + P T × S(j, T) ··· (7) As a result, when the accuracy of the learning model and the value of the normalized data 114 are large, the total value Fj of the multiplication results for each explanatory variable (data xj) becomes large. The selection unit 125 selects a predetermined n (for example, 5) explanatory variables with large total value Ej from the total values Fj. Note that the selection unit 125 corresponds to the "first selection unit".

[0078] <<<Explanatory variable selection process S100>>> FIG. 14 is a flowchart showing an example of the processing executed by each functional block of the information processing apparatus 21. First, the importance calculation unit 120 uses three learning models M L , M k , M T to calculate values indicating the importance of each of the data x1 to x10 which are explanatory variables (S200: calculation process). Note that the details of the process S200 are the same as, for example, the processes S30 to 32 in FIG. 10.

[0079] As a result, for example, the importance data 113 shown in FIG. 4 is obtained. Also, the normalization unit 121 normalizes the values of the importance data 113 for each of the three learning models M L , M k , M T (S201). Thereby, for example, the normalized data 114 shown in FIG. 5 is obtained.

[0080] The weight coefficient calculation unit 122 calculates the accuracy of the learning model Ma as the weight coefficient Wa (S202). When the weight coefficient Wa is calculated, the multiplication unit 123 calculates Pa×S(j,a), which is the product of the weight coefficient Wa (=Pa) and the value S(j,a) of the normalized data 114 (S203).

[0081] Then, the addition unit 124 uses the above-described formula (7) to add the multiplication results of the multiplication unit 54 corresponding to each of the learning models Ma (a = L, K, T) for each explanatory variable, and calculates the total value Fj (S204). When the total value Fj is calculated, the selection unit 56 selects a predetermined n (for example, 5) explanatory variables with large total value Fj from the total values Fj (S205: first selection process).

[0082] As a result, in the information processing apparatus 21, among the explanatory variables (data x1 to x10), n explanatory variables with high importance will be selected. In the second embodiment, since three learning models are used, "m" corresponds to three.

[0083] =====Third Embodiment (Case of Using Multiple Importance Calculation Methods for Each Learning Model)===== Next, an embodiment in which explanatory variables are selected using three importance calculation methods for each learning model will be described. The information processing apparatus 22 in FIG. 1 is the same apparatus as the information processing apparatus 20, and selects data important for the power demand of customers in power system A from among data x1 to x10.

[0084] The information processing apparatus 22 is a computer including a CPU 30, a memory 31, an input device 33, a display device 34, a communication device 35, and a storage device 101. Since the blocks other than the storage device 101 of the information processing apparatus 22 are the same as the respective blocks of the information processing apparatus 20, detailed description thereof is omitted here.

[0085] FIG. 15 is a diagram showing an example of information stored in the storage device 101. The storage device 100 stores a learning model 111, a control program 130, a data set 131, importance data 132, normalization data 133, and difference data 134. Note that the learning model 111 is the same as the learning model including the three learning models M L , M k , M T described above.

[0086] The control program 130 is a program for realizing various functions of the information processing apparatus 22, and includes, for example, an OS and the like.

[0087] The data set 131 is data including past data x1 to x10 and y output from the data processing apparatus 10, and is, for example, the same as the data set 41 shown in FIG. 3.

[0088] As shown in Fig. 16, importance data 132 is data that includes values indicating the importance of each of the explanatory variables x1 to x10 obtained using three importance calculation methods for each learning model. The importance data 132 of the present embodiment is for three learning models M L , M k , M T respectively, and is obtained by using three different variable selection methods C1 (for example, the neighborhood method), method C2 (for example, the Permutation Importance method), and method C3 (for example, the cooperative game theory method).

[0089] For example, in the learning model M L , the value indicating the importance of data x1 obtained by using method C1 (for example, the neighborhood method) is "9.2", and the value of data x2 is "5.6". In the present embodiment, the larger the value of the importance data 132, the more important it is for the target variable. Note that the variable selection methods C1 to C3 are the same for each of the three learning models M L , M k , M T in this example, but they may be different. Specifically, as method C3 for the three learning models M L , M k , the cooperative game theory method may be used, and as method C3 for the learning model M T , a method based on Gini impurity may be used. Also, in the present embodiment, in one learning model, three variable selection methods are used, but if multiple methods are used, it is not limited to three. Note that the values indicating the importance of each of the data x1 to x10 correspond to the "first value".

[0090] The normalized data 133 is data obtained by normalizing the importance data 132 corresponding to each of the three methods C1 to C3 for each learning model. In the present embodiment, the importance data 113 is subjected to a predetermined normalization process (described later) so that the value Sa(j, k) of the normalized data 133 falls within the range of 0 to 1. Here, "a" is a variable indicating any one of L, K, and T, "j" is a variable indicating any one of the data x1 to x10 that are explanatory variables, and "k" is a variable indicating any one of the methods C1 to C3. Therefore, here, ten values Sa(j, k) based on a predetermined learning model Ma and a predetermined method Ck are normalized. Note that each value of the normalized data 114 corresponds to the "second value".

[0091] FIG. 18 is a diagram showing an example of the difference data 134. The difference data 134 is data indicating how much the average value Ave1(a, j) of the normalized data 133 of one learning model deviates from the average value Ave2(j) of the normalized data 133 of the three learning models for each of the data x1 to x10, and includes the value da(j).

[0092] Here, the average value Ave1(a, j) of the normalized data 133 of one learning model can be calculated by the following formula (8). Ave1(a, j) = (Sa(j, 1) + Sa(j, 2) + Sa(j, 3)) / 3 ··· (8) For example, in the data x1, the average value Ave1(L, 1) of the normalized data 133 of one learning model M L is the average value of the values of the methods C1 to C3 in the first row of FIG. 17 of the learning model M L =(S L (1, 1) + S L (1, 2) + S L (1, 3)) / 3).

[0093] Also, the average value Ave2(j) of the normalized data 133 of the three learning models can be calculated by the following formula (9). Ave2(j) = (Ave1(L, j) + Ave1(K, j) + Ave1(T, j)) / 3 ··· (9) For example, in the case of data x1, the average value Ave2(1) of the normalized data 133 of the three learning models is the average value of the values of the normalized data 133 in the first row of FIG. 17 (=(Ave1(L,1)+Ave1(K,1)+Ave1(T,1)) / 3).

[0094] And the value da(j) indicating how far the average value Ave1(a,j) is from the average value Ave2(j) of the normalized data 133 of the three learning models can be calculated by Equation (10).

[0095] da(j)=1 / │Ave1(a,j)-Ave2(j)│···(10) Therefore, the value d L (1) is, in the case of data x1, the reciprocal of the absolute value of the difference between the average value Ave1(L,1) of one learning model M L and the average value Ave2(1) of the normalized data 133 of the three learning models. For this reason, the larger the difference between the average value Ave1(L,1) and the average value Ave2(1), the smaller the value d L (1) becomes. Thereby, it is possible to grasp how far the average value Ave1(a,j) is from the average value Ave2(j).

[0096] Note that in the present embodiment, the average value Ave1(a,j) corresponds to the "first average value", and the average value Ave2(j) corresponds to the "second average value".

[0097] ==Functional Blocks of Information Processing Apparatus 22== FIG. 18 is a diagram showing an example of functional blocks implemented in the information processing apparatus 22. When the CPU 30 of the information processing apparatus 22 executes the control program 130, an importance calculation unit 150, a normalization unit 151, an accuracy calculation unit 152, a weight coefficient calculation unit 153, a multiplication unit 154, an addition unit 155, and a selection unit 156 are implemented in the information processing apparatus 22.

[0098] The importance calculation unit 150 calculates importance data 132 corresponding to each of the three learning model methods C1 to C3 for each of the data x1 to x10. As a result, for example, the importance data 132 shown in FIG. 16 can be obtained.

[0099] The normalization unit 151 performs normalization processing on the importance data 132 in one learning model for each of the methods C1 to C3 to calculate the normalized data 133. Specifically, the normalization unit 151, for example, normalizes the 10 data x1 to x10 included in the column of method C1 of the learning model M L Note that as the normalization method, for example, the maximum value and the minimum value may be set to predetermined values, or the average and variance may be used.

[0100] The accuracy calculation unit 152 calculates values indicating the accuracy of each of the three learning models M L , M k , M T Specifically, the accuracy calculation unit 152 uses the past data x1 to x10 and y output from the data processing device 10 as test data, and calculates values P L , M k , M T indicating the accuracy of each of the three learning models M L , P k , P T .

[0101] The weight coefficient calculation unit 153 calculates weight coefficients Wa (a = L, K, T) corresponding to each of the three learning models M L , M k , M T The weight coefficient calculation unit 53 includes a first calculation unit 160, a second calculation unit 161, and a third calculation unit 162 as shown in FIG. 20.

[0102] The first calculation unit 160 calculates the value da(j) of the learning model Ma for each of the data xj using the above-described equation (10).

[0103] In the learning model Ma, the second calculation unit 161 calculates the total value Σda of the values da(1) to da(10). That is, in the learning model Ma, the second calculation unit 161 calculates the total value Σda of the ten values of da(1) to da(10) based on the following formula (11).

[0104] Σda = da(1) + da(2) + ··· + da(10) ··· (11) Thereby, it is possible to objectively grasp how much the values obtained by each of the three learning models M L , M k , M T deviate from the average of the methods obtained by the three learning models.

[0105] Based on the total value Σda in the learning model Ma, the third calculation unit 162 calculates the weight coefficient Wa for the method using the learning model Ma. Specifically, the third calculation unit 163 calculates the weight coefficient Wa using formula (12).

[0106] Wa = (1 / α) × Pa × Σda ··· (12) Here, "α" is a predetermined constant for normalizing the weight coefficient Wa, and "Pa" is a value indicating the accuracy of the learning model Ma described above. Therefore, the weight coefficient Wa increases as the accuracy of the learning model Ma increases. Furthermore, the weight coefficient Wa decreases as Σda of the learning model Ma decreases (that is, as the difference between the average value Ave1(a,j) and the average value Ave2(j) increases).

[0107] The multiplication unit 154 in FIG. 19 multiplies the value Sa(j,k) of the normalized data 133 of the learning model Ma by the weight coefficient Wa to calculate Wa × Sa(j,k). Therefore, for example, regarding S L of the normalized data 133 related to the learning model M L (j,1) to S L (j,3), the weight coefficient W L is multiplied for all of them.

[0108] The addition unit 155 calculates the total value Gj of the multiplication results for each explanatory variable (data xj) using Equation (13). Gj = W L ×(S L (j, 1) + S L (j, 2) + S L (j, 3)) + W K ×(S K (j, 1) + S K (j, 2) + S K (j, 3)) + W T ×(S T (j, 1) + S T (j, 2) + S T (j, 3)) ··· (13) The selection unit 156 selects a predetermined n (for example, 5) explanatory variables with large total value Gj from the total values Fj. Note that the selection unit 156 corresponds to the "first selection unit".

[0109] <<<Explanatory Variable Selection Process S110>>> Figure 21 is a flowchart showing an example of the process executed by each functional block of the information processing apparatus 22. First, the importance calculation unit 150 calculates values indicating the importance of each of the data x1 to x10, which are explanatory variables, using three methods C1 to C3 for each of the three learning models M L , M k , M T (S300: Calculation process).

[0110] Figure 22 is a diagram showing an example of the process S300 executed by the importance calculation unit 150. The importance calculation unit 150 first calculates the importance of the explanatory variables using the methods C1 to C3 in the learning model M L (S320). Also, the importance calculation unit 150 calculates the importance of the explanatory variables using the methods C1 to C3 in the learning model M K (S321). Further, the importance calculation unit 150 calculates the importance of the explanatory variables using the methods C1 to C3 in the learning model M T (S322).

[0111] As a result, for example, importance data 132 shown in FIG. 16 can be obtained. Further, the normalization unit 151 normalizes the value of the importance data 132 (S301). As a result, for example, normalized data 133 shown in FIG. 17 can be obtained.

[0112] Further, the accuracy calculation unit 152 calculates Pa indicating the accuracy of the learning model Ma (S302). Then, the weight coefficient calculation unit 153 calculates the weight coefficient Wa for each of the learning models Ma (S303). FIG. 23 is a flowchart showing an example of the calculation process S303 of the weight coefficient Wa.

[0113] First, the first calculation unit 160 calculates the value da(j) of the learning model Ma for each of the data xj using the above-described equation (10) (S330). Then, the second calculation unit 161 calculates the total value Σda (a = L, K, T) of the values da(1) to da(10) for each of the three learning models M L ,M k ,M T using the equation (11) (S331).

[0114] Thereafter, the third calculation unit 162 calculates the weight coefficient Wa for the method using the learning model Ma based on the total value Σda in the learning model Ma using the equation (12) (S332). As a result, for each of the three learning models, Wa = (1 / α) × Pa × Σda is calculated.

[0115] Further, the multiplication unit 154 multiplies the weight coefficient Wa and the value Sa(j,k) of the normalized data 133 of the learning model Ma to calculate Wa × Sa(j,k) (S304). The addition unit 155 adds the multiplication results for each explanatory variable (data xj) using the above-described equation (13) to calculate the total value Gj (S305). When the total value Gj is calculated, the selection unit 156 selects a predetermined n (for example, 5) explanatory variables with large total values Gj from the total values Gj (S306: first selection process).

[0116] As a result, in the information processing apparatus 22, among the explanatory variables (data x1 to x10), n explanatory variables with high importance will be selected.

[0117] ===Summary=== As described above, the information processing apparatus 20 of the present embodiment has been described. For example, when selecting important explanatory variables based on a specific single method, the selected explanatory variables are highly dependent on the specific single method, and it may not always be possible to select important explanatory variables with high accuracy. However, the selection unit 56 of the present embodiment selects n explanatory variables based on the values indicating the importance of the explanatory variables obtained based on a plurality of methods. Therefore, compared with the case of selecting important explanatory variables based on a specific single method, important explanatory variables can be selected with high accuracy.

[0118] Also, the normalization unit 51 normalizes the value of the importance data 42 for the method Mk (S21). Thereby, it is possible to easily compare the values indicating the importance obtained by different methods.

[0119] Also, the addition unit 55 adds the multiplication result of the weight coefficient Wk and the value S(j,k) of the normalized data 43 to calculate the total value Ej, and the selection unit 56 selects the explanatory variable based on the total value Ej. The weight coefficient Wk can be freely set based on, for example, the model used by the user. Therefore, in the present embodiment, for example, important explanatory variables can be selected while considering information such as the accuracy of the methods M1 to M3.

[0120] Also, generally, the accuracy of a method that is far from the average of the methods M1 to M3 tends to be low. The weight coefficient Wk of the present embodiment decreases as the deviation from the average of the methods M1 to M3 (that is, the total value Tk (k = 1 to 3)) increases. Therefore, in the present embodiment, since the weight coefficient Wk of the method with low accuracy can be reduced, important explanatory variables can be appropriately selected.

[0121] In addition, the selection unit 52 selects explanatory variables for which all values of the normalization data 43 for the explanatory variables are equal to or greater than a predetermined value τ (for example, 0.01). Therefore, in the process (for example, S23 to S26) when the selection unit 56 finally selects explanatory variables, the number of explanatory variables can be reduced, and thus the amount of calculation can be decreased.

[0122] Also, if L even if the value of S(j, L) calculated in the learning model M L is large, when the accuracy of the learning model M

[0123] is low, it is necessary to reduce the influence of the value of S(j, L). The weight coefficient calculation unit 122 realized in the information processing apparatus 21 of the present embodiment calculates a weight coefficient Wa that increases when the accuracy of the learning model Ma is high and decreases when the accuracy is low. Therefore, by using such a weight coefficient Wa, important explanatory variables can be selected. L , M k , M T In each of, for one explanatory variable x, values indicating nine (m = 9) importance levels may be calculated using three different importance calculation methods. In such a case, the weight coefficient Wa for each learning model may be calculated according to the difference between the average value Ave1(a, j) of the normalization data 133 of one learning model and the average value Ave2(j) of the normalization data 133 of the three learning models for each explanatory variable xj (for example, Equation (9)). Even in such a case where such a weight coefficient Wa is used, important explanatory variables can be selected.

[0124] Also, for example, in the information processing apparatus 22, as shown in Equation (12), the weight coefficient Wa is a value according to the accuracy Pa and the total value Σda, but it is not limited thereto. For example, the total value Σda may be used as the weight coefficient Wa without using the accuracy Pa. Even in such a case, important explanatory variables can be selected.

[0125] In addition, in the information processing apparatus 22, since the weight coefficient Wa is set to a value corresponding to the accuracy Pa and the total value Σda, it is possible to select important explanatory variables while considering the accuracy of the learning model Ma.

[0126] Also, the weight coefficient Wa used in the information processing apparatus 22 increases as the accuracy of the learning model Ma increases, and decreases as the Σda of the learning model Ma decreases (that is, as the difference between the average value Ave1(a,j) and the average value Ave2(j) increases). By using such a weight coefficient Wa, important explanatory variables can be selected with high accuracy.

[0127] In addition, in the information processing apparatus 22, the first calculation unit 160 to the third calculation unit 162 calculate the weight coefficient Wa by executing processes S330 to S332.

[0128] In this embodiment, for example, the power demand of the power system is used as an example for explanation, but it is not limited thereto, and it can be applied to any field as long as it is a technique for selecting the importance of explanatory variables for the target variable.

[0129] Also, in this embodiment, it is assumed that the target variable is continuous data, but it is not limited thereto, and the data y that is the target variable may be data indicating whether there is an abnormality in a predetermined device. In such a case, when the method Mk is executed, a classification model for classifying explanatory variables (for example, data x1 to x10) is used.

[0130] Also, for example, in the information processing apparatuses 21 and 22, it may be configured to include a functional block similar to the selection unit 52 of the information processing apparatus 20. The information processing apparatuses 21 and 22 may include such a selection unit 52 (second selection unit) and execute a process similar to the process S22 in FIG. 9.

[0131] The above embodiments are for facilitating the understanding of the present invention and are not for limiting the interpretation of the present invention. Further, the present invention can be changed and improved without departing from its gist, and it goes without saying that the equivalents of the present invention are included therein.

Explanation of Signs

[0132] 10 Data processing device 20, 21, 22 Information processing device 30 CPU 31 Memory 32, 100, 101 Storage device 33 Input device 34 Display device 35 Communication device 40, 110, 130 Control program 41, 112, 131 Data set 42, 113, 132 Importance data 43, 114, 133 Normalized data 44, 134 Difference data 50, 120, 150 Importance calculation unit 51, 121, 151 Normalization unit 52, 56, 125, 156 Selection unit 53, 122, 153 Weight coefficient calculation unit 54, 123, 154 Multiplication unit 55, 124, 155 Addition unit 56 Selection unit 60 Average value calculation unit 61 Difference calculation unit 62 Total value calculation unit 63 Coefficient calculation unit 111 Learning model 152 Accuracy calculation unit 160 First calculation unit 161 Second calculation unit 162 Third calculation unit

Claims

1. A calculation unit that calculates a first value indicating the importance of each of a plurality of explanatory variables with respect to a predetermined target variable using each of a plurality of m different methods; A normalization unit that calculates a plurality of second values obtained by normalizing each of the plurality of first values; A multiplication unit that multiplies, in each of the m different methods, the plurality of second values and a weight coefficient corresponding to each of the m different methods; An addition unit that adds, in each of the plurality of explanatory variables, the m multiplication results of the plurality of second values and the weight coefficient; A first selection unit that selects n explanatory variables based on an addition result obtained by adding the m multiplication results; comprising, at least one of the m different methods is a method using a learning model in which the relationship between the predetermined target variable and the plurality of explanatory variables is learned; An information processing apparatus characterized by the above.

2. The information processing apparatus according to claim 1, an average value calculation unit that calculates an average value of the m second values in each of the plurality of explanatory variables; a difference calculation unit that calculates a value corresponding to the difference between each of the m second values and the average value in each of the plurality of explanatory variables; a total value calculation unit that calculates a total value by adding the values corresponding to the plurality of differences in each of the m different methods; a coefficient calculation unit that calculates, for each of the m different methods, the weight coefficient that increases and decreases as the total value increases; An information processing apparatus characterized by including the above.

3. The information processing apparatus according to claim 1 or 2, including a second selection unit that selects the explanatory variables for which all of the second values calculated for each of the explanatory variables are equal to or greater than a predetermined value, wherein the multiplication unit multiplies the second value calculated for each of the m different methods for the selected explanatory variable and the weight coefficient. An information processing apparatus characterized by the above.

4. The information processing apparatus according to claim 1, including a weight coefficient calculation unit that calculates the weight coefficient, wherein the m different methods are methods using the m different learning models, and the calculation unit calculates the first value using each of the m different learning models, and the weight coefficient calculation unit calculates the weight coefficient according to the accuracy of each of the m different learning models. An information processing apparatus characterized by the above.

5. ​ A calculation unit that calculates a first value indicating the importance of each of a plurality of explanatory variables with respect to a predetermined target variable, using each of a plurality of m different methods; A normalization unit that calculates a plurality of second values obtained by normalizing each of the plurality of first values; A multiplication unit that multiplies, in each of the m different methods, the plurality of second values by weight coefficients corresponding to each of the m different methods; An addition unit that adds, in each of the plurality of explanatory variables, the m multiplication results of the plurality of second values and the weight coefficients; A first selection unit that selects n explanatory variables based on an addition result obtained by adding the m multiplication results; A weight coefficient calculation unit that calculates the weight coefficients; comprising; The calculation unit calculates the first value using the m different methods based on a plurality of learning models in which the relationship between the predetermined target variable and the plurality of explanatory variables is learned, and a plurality of calculation methods for calculating the importance of the explanatory variables applied to each of the plurality of learning models. The weight coefficient calculation unit calculates the weight coefficient based on a value corresponding to the difference between a first average value of the plurality of second values calculated by one of the plurality of learning models and a second average value of the plurality of second values obtained using the plurality of learning models, in each of the plurality of explanatory variables. An information processing apparatus characterized by the above. **Claim 6** The information processing apparatus according to claim 5, wherein the weight calculation unit a first calculation unit that calculates a value corresponding to the difference; a second calculation unit that adds the plurality of values corresponding to the differences in each of the plurality of learning models to calculate a total value; a third calculation unit that calculates the weight coefficient for each of the plurality of learning models based on the total value; An information processing apparatus characterized by including the above. **Claim 7** The information processing apparatus according to claim 5, including an accuracy calculation unit that calculates the accuracy of each of the plurality of learning models, wherein the weight calculation unit calculates the weight coefficient according to the difference of each of the plurality of explanatory variables and the accuracy of each of the plurality of learning models. An information processing apparatus characterized by the above. **Claim 8** The information processing apparatus according to claim 7, wherein the weight calculation unit calculates the weight coefficient such that the difference of each of the plurality of explanatory variables decreases as the difference increases, and increases as the accuracy of the learning model increases. An information processing apparatus characterized by the above. **Claim 9** The information processing apparatus according to claim 7 or claim 8, wherein the weight calculation unit a first calculation unit that calculates a value corresponding to the difference; a second calculation unit that calculates a total value by adding the values corresponding to the plurality of differences in each of the plurality of learning models; a third calculation unit that calculates the weight coefficient for each of the plurality of learning models based on the total value and the accuracy of the learning model; An information processing apparatus, characterized by including the above.

10. A computer performs a calculation process of calculating a first value indicating the importance of each of a plurality of explanatory variables with respect to a predetermined target variable using each of a plurality of m different methods; a normalization process of calculating a plurality of second values obtained by normalizing each of the plurality of first values; a multiplication process of multiplying, in each of the m different methods, the plurality of second values and the weight coefficients corresponding to the m different methods; an addition process of adding, in each of the plurality of explanatory variables, the m multiplication results of the plurality of second values and the weight coefficients; a first selection process of selecting n explanatory variables based on the addition result obtained by adding the m multiplication results; executes at least one of the m different methods is a method using a learning model in which the relationship between the predetermined target variable and the plurality of explanatory variables is learned. An information processing method.

11. To a computer perform a calculation process of calculating a first value indicating the importance of each of a plurality of explanatory variables with respect to a predetermined target variable using each of a plurality of m different methods; a normalization process of calculating a plurality of second values obtained by normalizing each of the plurality of first values; a multiplication process of multiplying, in each of the m different methods, the plurality of second values and the weight coefficients corresponding to the m different methods; an addition process of adding, in each of the plurality of explanatory variables, the m multiplication results of the plurality of second values and the weight coefficients; a first selection process of selecting n explanatory variables based on the addition result obtained by adding the m multiplication results; causes to execute at least one of the m different methods is a method using a learning model in which the relationship between the predetermined target variable and the plurality of explanatory variables is learned. A program.

Citation Information

Patent Citations

  • Learning device and learning method

    JP2018088080A

  • Analysis device, analysis method, and program

    JP2018151883A

  • Method and system for ontology-based dynamic learning and knowledge integration from measurement data and text

    JP2019507444A

  • Generic learning architecture for robust temporal and domain-based transfer learning

    US20190197550A1

  • Medical interview assistance system

    WO2017183085A1