Learning device and learning method

By finding constant terms and potential vectors in the prediction model, optimizing the objective function to select a subset of features, and generating a prediction model containing indicator functions, solving the problem of difficulty in reducing the prediction model in the existing technology, and achieving the scale reduction and efficiency improvement of the model.

JP2025074433APending Publication Date: 2025-05-14NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023185221
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-05-14

AI Technical Summary

Technical Problem

In the prior art, prediction models are usually large and difficult to narrow and optimize when large computers are not in time.

Method used

By finding the constant terms and latent vectors in a given prediction model, the inner product of the latent vector is calculated to obtain parameters, and selecting a subset of features by optimizing the objective function, thereby generating a prediction model containing the indicator function.

Benefits of technology

The scale reduction of the predictive model is achieved, and the portability and efficiency of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025074433000001_ABST
    Figure 2025074433000001_ABST
Patent Text Reader

Abstract

To provide a learning device with which it is possible to reduce the scale of a prediction model.SOLUTION: Learning means performs a first round of learning in a prescribed prediction model including a parameter of squared cross term of a feature amount and a constant term in order to determine the parameter and the constant term. Feature amount selection means optimizes an objective function including the parameter when it is assumed that the number of feature amounts included in a subset of feature amounts is L and the number of subsets is m, so as to create m subsets of feature amounts including L feature amounts and performs a first round of feature amount selection for selecting a prescribed number of feature amounts. The learning means performs a second round of learning similar to the first round of learning by using feature amounts included in each subset so as to determine the constant term and the parameter. The feature amount selection means performs a second round of feature amount selection similar to the first round of feature amount selection after the second round of learning.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a learning device, a learning method, and a learning program. [Background technology]

[0002] Non-Patent Document 1 describes a field-aware factorization machine (FFM). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Yuchin Juan et al., “Field-aware Factorization Machines for CTR Prediction”, [Retrieved September 6, 2023], Internet<URL : https: / / www.csie.ntu.edu.tw / ~cjlin / papers / ffm.pdf> Summary of the Invention [Problem to be solved by the invention]

[0004] Generally, it is not always possible to use a large-scale computer. Therefore, it is preferable that the prediction model (learning model) used in the computer is small-scale.

[0005] Therefore, an object of the present invention is to provide a learning device, a learning method, and a learning program that can reduce the scale of a prediction model. [Means for solving the problem]

[0006] The learning device according to the present disclosure is characterized in that, when determining a parameter and a constant term in a predetermined prediction model including a parameter of a second-order cross term of a feature and a constant term, the learning device is provided with: a learning means for performing a first learning to determine the parameter by determining the constant term and a latent vector based on training data and determining an inner product of the latent vector; and a feature selection means for performing a first feature selection to select a predetermined number of feature terms by optimizing an objective function including the parameter, where the number of features included in the feature subset is L and the number of subsets is m. The learning means performs a second learning similar to the first learning using the features included in each subset to determine the constant term and the parameter, and the feature selection means performs a second feature selection similar to the first feature selection after the second learning, and generates a prediction model including an indicator function based on the constant term obtained by the second learning and the feature selected by the second feature selection.

[0007] The learning method according to the present disclosure is characterized in that, when a computer determines a parameter and a constant term in a predetermined prediction model including a parameter of a second-order cross term of a feature and a constant term, the computer performs a first learning round to determine the parameter by determining the constant term and a latent vector based on training data and determining an inner product of the latent vector, creates m feature subsets including L features by optimizing an objective function including the parameter when the number of features included in the feature subset is L and the number of subsets is m, performs a first feature selection round to select a predetermined number of features, performs a second learning round similar to the first learning round using the features included in each subset to determine the constant term and the parameter, performs a second feature selection round similar to the first feature selection round after the second learning round, and generates a prediction model including an indicator function based on the constant term obtained by the second learning round and the feature selected by the second feature selection round.

[0008] A learning program according to the present disclosure causes a computer to function as a learning device comprising: learning means for performing a first learning to determine a parameter and a constant term in a predetermined prediction model including a parameter of a second-order cross term of a feature and a constant term, by determining the constant term and a latent vector based on training data and determining an inner product of the latent vector; and feature selection means for performing a first feature selection to select a predetermined number of feature terms, where the number of features included in the feature subset is L and the number of subsets is m, by optimizing an objective function including the parameter, creating m feature subsets including L features and performing a first feature selection, in which the learning means performs a second learning similar to the first learning using the features included in each subset to determine the constant term and the parameter, and the feature selection means performs a second feature selection similar to the first feature selection after the second learning, and generates a prediction model including an indicator function based on the constant term obtained by the second learning and the feature selected by the second feature selection. Effect of the Invention

[0009] According to the present disclosure, the prediction model can be made smaller in scale. [Brief description of the drawings]

[0010] [Figure 1] 1 is a schematic diagram showing a processing procedure of a learning device according to the present disclosure. [Diagram 2] 1 is a block diagram showing an example configuration of a learning device according to the present disclosure. [Diagram 3] FIG. 13 is a schematic diagram showing the behavior of loss when L is changed for each of the cases of m=1 to m=5. [Figure 4] This is a scatter plot of true values ​​and predicted values. [Diagram 5] Scatter plot of true and predicted values ​​for test data. [Figure 6] FIG. 11 is an explanatory diagram showing the change over time in the coefficient of determination obtained in Example 3. [Figure 7]FIG. 13 is an explanatory diagram showing the coefficient of determination when the number of epochs in the first learning is set to 20, 50, 75, 100, and 300. [Figure 8] FIG. 2 is a schematic block diagram showing an example of the configuration of a computer related to a learning device. [Figure 9] 1 is a block diagram showing an overview of a learning device according to the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Hereinafter, an embodiment of a learning device according to the present disclosure will be described. Note that in this specification, for convenience, a hat symbol may be written so as to be shifted from a variable.

[0012] The learning device according to the present disclosure generates a prediction model after repeating learning and feature selection twice for a secondary learning model. As the learning model, an FFM learning model is used. FFM is an extension of FM (Factorization Machine).

[0013] In FM, features are expanded in latent vectors to calculate cross terms of features. However, in FM, the latent vector of feature A is the same for both feature B and feature C, which may reduce the descriptive power of the learning model.

[0014] On the other hand, in FFM, features are treated as fields and each field is assigned a latent vector.

[0015] For example, suppose there are three features: "gender," "purchase history of product A," and "purchase history of product B." In FM, the latent vector of "gender" for "purchase history of product A" is the same as the latent vector of "gender" for "purchase history of product B." However, in FFM, the latent vector of "gender" for "purchase history of product A" is different from the latent vector of "gender" for "purchase history of product B," which increases the descriptive power of the learning model.

[0016] Furthermore, FM and FFM are explained.

[0017] For example, the objective variable is y, and there are three explanatory variables x1, x 2, The linear regression model with x3 is expressed as follows:

[0018] y = w1x1+ w2x2+ w3x3

[0019] In addition, regression models are not limited to the above examples, and regression models that include products of explanatory variables such as x1x2 are also used. Products of explanatory variables such as x1x2 are called cross terms. An example of such a regression model is shown below.

[0020] y = w1x1+ w2x2+ w3x3+ w 12 x1x2+ w 23 x2x3+ w 13 x1x3

[0021] In a regression model that includes cross terms, the descriptive power of the regression model increases as the number of parameters increases. However, a large number of cross terms can lead to problems such as an increase in the computational load or a lack of data corresponding to the cross terms. In FM, to solve these problems, the cross term parameters w l1l2 is approximated by the inner product of the latent vector v of K terms. This approximation is expressed by the following formula.

[0022]

number

[0023] If the number of explanatory variables is N, the number of cross terms is N(N-1) / 2, and the parameter required for the cross terms (w l1l2 The number of cross terms is also N(N-1) / 2. In FM, N(N-1) / 2 cross terms are divided into NK parameters (v l m ), where the parameter v l1 m is determined independently of the crossover pair (explanatory variable l2), which may reduce the descriptive power of the regression model.

[0024] Therefore, in FFM, we define a quantity called a field that combines several explanatory variables, and create a latent vector (v l1f(l2) m ) is set. f(l2) refers to the field to which the explanatory variable l2 belongs. In this case, the cross-term parameter w l1l2 The approximation is expressed by the following formula:

[0025]

number

[0026] For a more specific explanation, a data set showing the relationship between 10 features such as age, sex, BMI (Body Mass Index), and blood pressure of a diabetic patient and the progress of the disease after one year is used as an example. Then, all 442 pieces of data are randomly divided so that the ratio of training data to test data is 3:1. Such data is used in the examples described below. However, the items included in the data set, the number of pieces of data, and the ratio of training data to test data are not limited to the above example.

[0027] In optimization using the Ising model, implementation becomes easier if quantitative variables are divided into categories and treated as categorical variables.

[0028] For example, an ordinal scale is introduced for nine of the ten features (quantitative variables) excluding gender, and each feature (quantitative variable) is converted into a categorical variable consisting of four groups. Specifically, each quantitative variable is arranged in ascending order and divided into four groups by quartiles. Gender is a categorical variable from the beginning. In the embodiment described later, the quantitative variables are converted into categorical variables in this manner. However, one quantitative variable may be divided into a group other than four groups. Since the quantitative variables are made into categorical variables, the number of fields remains the original ten. Gender is divided into two groups, and the other nine features are divided into four groups. Therefore, the number of features is effectively 38. The number of cross terms of these features is 703.

[0029] Predicted value y for the i-th data when taking into account second-order cross terms i ^ is expressed by the following equation (1).

[0030]

number

[0031] Equation (1) is the original prediction model to be scaled down. In equation (1), q il is a 0 / 1 binary variable for the l-th feature of the i-th data, given by the input data. Similarly, q il1 is a 0 / 1 binary variable for the l1-th feature of the i-th data, given by the input data. q il2 is a 0 / 1 binary variable for the l2-th feature of the i-th data, given by the input data.

[0032] In addition, w0 (constant term) and w l ,w l1l2 are parameters determined by FFM learning based on the training data. More specifically, in FFM, w0,w l ,The latent vector v of the K-term expansion is obtained, and the quadratic parameter w is calculated as the inner product of the latent vector. l1l2 Find the second-order parameter w l1l2 is expressed as an approximation of the following equation (2) as the inner product of the latent vector v.

[0033]

number

[0034] Parameter w0,w l , latent vector v lf k is optimized by applying stochastic gradient descent to the residual sum of squares in equation (1). In this example, the parameters w0,w l , latent vector v lfk The total number of epochs for learning to obtain σ is assumed to be 300 unless otherwise specified. In the embodiment described later, the number of epochs is also set to 300. However, the number of epochs is not limited to 300.

[0035] As above, w0,w l ,w l1l2 is required, but w l is not used in the subsequent processing.

[0036] Let the number of features (number of elements) in the feature subset be L, and the number of subsets be m. l ,w l1l2 After calculating the feature vectors, we select m subsets S containing L features from the N features (38 in this example). g Select features by creating (g=1,…,m)

[0037] We also consider the cross terms that can be generated within each subset. This is equivalent to selecting two features from L features, so for each subset, the number of cross terms is L C2 = L(L-1) / 2. In this case, in order to treat the m subsets independently, a constraint is imposed that the number of subsets to which each selected feature belongs is at most one. A prediction model using cross terms generated from the sum of such subsets is given by the following equation (3).

[0038]

number

[0039] In this disclosure, we focus on the cross terms, so the regression model in equation (3) includes a constant term and cross terms, but no linear terms.

[0040] In equation (3), w (0) ,w p1p2 (2) are w0 and w in equation (1), respectively. l1l2 Also, q in equation (3) is the same as kp1 ,qkp2 are the q il1 ,q il2 is the same as:

[0041] In formula (3), S' is a set of cross terms that can be generated from each of the above-mentioned subsets, summed over m subsets. Cross terms that span two subsets are not included in S'.

[0042] In addition, χ in formula (3) s’ ([i,j]) is the indicator function, and if the (i,j) element of the cross term is included in the set S', then χ s’ ([i,j]) is 1, otherwise, χ s’ ([i,j]) is 0. That is, χ s’ ([i,j]) is expressed by the following equation (4).

[0043]

number

[0044] The cross terms with large absolute values ​​of the parameters are important. To select them, we use the parameter w l1l2 The objective function of optimization using the Ising model is defined as the following equation (5) so that the sum of the absolute values ​​of is maximized.

[0045]

number

[0046] In equation (5), w ij is the w in equation (1). l1l2 In addition, in equation (5), a minus sign is added to the first term on the right-hand side to handle the minimization problem.

[0047] In formula (5), x ig The i-th feature is a subset S g is a binary variable that is 1 if x belongs to the igis the variable to be optimized by the Ising model. The indicator function χ s’ ([i,j]) and variable x ig The relationship between them is expressed by the following equation (6).

[0048]

number

[0049] The first term on the right-hand side of equation (5) is the parameter w ij The second term on the right hand side of equation (5) corresponds to the constraint that the number of subsets to which the i-th feature belongs is at most one. The third term on the right hand side of equation (5) is a number that represents the constraint that the number of features included in each subset is L. The values ​​of L and m are specified externally.

[0050] By optimizing the objective function in (5), x in (5) ig This means that m subsets, each of which has L elements, are determined.

[0051] For example, let L = 14 and m = 2. Also, let us say that 38 features are assigned labels (identification information) from 0 to 37. By optimizing the objective function of equation (5), for example, the following two subsets are determined. S1= {1, 5, 6, 7, 8, 9, 10, 11, 13, 15, 23, 24, 27, 28} S2= {4, 12, 19, 22, 25, 29, 30, 31, 32, 33, 34, 35, 36, 37}

[0052] In other words, 28 features are selected. Note that S1 and S2 above are merely examples of the selection results.

[0053] FIG. 1 is a schematic diagram showing a processing procedure of a learning device according to the present disclosure. First, w0, w l ,wl1l2 (Process (a)). As mentioned above, based on the training data, w0,w l ,The latent vector v is obtained, and the quadratic parameter w is obtained by calculating the inner product of the latent vectors. l1l2 The operation of this process (a) corresponds to the first learning.

[0054] In equation (5), w ij is the w obtained in step (a) l1l2 After step (a), the first feature selection is performed (step (b)). In step (b), m feature subsets, each with L elements, are determined by optimizing the objective function of equation (5). The features included in each subset are the selected features.

[0055] After step (b), a second learning process similar to the first learning process is performed using the features selected in step (b) (step (c)). As a result, w0,w l ,w l1l2 is required.

[0056] After the process (c), a second feature selection similar to the first feature selection is performed (process (d)). That is, by optimizing the objective function of formula (5), m feature subsets each having L elements are determined. At this time, x ig is required.

[0057] Note that different values ​​may be specified as L in the first feature selection and L in the second feature selection, provided that L in the first feature selection is equal to or greater than L in the second feature selection.

[0058] After step (d), a prediction model expressed in the form of equation (3) is generated (step (e)). (0) is the w0 obtained in the second learning (process (c)). Also, x obtained in process (d) ig From equation (6), the indicator function χ s’ ([i,j]) and its indicator function χ s’A prediction model is generated by substituting ([i,j]) into equation (3).

[0059] 2 is a block diagram showing an example of the configuration of a learning device according to the present disclosure. The learning device according to the present disclosure includes a first input unit 1, a second input unit 2, a learning unit 3, a feature selection unit 4, and a model generation unit 5.

[0060] The first input unit 1 is a response variable y i (For example, a variable that indicates the progression of the disease after one year), explanatory variables (features) q ij , and an input interface for inputting the number N of feature quantities. The feature quantities are assumed to be categorical variables before being input to the first input unit 1. Alternatively, a feature quantity that has not been categoricalized may be input to the first input unit 1, and the first input unit 1 may have a function of categorizing the feature quantity. N is the number of feature quantities after being categoricalized.

[0061] The second input unit 2 is a function of the number m of feature subsets, the number L of features included in each subset in the first feature selection. (1) , and the number of features included in each subset in the second feature selection, L (2) This is an input interface for inputting L (1) , L (2) That's it. Here, L (1) =L (2) The case where the value is L will be described as an example.

[0062] When the learning unit 3 determines the parameters and constant terms in a predetermined prediction model (equation (1)) including the parameters of the second-order cross terms of the features and a constant term, the learning unit 3 performs a first learning to determine the parameters by determining the constant term and latent vector based on the training data and determining the inner product of the latent vector.

[0063] More specifically, the learning unit 3 minimizes the loss function by applying the stochastic gradient descent method to the residual sum of squares in equation (1). l, and the latent vector v are calculated. Then, by calculating the inner product of the latent vector v using equation (2), the cross-term parameter w l1l2 Request.

[0064] The feature selection unit 4 optimizes an objective function (equation (5)) including the parameters, thereby creating m feature subsets including L features, and performs a first feature selection in which a predetermined number of features are selected. At this time, a constraint is imposed that the number of subsets to which each selected feature belongs is at most one. This constraint is also applied to the second feature selection.

[0065] Specifically, the feature selection unit 4 optimizes the objective function of the formula (5) to obtain x ig As mentioned above, this means determining m subsets, each of which has L elements.

[0066] After the first feature selection, the learning unit 3 performs a second learning process similar to the first learning process, using the features included in each subset.

[0067] After the second learning, the feature selection unit 4 performs a second feature selection similar to the first feature selection.

[0068] The model generation unit 5 generates a prediction model (equation (3)) including an indicator function, based on the constant term obtained by the second learning and the feature quantity selected by the second feature quantity selection.

[0069] In equation (3), w (0) is a constant term obtained in the second learning. The model generation unit 5 then uses the x ig From equation (6), the indicator function χ s’ ([i,j]) and its indicator function χ s’ A prediction model is generated by substituting ([i,j]) into equation (3).

[0070] The learning unit 3, the feature selection unit 4, and the model generation unit 5 are realized, for example, by a CPU (Central Processing Unit) of a computer that operates according to a learning program. In this case, the CPU reads the learning program from a program recording medium such as a program storage device of the computer, and operates as the learning unit 3, the feature selection unit 4, and the model generation unit 5 according to the learning program.

[0071] The prediction model generated by the model generation unit 5 includes an indicator function. Therefore, it can be said that the prediction model is small-scaled. Therefore, according to the learning device according to the present disclosure, it is possible to reduce the size of the prediction model.

[0072] As described above, L in the first feature selection is equal to or greater than L in the second feature selection. That is, the number of features included in the subset in the first feature selection may be equal to the number of features included in the subset in the second feature selection, or may be greater than the number of features included in the subset in the second feature selection. When the number of features included in the subset in the first feature selection is greater than the number of features included in the subset in the second feature selection, the number of features included in the subset in the second feature selection may be, for example, 14 to 28% greater than the number of features included in the subset in the second feature selection. However, this value is merely an example and is not limited to this value.

[0073] In addition, the number of epochs in the first learning may be smaller than the number of epochs in the second learning. For example, the number of epochs in the first learning may be 1 / 3 of the number of epochs in the second learning, or may be smaller than 1 / 3 of the number of epochs in the second learning.

[0074] In the above embodiment, the case where FFM learning is performed has been described as an example, but the learning device according to the present disclosure can also be applied to the case where FM learning is performed.

[0075] Examples are shown below. In each example, an equation without the second term on the right-hand side of equation (1) was used instead of equation (1).

[0076] [Example 1] To confirm the extent to which the original learning model can be reproduced by the sum of cross terms generated from the subsets, we investigated the relationship between (L,m) and the loss function. Here, the sum of squared residuals between the true value and the predicted value for the training data is defined as the loss. Figure 3 is a schematic diagram showing the behavior of the loss when L is changed for each case of m=1 to m=5. The dashed horizontal line is the value from a model that takes into account all non-zero matrix elements (i.e., the original FFM learning model). The maximum value of L that can be taken for each m is L max (m)=floor(N / m). Note that floor() represents the floor function. As L gets larger, the loss decreases. If L is the same, the loss decreases as m gets larger. However, there is an upper limit for L, L. max (m), and when m is increased, L max (m) decreases. When m is increased, the total number of matrix elements of cross terms generated from the subsets decreases because it is asymptotically expressed as N(N-1) / (2m). For example, when (L,m)=(7,5), only about 15% of the original matrix elements are considered, so a model that sufficiently reduces loss cannot be constructed. Basically, the maximum value of L that can be taken within the constraints of hardware, etc. is set as L. max H Then, m H =floor(N / L max H ) is the best way to do this. Therefore, in the following, (L max H ,m H Let us focus on the case where θ = (14,2).

[0077] In the first embodiment, a prediction model was generated after the first learning and the first feature selection.

[0078] Figure 4 shows a scatter plot of true values ​​and predicted values. Figure 4(A) shows a scatter plot of true values ​​and predicted values ​​for training data. Figure 4(B) shows a scatter plot of true values ​​and predicted values ​​for test data. Figures 4(A) and 4(B) show the results of the simulation for (y k ,y k S’ ^) and (yk ,y k A ^) represents the k S’ ^ is the predicted value for the kth data using only the constant term and the cross term generated from the subset, and is calculated using equation (3). On the other hand, y k A ^ is calculated using the original FFM model (in the embodiment, the equation without the second term on the right-hand side of equation (1)) when all non-zero cross terms before subset selection are considered. The 45° line corresponds to a perfect prediction. In comparison, in both the training data and the test data, the predicted value y^ tends to be underestimated in the region where the objective variable y (degree of progression of disease after one year) is large. In particular, y k S’ ^, this tendency is even stronger. In addition, the coefficient of determination (0.398) of the model of equation (3) for the test data is slightly lower than that of the original FFM model (0.428). Therefore, simply generating a predictive model after the first learning and the first feature selection is insufficient to reproduce the original model. The spin coupling matrix (weight matrix of cross terms) is a parameter determined by learning, so it can be considered as a hyperparameter in optimization using the Ising model.

[0079] [Example 2] In Example 2, a predictive model (referred to as predictive model A) was generated after the first learning and the first feature selection, and a predictive model (referred to as predictive model B) was also generated after the second learning and the second feature selection.

[0080] Figure 5 is a scatter plot of true values ​​and predicted values ​​for the test data. The circle markers show the relationship between true values ​​and predicted values ​​when the original FFM model is used. The triangle markers show the relationship between true values ​​and predicted values ​​when prediction model A is used. The diamond markers show the relationship between true values ​​and predicted values ​​when prediction model B is used. Figure 5 shows that when prediction model B generated after the second learning and second feature selection is used, the tendency for the prediction value to be underestimated is improved in areas where the value of the objective variable y is large. In addition, when prediction model B is used, the value of the coefficient of determination is improved from 0.398 mentioned above to 0.420.

[0081] In the following Examples 3 and 4, a prediction model was generated after the first learning, the first feature selection, the second learning, and the second feature selection.

[0082] [Example 3] In the third embodiment, the second feature selection is set to (L,m)=(14,2) because of the constraints of hardware such as a quantum computer. However, the first feature selection does not need to be limited to (L,m)=(14,2). The number of features in the embodiment is 38. Therefore, if two subsets are created and the number of subsets to which each selected feature belongs is at most one, the number of elements of each subset can be set to a maximum of 19. Therefore, the transition of the coefficient of determination of the prediction model generated after the second feature selection was obtained when the value of L was changed from 14 to 19 in the first feature selection. In the third embodiment, the transition of the coefficient of determination for the training data and the transition of the coefficient of determination for the test data were obtained. FIG. 6 is an explanatory diagram showing the transition of the coefficient of determination obtained in the third embodiment. The coefficient of determination for the test data can be considered as an index of generalization performance. When L in the first feature selection is 18, the coefficient of determination for the test data is a maximum value of 0.464. Furthermore, when L in the first feature selection is between 16 and 18, the coefficient of determination for the test data is higher than the coefficient of determination (0.428) of the original FFM model. Therefore, in this example, it is preferable to set L in the first feature selection to any value between 16 and 18, and L in the second feature selection to 14. In this case, the number of features included in the subset in the first feature selection is 14 to 28% higher than the number of features included in the subset in the second feature selection.

[0083] [Example 4] The first feature selection is a process for selecting features that are effective for the second learning, so it does not necessarily need to be calculated with high accuracy. Therefore, the number of epochs when calculating the parameter w0 and latent vector in the first learning does not need to be the number of epochs that will allow the loss to converge sufficiently. In other words, early termination may be performed when calculating the parameter w0 and latent vector in the first learning. Then, the parameter w of the cross term is calculated by the inner product of the latent vector. l1l2Then, the first feature selection may be performed, followed by the second learning. In the second learning, the number of epochs for obtaining the parameter w0 and the latent vector is set to a number of epochs that allows the loss to converge sufficiently (300 in this example).

[0084] In Example 4, the coefficient of determination for the test data was confirmed when the number of epochs in the first learning was less than the number of epochs in the second learning. In this case, in the first feature selection, (L, m) = (18, 2). In the second learning, the number of epochs when obtaining the parameter w0 and the latent vector was set to 300, and in the second feature selection, (L, m) = (14, 2). In this case, in the first learning, the number of epochs when obtaining the parameter w0 and the latent vector was set to 20, 50, 75, 100, and 300, and the coefficient of determination in each case was examined. FIG. 7 is an explanatory diagram showing the coefficient of determination when the number of epochs in the first learning is set to 20, 50, 75, 100, and 300. When the number of epochs in the first learning is set to 100, the same coefficient of determination is obtained as when the number of epochs is set to 300. Even when the number of epochs in the first learning is set to 20, the loss does not converge sufficiently, but the coefficient of determination of the final result does not deteriorate significantly. Thus, the number of epochs in the first learning may be 1 / 3 of the number of epochs in the second learning, or may be less than 1 / 3 of the number of epochs in the second learning.

[0085] 8 is a schematic block diagram showing an example of the configuration of a computer related to a learning device. The computer 2000 includes, for example, a CPU 2001, a main memory device 2002, an auxiliary memory device 2003, an interface 2004, and an input interface 2005.

[0086] The learning device according to the present disclosure is realized, for example, by a computer 2000. The operation of the learning device is stored in the form of a program (learning program) in an auxiliary storage device 2003. A CPU 2001 reads the program from the auxiliary storage device 2003, loads the program in a main storage device 2002, and executes the processing described in the above embodiment according to the program.

[0087] The auxiliary storage device 2003 is an example of a non-transient tangible medium. Other examples of non-transient tangible media include a magnetic disk, a magneto-optical disk, a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), a semiconductor memory, and the like, which are connected via an interface 2004.

[0088] Next, an overview of the learning device according to the present disclosure will be described. Fig. 9 is a block diagram showing an overview of the learning device according to the present disclosure. The learning device includes a learning unit 13, a feature quantity selection unit 14, and a model generation unit 15.

[0089] When the learning means 13 (e.g., the learning unit 3) determines a parameter and a constant term in a predetermined prediction model including a parameter of a second-order cross term of a feature and a constant term, the learning means 13 performs a first learning to determine the parameter by determining the constant term and a latent vector based on training data and determining the inner product of the latent vector.

[0090] The feature selection means 14 (e.g., feature selection unit 4) creates m feature subsets, each including L features, by optimizing an objective function including its parameters, where the number of features included in the feature subset is L and the number of subsets is m, and performs a first feature selection in which a predetermined number of features are selected.

[0091] The learning means 13 performs a second learning process similar to the first learning process using the feature amounts included in each subset, thereby obtaining constant terms and parameters.

[0092] After the second learning, the feature selection means 14 performs a second feature selection similar to the first feature selection.

[0093] The model generation means 15 (for example, the model generation unit 5) generates a prediction model including an indicator function based on the constant term obtained by the second learning and the feature quantity selected by the second feature quantity selection.

[0094] Such a configuration allows the prediction model to be scaled down.

[0095] The above-described embodiment of the present invention can be described as follows, but is not limited to the following.

[0096] (Appendix 1) a learning means for performing a first learning for determining a parameter of a quadratic cross term of a feature quantity and a constant term in a predetermined prediction model including the parameter and the constant term, by determining the constant term and a latent vector based on training data and determining an inner product of the latent vector; a feature selection means for optimizing an objective function including the parameter, where the number of feature quantities included in the feature subset is L and the number of the subsets is m, thereby creating m feature subsets including the L feature quantities and performing a first feature selection of selecting a predetermined number of feature quantities; The learning means includes: performing a second learning round similar to the first learning round using the feature amounts included in each subset to obtain the constant terms and the parameters; The feature selection means After the second learning, a second feature selection is performed similar to the first feature selection, A model generating means is provided for generating a prediction model including an indicator function based on the constant term obtained by the second learning and the feature quantity selected by the second feature quantity selection. A learning device characterized by:

[0097] (Appendix 2) The number of features included in the subset in the first feature selection is greater than the number of features included in the subset in the second feature selection. 2. A learning device as described in appendix 1.

[0098] (Appendix 3) In the first feature selection and the second feature selection, the number of subsets to which each selected feature belongs is at most one. 3. A learning device according to claim 1 or 2.

[0099] (Appendix 4) The number of epochs in the first learning round is less than the number of epochs in the second learning round. 3. A learning device according to claim 1 or 2.

[0100] (Appendix 5) The computer When calculating a parameter and a constant term in a predetermined prediction model including a parameter of a second-order cross term of a feature and a constant term, the constant term and a latent vector are calculated based on training data, and an inner product of the latent vector is calculated, thereby performing a first learning for calculating the parameter; a first feature selection process for selecting a predetermined number of features by optimizing an objective function including the parameter, where the number of features included in the feature subset is L and the number of the subsets is m; performing a second learning round similar to the first learning round using the feature amounts included in each subset to obtain the constant terms and the parameters; After the second learning, a second feature selection is performed similar to the first feature selection, A predictive model including an indicator function is generated based on the constant term obtained by the second learning and the feature quantity selected by the second feature quantity selection. A learning method comprising:

[0101] (Appendix 6) The number of features included in the subset in the first feature selection is greater than the number of features included in the subset in the second feature selection. The learning method described in Appendix 5.

[0102] (Appendix 7) In the first feature selection and the second feature selection, the number of subsets to which each selected feature belongs is at most one. The study method described in Appendix 5 or Appendix 6.

[0103] (Appendix 8) The number of epochs in the first learning round is less than the number of epochs in the second learning round. The study method described in Appendix 5 or Appendix 6.

[0104] (Appendix 9) Computer, a learning means for performing a first learning for determining a parameter of a quadratic cross term of a feature quantity and a constant term in a predetermined prediction model including the parameter and the constant term, by determining the constant term and a latent vector based on training data and determining an inner product of the latent vector; a feature selection means for optimizing an objective function including the parameter, where the number of feature quantities included in the feature subset is L and the number of the subsets is m, thereby creating m feature subsets including the L feature quantities and performing a first feature selection of selecting a predetermined number of feature quantities; The learning means includes: performing a second learning round similar to the first learning round using the feature amounts included in each subset to obtain the constant terms and the parameters; The feature selection means After the second learning, a second feature selection is performed similar to the first feature selection, a learning device including a model generating means for generating a predictive model including an indicator function based on a constant term obtained by the second learning and a feature quantity selected by the second feature quantity selection, A learning program to help you function as a

[0105] In addition, some or all of the configurations described in Supplementary Notes 2 to 4 that are dependent on Supplementary Note 1 may also be dependent on Supplementary Note 9 in the same dependent relationship as Supplementary Notes 2 to 4. Furthermore, not limited to Supplementary Notes 1, 5, and 9, various hardware, software, various recording means for recording software, or systems may also be made to be dependent on some or all of the configurations described as supplementary notes, within the scope of the above-mentioned embodiment.

[0106] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-mentioned embodiments. Various modifications that can be understood by a person skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be appropriately combined with other embodiments. [Industrial Applicability]

[0107] The present invention can be suitably applied to a learning device that generates a predictive model. [Explanation of symbols]

[0108] 1. First input section 2 Second input section 3. Learning Department 4. Feature Selection Section 5. Model Generation

Claims

1. a learning means for performing a first learning for determining a parameter and a constant term in a predetermined prediction model including a parameter of a second-order cross term of a feature and a constant term, by determining the constant term and a latent vector based on training data and determining an inner product of the latent vector; a feature selection means for optimizing an objective function including the parameter, where the number of feature quantities included in the feature subset is L and the number of the subsets is m, thereby creating m feature subsets each including the L feature quantities, and performing a first feature selection of selecting a predetermined number of feature quantities; The learning means includes: performing a second learning process similar to the first learning process using the features included in each subset to obtain the constant term and the parameters; The feature selection means After the second learning, a second feature selection is performed similarly to the first feature selection; A model generating means for generating a prediction model including an indicator function based on the constant term obtained by the second learning and the feature quantity selected by the second feature quantity selection is provided. A learning device characterized by:

2. The number of features included in the subset in the first feature selection is greater than the number of features included in the subset in the second feature selection. The learning device according to claim 1 .

3. In the first feature selection and the second feature selection, the number of subsets to which each selected feature belongs is at most one. The learning device according to claim 1 or 2.

4. The number of epochs in the first learning is less than the number of epochs in the second learning. The learning device according to claim 1 or 2.

5. The computer When calculating a parameter and a constant term in a predetermined prediction model including a parameter of a second-order cross term of a feature and a constant term, the constant term and a latent vector are calculated based on training data, and an inner product of the latent vector is calculated, thereby performing a first learning for calculating the parameter; a first feature selection process is performed to select a predetermined number of features from the feature subsets, the first feature selection process being performed by optimizing an objective function including the parameter, the first feature selection being performed by optimizing an objective function including the parameter, the first feature selection being performed by optimizing an objective function including the parameter, the number of feature subsets being m, the number of feature subsets being m, the number of feature subsets being m, the number of feature subsets being m, performing a second learning process similar to the first learning process using the features included in each subset to obtain the constant term and the parameters; After the second learning, a second feature selection is performed similarly to the first feature selection; A prediction model including an indicator function is generated based on the constant term obtained by the second learning and the feature quantity selected by the second feature quantity selection. A learning method comprising:

6. Computer, a learning means for performing a first learning for determining a parameter and a constant term in a predetermined prediction model including a parameter of a second-order cross term of a feature and a constant term, by determining the constant term and a latent vector based on training data and determining an inner product of the latent vector; a feature selection means for optimizing an objective function including the parameter, where the number of feature quantities included in the feature subset is L and the number of the subsets is m, thereby creating m feature subsets each including the L feature quantities, and performing a first feature selection of selecting a predetermined number of feature quantities; The learning means includes: performing a second learning process similar to the first learning process using the features included in each subset to obtain the constant term and the parameters; The feature selection means After the second learning, a second feature selection is performed similarly to the first feature selection; a learning device including a model generating means for generating a predictive model including an indicator function based on a constant term obtained by the second learning and a feature quantity selected by the second feature quantity selection; A learning program to help you function as a