Regression analysis method, regression analysis system, and regression analysis program

JP2024102632A5Pending Publication Date: 2025-12-22THE UNIV OF TOKYO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023006647
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-01-19
Publication Date
2025-12-22

AI Technical Summary

Technical Problem

Existing regression analysis methods face challenges in constructing a regression model with a correspondence relationship between explanatory and objective variables when the number of training data is small, requiring time and effort to prepare appropriate constraints.

Method used

A regression analysis method that uses a cost function with regularization terms to impose adaptive sign constraints, ensuring the direction of variation in explanatory and objective variables align, allowing for high-accuracy machine learning even with limited data.

Benefits of technology

Enables the construction of a more general-purpose regression model with a robust correspondence between explanatory and objective variable changes, achieving higher prediction performance without the need for extensive parameter tuning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a more versatile technology for constructing a regression model that has a correspondence relationship between a change in an explanatory variable and a change in an objective variable.SOLUTION: In a regression analysis method, a computer reads training data from a storage device that stores training data used as objective and explanatory variables of a regression model, and trains the regression model through machine learning using the training data so as to minimize a cost function that includes a regularization term. The regularization term includes a first term that increases cost in a section where coefficients are positive to be greater than in a section where coefficients are negative, and a second term that increases cost in a section where coefficients are negative to be greater than in a section where coefficients are positive.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a regression analysis method, a regression analysis system, and a regression analysis program. [Background technology]

[0002] Conventionally, when estimating parameters of a regression model using the least squares method, there was a problem that an appropriate least squares estimator could not be obtained, for example, when the number of data samples was small. To address this problem, a method of imposing a constraint condition called the L1 norm was proposed (for example, Non-Patent Document 1). According to LASSO (Least Absolute Shrinkage and Selection Operator), a parameter estimation method that uses the L1 norm as a constraint condition, it is possible to select explanatory variables suitable for explaining the objective variable. The selection and coefficient determination are carried out together.

[0003] Also, a regression analysis device has been proposed that includes a data acquisition unit that reads out training data and constraint conditions from a storage device that stores training data used as the objective variable and explanatory variable of the regression model, and constraint conditions that define in advance whether the explanatory variable should be changed positively or negatively in order to change the objective variable in a positive or negative direction, and a coefficient update unit that repeatedly updates the coefficients of the explanatory variables in the regression model using the training data so as to minimize a cost function that includes a regularization term that increases the cost when the constraint conditions are violated (sign constraint regularization) (for example, Patent Document 1). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] International Publication No. 2021 / 157670 [Non-patent literature]

[0005] [Non-Patent Document 1] Robert Tibshirani, “Regression Shrinkage and Selection via the Lasso”, Journal of the Royal Statistical Society. Series B (Methodological) Vol. 58, No. 1 (1996), pp. 267-288 Summary of the Invention [Problem to be solved by the invention]

[0006] In the past, a regression analysis device was proposed that can perform machine learning with high accuracy even when the training data is relatively small by predefining constraint conditions that define whether the explanatory variables should be changed positively or negatively in order to change the objective variable in a positive or negative direction. However, it is time-consuming to prepare appropriate constraint conditions depending on the target of regression analysis. Therefore, the present technology aims to provide a more versatile technology for constructing a regression model that has a correspondence between the change in the explanatory variables and the change in the objective variable. [Means for solving the problem]

[0007] (Aspect 1) In the regression analysis method, a computer reads out training data from a storage device that stores the training data used as a response variable and explanatory variables of the regression model, and performs machine learning using the regression model using the training data so as to minimize a cost function including a regularization term. The regularization term includes a first term that increases the cost in a range where the coefficient of the explanatory variable is positive more than in a range where the coefficient is negative, and a second term that increases the cost in a range where the coefficient is negative more than in a range where the coefficient is positive.

[0008] (Aspect 2) In the above-mentioned aspect 1, the first term and the second term may include a parameter for making either the first term or the second term zero or approximating it to zero, and in the machine learning, the regression coefficients and the parameter in the regression model may be repeatedly updated.

[0009] (Aspect 3) In the above aspect 2, the parameter may be a binary variable, and the parameter may constitute a factor in a first term that is zero when it takes one value, and a factor in a second term that is zero when it takes the other value.

[0010] (Aspect 4) In the above-mentioned third aspect, the binary variables may be values ​​approximating a growth curve, and regression coefficients and parameters may be calculated so as to minimize the cost function.

[0011] (Aspect 5) In the above-mentioned third or fourth aspect, the binary variable may be a value determined based on the magnitude relationship between a predetermined value expressed using a regression coefficient and the random variable, and the regression coefficient and the parameter may be calculated so as to minimize the cost function.

[0012] The regression analysis method may be applied to a facility in a manufacturing or non-manufacturing industry. The explanatory variables may be sensing data output by a production facility. The objective variables may be the operating conditions of the facility, or a predetermined characteristic value or abnormality level. The contents described in the means for solving the problem may be combined as much as possible without departing from the problem and technical idea of ​​the present disclosure. The above aspects 1 to 5 may be provided as a device such as a computer or a system including multiple devices, a method executed by a computer or a system, or a program for executing the method by a computer or a system. A recording medium for storing the program may be provided. Effect of the Invention

[0013] According to the disclosed technique, it is possible to provide a more versatile technique for constructing a regression model having a correspondence between the variation in an explanatory variable and the variation in a dependent variable. [Brief description of the drawings]

[0014] [Figure 1] FIG. 1 is a diagram showing an example of observed values ​​used to create a regression equation. [Diagram 2] FIG. 2 is a schematic diagram for explaining the factor “R+(w)”. [Diagram 3] FIG. 3 is a schematic diagram for explaining the factor "R-(w)." [Figure 4] FIG. 4 is a diagram for explaining the approximation of the spin variable Sk. [Diagram 5] FIG. 5 is a block diagram showing an example of the configuration of a regression analysis device. [Figure 6] FIG. 6 is a process flow diagram showing an example of the regression analysis process. [Figure 7] FIG. 7 is a diagram showing the verification results. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] Hereinafter, an embodiment of a regression analysis device will be described with reference to the drawings.

[0016] <First embodiment> The regression analysis device according to this embodiment constructs a regression equation (regression model) that represents the relationship between one or more explanatory variables (independent variables) and one objective variable (dependent variable). The process of determining the regression coefficients of the regression equation is called "learning." In this embodiment, the regression equation is constructed so that a constraint (called a "sign constraint") is adaptively imposed such that the direction of variation (positive or negative) of the explanatory variable and the direction of variation (positive or negative) of the objective variable have a predetermined correspondence. In other words, learning is performed so that the sign of the regression coefficient becomes appropriate. Generally, when the number of training data is If there are sufficient numbers of them, the regression equation constructed using them will have the regression coefficients converged so that the direction of variation of the explanatory variables and the direction of variation of the objective variable have a predetermined correspondence relationship. On the other hand, when the number of training data is relatively small, a regression equation may be created in which the correspondence relationship is reversed. The "sign constraint" in this embodiment is imposed so that the regression equation created has an appropriate correspondence relationship even when the number of training data is relatively small.

[0017] FIG. 1 is a diagram showing an example of observed values ​​used to create a regression equation. The table in FIG. 1 shows K inputs x(x 1 ~x K ) and a column of outputs y. The inputs x are used as explanatory variables, and the outputs y are used as response variables. Also, each sample data point t(t 1 ~t T A regression equation is created using T records (training data) from among the multiple records representing the t-th data points, and the regression equation is then applied to the created regression equation. T+1 We input the value of and predict the unknown output y.

[0018] The regression equation is expressed, for example, by the following equation (1).

number

[0019] To determine the regression coefficients and constant terms, the cost function (error function) E(w,s) expressed by the following equation (2) can be used.

number

[0020] Moreover, the regularization term S(w, s) is expressed by the following equation (3).

number

[0021] Figure 2 shows the factor "R + The graph in FIG. 2 is a schematic diagram for explaining "(w)". The vertical axis of the graph is R + (w) and the horizontal axis represents w. The above equation (3) is k Coefficient of w k When is greater than or equal to zero, R + (w)=w, and E(w,s) increases as w increases. On the other hand, the input x k Coefficient of w k When is less than zero, R + (w)=+∞, which causes the cost to diverge to positive infinity. That is, the regularization term in this embodiment is the spin variable Sk If =+1, then input x k Coefficient of w k When is less than zero, the cost diverges to infinity. For the purposes of processing by the regression analysis device, it is sufficient to set it to a sufficiently large value. On the other hand, the input x k Coefficient of w k When is greater than or equal to zero, the cost is increased according to the size of the regression coefficient w. That is, R + The first term containing (w) increases the cost in intervals with negative coefficients more than in intervals with positive coefficients. Here, the coefficient w k When is greater than or equal to zero, the input x of the regression equation shown in equation (1) k As the spin variable S increases, the predicted value y by the regression equation also increases. k If =+1, then input x k When the value of increases, the value of the predicted value y also increases, and the regularization term becomes smaller, and the input x k The cost function is defined such that the regularization term becomes larger when the value of the predicted value y decreases as the value of y increases.

[0022] Figure 3 shows the factor "R - FIG. 3 is a schematic diagram for explaining "(w)". In the graph of FIG. 3, the vertical axis is R - (w) and the horizontal axis represents w. The above equation (3) is k Coefficient of w k When is greater than or equal to zero, R - (w)=+∞, which causes the cost to diverge to positive infinity. For the purposes of regression analysis, this should be a sufficiently large value. On the other hand, the input x k Coefficient of w k When is less than zero, R - (w) = -w, and E(w,s) is increased as w decreases. That is, R - The second term, which contains (w), increases the cost in intervals where the coefficient is positive more than in intervals where the coefficient is negative. Here, the coefficient w k When is less than zero, the input x of the regression equation shown in equation (1) k As the spin variable S increases, the predicted value y by the regression equation decreases. k If =-1, then input x kAs the value of increases, the value of the predicted value y decreases, and the regularization term becomes smaller, and the input x k The cost function is defined such that the regularization term becomes larger when the value of the predicted value y increases as the value of y increases.

[0023] That is, equation (3) is the spin variable S k Depending on the value of , the regression coefficient w k is non-negative or non-positive. And the spin variable S k The optimization of equation (2) is as follows: w={W 0 ,W 1 ,···,W K} and s = {S 1 ,S 2 ,···,S K This is achieved by minimizing the coefficients w and s. The update of the parameters w and so on to minimize E(w,s) can be performed, for example, by the gradient method. In other words, the variables w and so on in a certain step are updated based on the gradient of the cost function E(w,s) for the variables w and so on in the subsequent step, and such processing is repeated until w and so on converge. In this way, the coefficients w and so on can be calculated in advance. k Even if the sign of the spin variable S k The value of is appropriately determined, and the sign constraint is adaptively imposed. According to the above-described regularization term, it is possible to perform regression analysis by imposing a constraint that the direction of change of the explanatory variable and the direction of change of the objective variable have a certain correspondence relationship.

[0024] Spin variable S k is a continuous variable u k (-∞ k <+∞). The spin variable S k can be approximated by the following equation (4).

number

[0025] The error function shown in equation (3) is w and u={u 1 ,u 2 ,···,u k} can be rewritten as follows:

number

number

number

[0026] Based on the steepest gradient method formula, w k and u k Repeat the update of w k and u kThe iteration ends when the value of stops changing (minimization is achieved). At this time, the value of the inverse temperature parameter β is adjusted according to the "simulated annealing" method. In the iterative calculation of the gradient method, β shown in Fig. 4 is initially set to a sufficiently small value, and as the iterations proceed, the value is gradually increased, and finally the spin variable S in equation (4) is k takes on only two values, approximately ±1. The value of β can be increased (annealing schedule) by, for example, multiplying it by a fixed number γ greater than 1 at each iteration step (β←γβ). It is known that if the value of β is increased sufficiently slowly, an optimal solution can be obtained with a probability of 1 without falling into a local solution.

[0027] By doing the above, the coefficient w is subject to an appropriate sign constraint and converges to a value that contributes to the objective variable, and if there is no such value, the coefficient w approaches zero. In other words, if there is no value that satisfies the sign constraint, the penalty effect of regularization will come into play as shown in Figures 2 and 3, pulling back values ​​that violate the sign constraint, ultimately converging to zero. Therefore, like the so-called LASSO, some of the regression coefficients can be estimated to be zero. In this way, the solution to minimizing the error function (5) is s k (k=1,2,...,K) and w k (k=1,2,...,K) is calculated.

[0028] Also, I was asked to k By using these as the regression coefficients in regression equation (1), the new input x K+1 It is possible to predict the output y for the error function (5). k (k=1,2,...,K), we define the sign constraints and use the sign constraint regularization of the conventional technique to obtain w k (k=1,2,...,K). That is, the spin variable S k If >0, we impose the non-negativity constraint shown in Figure 2 and the spin variable S k If <0, we impose the non-positive constraint shown in Figure 3 and the regression coefficient wk Request.

[0029] <Device configuration> FIG. 5 is a block diagram showing an example of the configuration of the regression analysis device 1 that performs the above-mentioned regression analysis. The regression analysis device 1 is a computer, and includes a communication interface (I / F) 11, a storage device 12, an input / output device 13, and a processor 14, and the above components are connected via a bus or the like. The communication I / F 11 may be, for example, a network card or a communication module, and communicates with other computers based on a predetermined protocol. The storage device 12 may be a main storage device such as a random access memory (RAM) or a read only memory (ROM), and an auxiliary storage device (secondary storage device) such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory. The main storage device temporarily stores a program read by the processor 14 and information processed by the program. The auxiliary storage device stores a program executed by the processor 14 and information processed by the program. In this embodiment, the storage device 12 temporarily or permanently stores training data and information representing constraint conditions. The input / output device 13 may be, for example, a keyboard, a pointer ... The processor 14 is a user interface including an input device such as a display device, an output device such as a monitor, and an input / output device such as a touch panel. The processor executes a program to implement each of the present embodiment. 1, functional blocks are shown in the processor 14. That is, the processor 14 executes a predetermined program to function as a data acquisition unit 141, a regression analysis unit 142, a verification processing unit 143, and an operation processing unit 144.

[0030] The data acquisition unit 141 acquires information representing training data and constraint conditions from the storage device 12. The regression analysis unit 142 updates the regression coefficients as a solution to minimize the error function described above, and judges whether the updated regression coefficient values ​​have converged. If it is judged that the regression coefficients have not converged, the regression analysis unit 142 repeats updating the coefficients. If it is judged that the regression coefficients have converged, for example, the regression analysis unit 142 stores the finally generated coefficients in the storage device 12. Furthermore, the verification processing unit 143 evaluates the regression equation created based on a predetermined evaluation index. For example, the verification processing unit 143 calculates the coefficient of determination R 2 The operation processing unit 144 calculates a predicted value using the created regression equation and, for example, a newly acquired observation value (for example, the K+1th data in FIG. 1). The operation processing unit 144 may also calculate a predicted value when conditions are changed using the created regression equation and an arbitrary value. Here, the arbitrary value may be, for example, a value input by a user via the communication I / F 11 or the input / output device 13. For example, when the operating conditions of the plant are changed, it becomes possible to predict the process state after the change.

[0031] <Regression analysis processing> FIG. 6 is a process flow diagram showing an example of regression analysis processing executed by the regression analysis device. The data acquisition unit 141 of the regression analysis device 1 reads out training data and information representing constraint conditions from the storage device 12 (FIG. 6: S1). In this step, for example, values ​​of input x and output y as shown in FIG. 1 are read out as training data. It is assumed that the input x is treated as an explanatory variable and the output y is treated as a target variable. Also, in FIG. 1, a positive or negative code registered in association with the input x is read out as information representing the constraint conditions. The regression analysis device 1 uses the read out code as the above-mentioned constraint code. It is to be noted that in this embodiment, a regression equation as shown in Equation (1) is used.

[0032] Furthermore, the regression analysis unit 142 of the regression analysis device 1 updates the regression coefficients under the above-mentioned sign constraint (FIG. 6: S2). In this step, the regression analysis unit 142 updates the coefficient w and the parameter s (specifically, for example, the parameter "u" in equation (5)) so as to minimize the cost function E(w, s) shown in equation (2), for example. Specifically, the regression analysis unit 142 can update the coefficient w, etc. based on the equation for the steepest gradient method described above.

[0033] The regularization term of the cost function E(w,s) according to this embodiment is defined so that the cost increases when the adaptively imposed constraint condition is not satisfied. That is, the regularization term decreases the value of the cost function E(w,s) when the direction of variation of the explanatory variable and the direction of variation of the objective variable have a plausible correspondence. Furthermore, the regression analysis unit 142 sets the coefficient w to zero when the coefficient does not converge to a value that satisfies the constraint condition. In this manner, the explanatory variables may be selected.

[0034] Furthermore, the regression analysis unit 142 of the regression analysis device 1 determines whether all the coefficients w, etc. have converged or been set to zero (FIG. 6: S3). In this step, the regression analysis unit 142 determines that convergence has occurred when the gradient of the updated coefficients w, etc., is sufficiently close to zero. Specifically, the regression analysis unit 142 determines that convergence has occurred when the values ​​of the coefficient w and parameter u in the above-mentioned steepest gradient method equation no longer change.

[0035] If it is determined that the coefficients w etc. have not converged or been set to zero (S3: NO), the process returns to S2 and is repeated. On the other hand, if it is determined that the coefficients w etc. have converged or been set to zero (S3: YES), the regression analysis unit 142 stores the regression equation in the storage device 12 (FIG. 6: S4). In this step, the regression analysis unit 142 stores the updated coefficients w etc. in the storage device 12.

[0036] Furthermore, the verification processing unit 143 of the regression analysis device 1 may verify the accuracy of the created regression equation (FIG. 6: S5). In this step, the verification processing unit 143 verifies the accuracy of the regression equation using test data, for example, by cross-validation. The verification processing unit 143 can perform verification based on a predetermined evaluation index such as a correlation coefficient or a coefficient of determination. Note that this step may be omitted.

[0037] In addition, the operation processing unit 144 of the regression analysis device 1 performs operation processing using the created regression equation (FIG. 6: S6). In this step, the operation processing unit 144 performs operation processing using the data number t T+1 A predicted value of the output y for a new input x is calculated, as in the record of step S2. Note that this step may be performed by a device (not shown) other than the regression analysis device 1, using the regression equation stored in S4.

[0038] <Example> The regression analysis process according to this technology was applied to sensing data obtained from a chemical process plant. The number of training data, N, was set to 14, and 14 pairs of inputs x and outputs y were prepared. Each pair contains 12 inputs x and one output y. The constructed regression equation was evaluated using leave-one-out cross validation (LOOCV). In other words, N-1 of the N pieces of data were used as training data, and the remaining one was used as evaluation data. The validation was repeated N times so that all 14 pairs of data were used as evaluation data once. Then, the coefficient of determination (R 2 ) is calculated. 2 ) is an index of prediction accuracy and is given by the following equation (8):

number

[0039] This embodiment is applied to various values ​​of α, and the coefficient of determination R 2was calculated. In addition, a similar verification was performed using L1 regularization (Lasso) and the least squares method. Fig. 7 is a diagram showing the verification results. The solid line indicates the coefficient of determination R 2 The dashed line shows the coefficient of determination R of the regression equation created by L1 regularization. 2 The dashed horizontal line indicates the coefficient of determination R 2 In the case of this embodiment, when α is 0.0001 or less, a high R of 0.6 or more is obtained. 2 In the case of L1 regularization, when α is 0.0001, the R value is the same as that of this embodiment. 2 However, when α is larger or smaller than this, the R 2 The value drops sharply to near zero. From the above, it can be seen that this embodiment can stably guarantee higher prediction performance (generalization performance) than L1 regularization. Furthermore, this embodiment and L1 regularization obtained higher prediction performance than the least squares method.

[0040] As shown in Figure 7, the prediction performance with L1 regularization is maximized when α is a specific value (around 0.00011 in the example in Figure 7). When the value of α becomes larger or smaller than that, This means that the predictive performance when using L1 regularization is sensitive to the value of α. Therefore, in order to obtain the effect of L1 regularization, it is essential to tune the parameter α by cross-validation or other methods.

[0041] On the other hand, the prediction performance when this embodiment is used is maximized when α is a specific value (near 0.00011 in the example of FIG. 7), but even if the value of α is made smaller than this, it does not decrease as much as in the case of L1 regularization. Therefore, even if the value of α is set to zero a priori without undergoing cross-validation, sufficiently high prediction performance can be guaranteed. Although cross-validation is a procedure that requires a large amount of calculation, according to this embodiment, it can be said that sufficient prediction performance can be obtained without undergoing cross-validation by setting the value of α to zero from the beginning.

[0042] <Second embodiment> In this embodiment, the spin variable s k is treated as a random variable rather than being approximated by a continuous variable. That is, instead of equation (4) in the first embodiment, the spin variable s k Strictly speaking, x can only take the value of +1 or -1, and which value it will take is determined probabilistically. The following describes the differences from the first embodiment.

[0043] In learning by the steepest descent method, the quantity given by the following equation (9) is defined.

number

[0044] Also, the spin variable s k Update (k=1,2,...,K). (i) Generate a uniform random number r in the range [0,1]. (ii) If ρ ≧ r, s k =1; (iii) else if ρ <r, s k =-1. First, generate a uniform random number r in the range of 0 to 1. Second, if the value of ρ defined in equation (9) is equal to or greater than r, the spin variable s k Third, when the value of ρ defined in equation (9) is less than r, the spin variable s k Let be -1.

[0045] In addition, the error function shown in equation (2) is minimized (learned) in the following procedure. The regression coefficient w k The minimization of can be done by the gradient descent method. At each step, the regression coefficient w k is updated by the following equation (10), which is the same as the usual steepest descent method.

number

[0046] As in the first embodiment, the value of the inverse temperature parameter β is adjusted according to the “simulated annealing method”.

[0047] The second embodiment as described above also does not require any constraint conditions to be defined in advance, and can provide a more versatile technique for constructing a regression model having a corresponding relationship between the variation in the explanatory variables and the variation in the target variable.

[0048] <Modification> The configuration of the computer shown in FIG. 5 is an example, and is not limited to such an example. For example, at least a part of the functions of the regression analysis device 1 may be distributed to multiple devices to provide a regression analysis system that realizes the above-mentioned processing as a whole. Also, for at least a part of the functions of the regression analysis device 1, a regression analysis system in which multiple devices execute the same functions in parallel may be provided. Also, at least a part of the functions of the regression analysis device 1 may be provided on a so-called cloud. Also, the regression analysis device 1 may not have some components, such as the verification processing unit 143. Also, a system that holds the created trained model and has only the operation processing unit 144 may be provided.

[0049] In addition, the cost function shown in formula (2) performs L1 regularization on either the positive or negative side, but it can also work with L2 norm or other convex functions. That is, instead of the sum of the absolute values ​​of the coefficients, a sum of squares of the coefficients on either the positive or negative side or other terms that impose a penalty may be used.

[0050] The parameter S included in the regularization term k is not restricted to spin variables, but is included in the regularization term. The first term, (1 / 2)(1+s k )R + (w k ) and the second term "(1 / 2)(1-s k )R - (w k )” to zero or to approximate zero. For example, the parameter S k can be a binary variable that takes either 0 or +1. In this case, R in Equation (3) ASL (w k ,s k ) can be defined as follows:

number

[0051] The content of the data analyzed by the regression analysis device 1 is not particularly limited. In addition to prediction of characteristic values ​​such as quality in the manufacturing industry that produces chemicals and the like as described in the embodiment, the regression analysis device 1 may be applied to process control such as processing, diagnosis, etc., and may be applied to non-manufacturing industries such as power generation and wastewater treatment, and various other fields. For example, it is possible to predict, as objective variables, predetermined characteristic values ​​representing mechanical, physical, or chemical properties, preferable operating conditions depending on the situation, abnormality levels, etc., using information including sensing data output by production equipment as explanatory variables. In non-manufacturing industries, it is also possible to predict preferable operating conditions of equipment, etc., using various information as explanatory variables.

[0052] The present disclosure also includes a method and a computer program for executing the above-mentioned processes, and a computer-readable recording medium having the program recorded thereon. The recording medium having the program recorded thereon enables the above-mentioned processes by causing a computer to execute the program.

[0053] Here, a computer-readable recording medium refers to a recording medium that stores information such as data and programs through electrical, magnetic, optical, mechanical, or chemical action and can be read by a computer. Among such recording media, those that can be removed from a computer include flexible disks, magneto-optical disks, optical disks, magnetic tapes, memory cards, etc. Furthermore, recording media that are fixed to a computer include HDDs, SSDs (Solid State Drives), ROMs, etc.

[0054] Each configuration and their combinations in each embodiment are merely examples, and addition, omission, substitution, and other modifications of the configurations are possible as appropriate within the scope of the present disclosure. The present disclosure is not limited by the embodiments, but is limited only by the scope of the claims. In addition, each aspect disclosed in this specification can be combined with any other feature disclosed in this specification. [Explanation of symbols]

[0055] 1: Regression analysis device 11: Communication I / F 12: Storage device 13: Input / output devices 14: Processor 141: Data Acquisition Section 142: Regression analysis section 143: Verification processing section 144: Operation Processing Unit

Claims

1. The computer reading out training data from a storage device that stores the training data to be used as a response variable and an explanatory variable of the regression model; Machine learning is performed using the regression model using the training data so as to minimize a cost function including a regularization term.

1. A regression analysis method comprising: The regularization term includes one term that increases the cost in an interval where the coefficient of the explanatory variable is positive more than in an interval where the coefficient is negative, and another term that increases the cost in an interval where the coefficient is negative more than in an interval where the coefficient is positive. Regression analysis methods.

2. the one term and the other term include a parameter for making either the one term or the other term zero or approximating it to zero, In the machine learning, the regression coefficients in the regression model and the parameters are repeatedly updated. The regression analysis method according to claim 1 .

3. The parameter is a binary variable, and when the parameter takes one value, it constitutes a factor that is zero in one term, and when the parameter takes the other value, it constitutes a factor that is zero in the other term. The regression analysis method according to claim 2 .

4. The binary variables are used as values ​​that approximate a growth curve, and the regression coefficients and parameters are calculated so as to minimize the cost function. The regression analysis method according to claim 3 .

5. The binary variables are values ​​determined based on the magnitude relationship between a predetermined value expressed using the regression coefficients and a random variable, and the regression coefficients and the parameters are used to minimize the cost function. The meter and The regression analysis method according to claim 3 .

6. A regression analysis system comprising one or more computers that execute the regression analysis method according to any one of claims 1 to 5.

7. A regression analysis program for causing one or more computers to execute the regression analysis method according to any one of claims 1 to 5.