Regression Analysis Device, Regression Analysis Method, and Recording Medium

By introducing regularization terms in the regression model, ensuring that the coefficients comply with predefined constraints, the problem of difficult to infer regression model parameters when the number of data samples is small, the corresponding relationship between variable changes and target variable changes is realized, and the calculation cost is reduced.

CN115053216BActive Publication Date: 2025-05-30THE UNIV OF TOKYO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180012527.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-04
Filing Date
2021-02-04
Publication Date
2025-05-30
Estimated Expiration
2041-02-04

AI Technical Summary

Technical Problem

When the prior art estimates the parameters of the regression model by the least squares method, it is difficult to find the least squares estimation quantity when the number of data samples is small, and the calculation cost is high when the combination of the variables is changed to perform repeated simulations.

Method used

By introducing regularization terms into the regression model, we ensure that the coefficients of the variables in the model minimize the cost function without violating predefined constraints, so as to build a regression model that has a corresponding relationship between the variable changes and the target variable changes.

Benefits of technology

It is realized that the direction of change of the variable required for the target variable to change in the positive or negative direction without selecting a coefficient that violates the constraints is not selected, which reduces the calculation cost, and a regression model with clear correspondence is constructed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115053216B_ABST
    Figure CN115053216B_ABST
Patent Text Reader

Abstract

The present disclosure constructs a regression model that illustrates a correspondence relationship between changes in explanatory variables and changes in a target variable. The regression analysis device includes: a data acquisition unit that reads out training data and constraint conditions from a storage device that stores the training data used as the target variable and explanatory variables of the regression model, and constraint conditions that are predefined to indicate whether the explanatory variable should change in the positive or negative direction in order to change the target variable in the positive or negative direction; and a coefficient update unit that repeatedly updates the coefficients of the explanatory variables in the regression model using the training data in such a way as to minimize a cost function including a regularization term, where the regularization term increases the cost in the case of violating the constraint conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a regression analysis device, a regression analysis method, and a program. Background Art

[0002] Conventionally, when estimating parameters of a regression model by the least squares method, there has been the following problem: for example, when the number of data samples is small, the least squares estimator cannot be obtained. Therefore, a method of imposing a constraint condition called the L1 norm has been proposed (for example, Non-Patent Document 1). According to LASSO (Least Absolute Shrinkage and Selection Operator), which is a parameter estimation method using the L1 norm as a constraint condition, selection of explanatory variables applicable to explaining a target variable and determination of coefficients are performed together.

[0003] In addition, regarding LASSO, various improved methods have been proposed, such as pre-grouping or clustering highly correlated explanatory variables.

[0004] Prior Art Documents

[0005] Non-Patent Documents

[0006] Non-Patent Document 1: Robert Tibshirani, “Regression Shrinkage and Selection via the Lasso”, Journal of the Royal Statistical Society. Series B (Methodological) Vol. 58, No. 1 (1996), pp. 267 - 288 Summary of the Invention

[0007] Problems to be Solved by the Invention

[0008] Conventionally, for example, in a case of performing control so as to obtain a desired result, even when solving an inverse problem using a prediction model, an appropriate result is sometimes not obtained. That is, it is not clear how to change the values of explanatory variables in order to make an estimated value based on the prediction model close to a desired value. However, a calculation cost is incurred in a method of repeatedly performing simulation by changing combinations of explanatory variables. Therefore, an object of the present technology is to construct a regression model having a correspondence relationship between changes in explanatory variables and changes in a target variable.

[0009] Technical Solution

[0010] The regression analysis device includes: a data acquisition unit that reads training data and constraint conditions from a storage device that stores the training data for the target variable and explanatory variables used as a regression model, and the constraint conditions predefined to determine whether the explanatory variable should change in the positive or negative direction to cause the target variable to change in the positive or negative direction; and a coefficient update unit that repeatedly updates the coefficients of the explanatory variables in the regression model using the training data in such a way as to minimize the cost function including a regularization term, where the regularization term increases the cost when the constraint conditions are violated.

[0011] Through the above regularization term, a regression model can be created as follows: Without selecting coefficients that violate the constraint conditions, it is known that in order to cause the target variable to change in the positive or negative direction, it is only necessary to change the explanatory variable in either the positive or negative direction. That is, a regression model having a correspondence relationship between the change in the explanatory variable and the change in the target variable can be constructed.

[0012] In addition, it may be that the regularization term increases the cost according to the sum of the absolute values of the coefficients in the positive or negative interval corresponding to the constraint conditions. For example, in one side where the coefficient is positive or negative, a regression model using L1 regularization can be constructed. In addition, it may be that the regularization term increases the cost according to the sum of the absolute values of the coefficients in one of the positive or negative intervals corresponding to the constraint conditions, and sets the cost to infinity in the other.

[0013] In addition, it may be that the coefficient update unit sets the coefficient to 0 when the coefficient does not converge to a value that satisfies the constraint conditions. By doing so, the explanatory variables that do not contribute to the target variable under the above constraint conditions can be deleted from the regression model, achieving sparse modeling.

[0014] In addition, it may be that the coefficient update unit updates the coefficient by the proximal gradient method. By doing so, the non-differentiable points of the regularization term in the convergence calculation are avoided. Therefore, the time required for convergence can be shortened.

[0015] It should be noted that the content described in the technical solution can be combined as much as possible without departing from the subject matter and technical idea of the present disclosure. In addition, the content of the technical solution can be provided as a computer or other device, or a system including multiple devices, a method executed by a computer, or a program that causes a computer to execute. It should be noted that a recording medium storing the program can also be provided.

[0016] Advantages of the Invention

[0017] According to the disclosed technology, a regression model having a correspondence relationship between the change in the explanatory variable and the change in the target variable can be constructed. Brief Description of the Drawings

[0018] Figure 1 It is a diagram showing an example of training data for creating a regression formula.

[0019] Figure 2A It is a schematic diagram for explaining the constraints imposed on the regression coefficients.

[0020] Figure 2B It is a schematic diagram for explaining the constraints imposed on the regression coefficients.

[0021] Figure 3 It is a diagram for explaining the update of parameter w.

[0022] Figure 4 It is a diagram for explaining the update of parameter η.

[0023] Figure 5 It is a block diagram showing an example of the configuration of the regression analysis device 1 that performs the above regression analysis.

[0024] Figure 6 It is a processing flow chart showing an example of the regression analysis process executed by the regression analysis device.

[0025] Figure 7A It is a diagram showing the relationship between the parameter α representing the strength of the constraint and the correlation coefficient r.

[0026] Figure 7B It is a diagram showing the relationship between the parameter α representing the strength of the constraint and the correlation coefficient r.

[0027] Figure 8 It is a diagram showing the relationship between the parameter α representing the strength of the constraint and the coefficient of determination E.

[0028] Figure 9 It is a diagram showing the relationship between the number of data T for learning and the correlation coefficient r.

[0029] Figure 10 It is a diagram showing the relationship between the number of data T for learning and the coefficient of determination E.

[0030] Figure 11A It is a schematic diagram for explaining the constraints imposed on the regression coefficients.

[0031] Figure 11B It is a schematic diagram for explaining the constraints imposed on the regression coefficients.

[0032] Figure 12 It is a diagram showing the relationship between parameter β and the correlation coefficient r.

[0033] Figure 13 It is a diagram showing the relationship between parameter β and the coefficient of determination R 2 .

[0034] Figure 14 It is a graph showing the relationship between parameter β and RMSE. Detailed implementation mode

[0035] Hereinafter, with reference to the drawings, an implementation mode of the regression analysis device will be described.

[0036] <Implementation mode>

[0037] The regression analysis device of the present implementation mode constructs a regression formula (regression model) representing the relationship between one or more explanatory variables (independent variables) and one target variable (dependent variable). At this time, in at least any one of the explanatory variables, a constraint (referred to as "sign constraint") is imposed such that the positive or negative direction of the change of the explanatory variable has a certain correspondence with the positive or negative direction of the change of the target variable, and the regression formula is made.

[0038] Figure 1 It is a graph showing an example of the observed values (training data) for the production of the regression formula. Figure 1 The table of... includes columns of K kinds of inputs x (x 1 ~x K ) and a column of output y. The input x corresponds to the explanatory variable, and the output y corresponds to the target variable. In addition, T records out of a plurality of records representing data points t (t 1 ~t T ,...) of each training data are used to produce the regression formula. In addition, at least a part of the K kinds of inputs x is made to correspond to a positive or negative sign (information indicating the constraint condition of the present implementation mode, referred to as "constraint sign"). The constraint sign corresponding to each input x is information for predefining in the constructed regression formula that in order to make the output y change in the positive direction, it is sufficient to make the input x change in either the positive or negative direction.

[0039] The regression formula is represented by the following formula (1), for example.

[0040] [Mathematical formula 1]

[0041]

[0042] Note that w k is the regression coefficient, and w 0 is the constant term. In addition, w k is determined according to the pre-determined constraint sign.

[0043] For the determination of the regression coefficient and the constant term, a cost function represented by the following formula (2) can be used. The regression formula is determined by selecting the coefficient w k so as to minimize the cost function E(w).

[0044] [Formula 2]

[0045]

[0046] where

[0047]

[0048] αR is a regularization term (penalty term), and its coefficient α is a parameter representing the strength of the constraint. In the table of Figure 1 , when the constraint sign of x k is positive, take the value of R + (w), and when the constraint sign is negative, take the value of R - (w). In this way, the regularization term αR of the present embodiment imposes a sign constraint determined based on L1-type regularization unidirectionally on the positive or negative side. That is, the regularization term increases the cost according to the sum of the absolute values of the coefficients in the interval where the coefficient w k is either positive or negative corresponding to the constraint sign.

[0049] Figure 2A and Figure 2B are schematic diagrams for explaining the constraints imposed on one regression coefficient w. In the chart of Figure 2A , the vertical axis represents R + (w), and the horizontal axis represents w. In addition, the arrow schematically indicates that in the interval where w is negative, the regularization term is defined in such a way that the larger the value of α, the further the value of R + (w) increases. The above formula (2) corresponds to the case where the constraint sign established with the input x k is positive. When the coefficient w k of the input x k is 0 or more, R + (w) = 0 and E(w) does not increase. On the other hand, when the coefficient w k of the input x k is less than 0, R + (w) = -w and E(w) increases. Here, when the coefficient w k is 0 or more, the larger the input x k of the regression equation shown in formula (1), the larger the predicted value μ based on the regression equation. That is, in the case where the constraint sign established with x k is positive, the cost function is defined in such a way that the regularization term becomes smaller when the value of the input x k increases and the predicted value μ also increases, and the regularization term becomes larger when the value of the input x k increases and the predicted value μ decreases.

[0050] In Figure 2B , the vertical axis represents R -(w), the horizontal axis represents w. In addition, the arrow schematically indicates that in the interval where w is positive, the larger the value of α, the further R increases - (w) defines the regularization term in a way that the value of k is negative for the constraint sign of the input x k When the coefficient w k of the input x is 0 or more, R - (w) = w and E(w) is increased. On the other hand, when the coefficient w k of the input x k is less than 0, R - (w) = 0 and E(w) is not increased. Here, when the coefficient w k is less than 0, as the input x k shown in the regression equation of formula (1) increases, the predicted value μ based on the regression equation decreases. That is, in the case where the constraint sign corresponding to the input x k is negative, the regularization term becomes smaller when the predicted value μ decreases as the value of the input x k increases, and the regularization term becomes larger when the predicted value μ also increases as the value of the input x k increases. The cost function is defined in this way.

[0051] According to the above regularization term, regression analysis is performed by imposing a constraint such that there is a certain correspondence between the positive or negative direction of the change in the explanatory variable and the positive or negative direction of the change in the target variable.

[0052] In addition, the partial derivative of the variable w of the cost function E(w) is represented by the following formula (3).

[0053] [Formula 3]

[0054]

[0055] where

[0056]

[0057] For example, the update of the parameter w to minimize E(w) can also be performed by the gradient method using the following formula (4).

[0058] [Formula 4]

[0059]

[0060] Figure 3 is a diagram for explaining the update of the parameter w. Based on the gradient of the variable w of the cost function E(w) at a certain step s, the variable w at the subsequent step s + 1 is updated, and such processing is repeated until w converges.

[0061] However, as shown in Equation (3), when the constraint symbol corresponding to the input x k is any one, differentiation cannot be performed with w = 0. For example, the value corresponding to the constraint symbol can also be calculated for each input x k and the sum thereof can be used as a regularization term to perform regression based on the steepest descent method. However, the calculation becomes unstable. Therefore, for example, the proximal gradient method can also be used. In the proximal gradient method, for example, w that minimizes the above Equation (2) is obtained. If the sum of squared errors in Equation (2) is set to f(x) in advance and the regularization term is set to g(w) in advance, the update formula for w is represented by the following Equation (5).

[0062] [Equation 5]

[0063]

[0064] where

[0065]

[0066] η is the step size that determines the magnitude of the coefficient w updated in one step (one iteration). (W(t)) is the gradient. The update is repeated until the gradient is sufficiently close to 0. When the gradient is sufficiently close to 0, it is determined that convergence has occurred and the update ends.

[0067] More specifically, the update formula for w is represented by the following Equation (6).

[0068] [Equation 6]

[0069]

[0070] When the constraint symbol is positive, the calculation can be performed as shown in the following Equation (7).

[0071] [Equation 7]

[0072]

[0073] When the constraint symbol is negative, the calculation can be performed as shown in the following Equation (8).

[0074] [Equation 8]

[0075]

[0076] Through the above processing, the coefficient w can be determined. The coefficient w satisfies the sign constraint and converges to a value that contributes to the target variable. If there is no such value, the coefficient w approaches 0. That is, when there is no value that satisfies the sign constraint, as Figure 2A and Figure 2BAs shown, the penalty effect based on regularization takes effect, pulling back the values that violate the sign constraint, and as a result, it converges to 0. Therefore, a part of the regression coefficients can be estimated to be 0 in the same way as the so-called LASSO.

[0077] It should be noted that the value of η can also be appropriately updated in each step repeated in the process of updating the coefficients. Figure 4 An example of a schematic code for searching for an appropriate η is shown. For example, as Figure 4 shown, the process like that is executed for η in each step 0 which is a predetermined initial value. β is a positive value less than 1, for example, and is updated in a way that reduces η. In this way, by adjusting η which is the step size for updating the coefficient w, the coefficient w can be appropriately converged.

[0078] <Device Configuration>

[0079] Figure 5 It is a block diagram showing an example of the configuration of a regression analysis device 1 that performs the above regression analysis. The regression analysis device 1 is a general computer, and includes a communication interface (I / F) 11, a storage device 12, an input / output device 13, and a processor 14. The communication I / F 11 can be, for example, a network card or a communication module, and communicates with other computers based on a prescribed protocol. The storage device 12 can be a main storage device such as a RAM (Random Access Memory) and a ROM (Read Only Memory), and an auxiliary storage device (secondary storage device) such as an HDD (Hard-Disk Drive), an SSD (Solid State Drive), and a flash memory. The main storage device temporarily stores the programs read by the processor 14 and the information processed by the programs. The auxiliary storage device stores the programs executed by the processor 14, the information processed by the programs, etc. In the present embodiment, training data and information representing constraint conditions are stored temporarily or permanently in the storage device 12. The input / output device 13 is, for example, a user interface such as an input device like a keyboard and a mouse, an output device like a monitor, and an input / output device like a touch panel. The processor 14 is an arithmetic processing device such as a CPU (Central Processing Unit), and performs each process of the present embodiment by executing programs. In Figure 1 the example of, functional blocks are shown inside the processor 14. That is, the processor 14 functions as a data acquisition unit 141, a coefficient update unit 142, a convergence determination unit 143, a verification processing unit 144, and an application processing unit 145 by executing a prescribed program.

[0080] The data acquisition unit 141 acquires training data and information representing constraint conditions from the storage device 12. The coefficient update unit 142 updates the coefficients of the regression equation under the above-mentioned constraint conditions. In addition, the convergence determination unit 143 determines whether the values of the updated coefficients converge. Note that, in the case where it is determined that convergence has not occurred, the coefficient update unit 142 repeatedly updates the coefficients. In the case where it is determined that convergence has occurred, for example, the coefficient update unit 142 stores the finally generated coefficients in the storage device 12. In addition, the verification processing unit 144 evaluates the produced regression equation based on a prescribed evaluation index. The application processing unit 145 calculates a predicted value using the produced regression equation and, for example, newly acquired observed values. In addition, the application processing unit 145 may also calculate a predicted value in the case where conditions are changed using the produced regression equation and arbitrary values. Here, the arbitrary values may be, for example, values input by the user via the communication I / F 11 or the input / output device 13. Regarding the regression equation produced in the present embodiment, it is explained that there is a certain correspondence between the direction of change of the variable and the direction of change of the target variable. Therefore, for example, the user can easily presume whether it is sufficient to increase the input value or to decrease the input value in order to make the predicted value close to the desired value. Therefore, for example, in the case of performing a certain control based on the presumed value, the regression equation of the present embodiment is effective.

[0081] The above-described components are connected via the bus 15.

[0082] <Regression analysis processing>

[0083] Figure 6 is a processing flowchart showing an example of the regression analysis processing executed by the regression analysis device. The data acquisition unit 141 of the regression analysis device 1 reads out training data and information representing constraint conditions from the storage device 12 ( Figure 6 : S11). In this step, for example, values of input x and output y as shown Figure 1 are read out as training data. Note that the input x is treated as an explanatory variable, and the output y is treated as a target variable. In addition, in Figure 1 , the positive or negative sign registered in correspondence with the input x is read out as information representing constraint conditions. The regression analysis device 1 uses the read-out sign as the above-mentioned constraint sign. Note that, in the present embodiment, a regression equation as shown in Equation (1) is used.

[0084] In addition, the coefficient update unit 142 of the regression analysis device 1 updates the regression coefficients under the above-mentioned sign constraint ( Figure 6 : S12). In this step, the coefficient update unit 142, for example, as in Figure 3As shown by the arrow on the upper side, the coefficient w is updated in such a way as to minimize the cost function E(w) shown in Equation (2). Specifically, the coefficient update unit 142 can update the coefficient w based on Equations (6) to (8).

[0085] The regularization term of the cost function E(w) in the present embodiment is defined to increase the cost when the constraint condition obtained in S11 is not satisfied. That is, the regularization term reduces the value of the cost function E(w) when there is a predetermined correspondence relationship between the positive or negative direction of the change in the explanatory variable and the positive or negative direction of the change in the target variable. In addition, when the coefficient does not converge to a value that satisfies the constraint condition, the coefficient update unit 143 sets the coefficient to 0.

[0086] In addition, the convergence determination unit 143 of the regression analysis device 1 determines whether the coefficient w converges or whether the coefficient w is set to 0( Figure 6 : S13). In this step, the convergence determination unit 143 determines convergence when the gradient of the updated coefficient w is sufficiently close to 0. Specifically, the convergence determination unit 143 determines convergence when the value of the coefficient w does not change in Equation (7) or Equation (8).

[0087] When it is determined that the coefficient w does not converge and is not set to 0 (S13: No (NO)), the process returns to S12 and is repeated. On the other hand, when it is determined that the coefficient w converges or is set to 0 (S13: Yes (YES)), the convergence determination unit 143 stores the regression equation in the storage device 12( Figure 6 : S14). In this step, the convergence determination unit 143 stores the updated coefficient w in the storage device 12.

[0088] In addition, the verification processing unit 144 of the regression analysis device 1 may verify the accuracy of the produced regression equation( Figure 6 : S20). In this step, the verification processing unit 144 verifies the accuracy of the regression equation using test data, for example, by cross-validation. In addition, the verification processing unit 144 can perform verification based on a specified evaluation index such as a correlation coefficient and a specified determination coefficient. It should be noted that, as described later, this step may be omitted.

[0089] Then, the operation processing unit 145 of the regression analysis device 1 performs operation processing using the produced regression equation( Figure 6 : S30). In this step, the operation processing unit 145, for example, calculates the predicted value of the output y with respect to the new input x in the form of a record with the data number t Figure 1 shown. It should be noted that this step uses the regression equation stored in S14 and may also be performed by a device (not shown) other than the regression analysis device 1. T+1 ​

[0090] <Example>

[0091] Using the sensing data obtained from the production equipment to construct a regression equation and evaluate the accuracy. For each of the Figure 1 input and output shown, the output values of different sensors are used. In addition, regarding the sensing data continuously output from the sensor, the number T of the most recent data is set as the learning interval. In addition, the constraint symbols are preset based on the knowledge related to the production equipment.

[0092] The correlation coefficient r used as an evaluation index is obtained by the following formula (9).

[0093] [Formula 9]

[0094]

[0095] where

[0096] is the observed value

[0097]

[0098] That is, the numerator of formula (9) is the covariance between the predicted value μ and the measured value y of the training data. The denominator of formula (9) is the product of the standard deviation of the predicted value μ and the standard deviation of the measured value y of the training data.

[0099] In addition, the coefficient of determination E used as another evaluation index is obtained by the following formula (10).

[0100] [Formula 10]

[0101]

[0102] The coefficient of determination E is a value representing the magnitude of the distribution of the predicted value relative to the distribution of the observed value. In the case where the distribution of the observed value coincides with the distribution of the predicted value through standardization, E = 1. In addition, in the case where the distribution of the predicted value is narrower than the distribution of the observed value, E < 1. And, in the case where the distribution of the predicted value is wider than the distribution of the observed value, E > 1.

[0103] Figure 7A and Figure 7B are graphs showing the relationship between the parameter α representing the strength of the constraint and the correlation coefficient r for models constructed by multiple methods. Figure 8 is a graph showing the relationship between the parameter α representing the strength of the constraint and the coefficient of determination E for models constructed by multiple methods. Figure 7A and Figure 7B The horizontal axis of the curve graphs of and represents the parameter α, and the vertical axis represents the correlation coefficient r. Figure 7A and Figure 7B have different scales on the horizontal axis. In addition,Figure 8 The horizontal axis of the curve graph represents α, and the vertical axis represents the determination coefficient E. The solid line represents the method disclosed in the embodiment, the dashed line represents the comparative example in which a part of the symbol constraints of the embodiment is randomly selected and the positive and negative are reversed, the single-dot dashed line represents L1 regularization (LASSO), and the double-dot dashed line represents the results without regularization. It should be noted that in each method, the number of data T is set to 40 to construct the model. In addition, as described above, the constraint symbols are preset based on the knowledge related to the production equipment. However, generally speaking, inappropriate settings can be included. It can be said that the comparative example simulates the symbol constraints with errors.

[0104] As Figure 7A and Figure 7B shown, in the order of the method of the present disclosure, the comparative example, LASSO, and no constraint, the value of the correlation coefficient r becomes higher. In addition, as Figure 8 shown, in the order of LASSO, the method of the present disclosure and the comparative example, and no constraint, the value of the determination coefficient E approaches 1. From Figure 7A and Figure 7B it is also known that in general LASSO, if the parameter α is increased excessively, the accuracy decreases. That is, in LASSO, α is a so-called hyperparameter and requires adjustment based on cross-validation. On the other hand, according to the method of the present disclosure, by taking α sufficiently large, the accuracy can be improved. This has the effect of not requiring manual parameter adjustment. In addition, in the case of randomly assigning symbol constraints as in the comparative example, for example, the correlation coefficient r is lower than that of the method of the embodiment. That is, it can be said that the method of the embodiment can create a model as follows: The data to be analyzed has a certain correspondence relationship in explaining the variation of the variable and the variation of the target variable, and the fitting is particularly good when the corresponding symbol constraints are given. In addition, from Figure 7A and Figure 7B it can be seen that even in the case of the comparative example of randomly assigning symbol constraints represented by the dashed line, the correlation coefficient is higher than that in the case of no regularization represented by the double-dot dashed line. This shows that a model can be created as follows: Even if inappropriate symbol constraints are imposed on a part of the explanatory variables, the fitting is still good. In reality, there are often cases where the knowledge about the correspondence relationship between the variation of the explanatory variable and the variation of the target variable is not complete. Even in such a case, according to the method of the embodiment, the effect of being able to create a model with better fitting than the case without regularization can be achieved.

[0105] Figure 9 is a graph showing the relationship between the number of data T for learning and the correlation coefficient r for models constructed by multiple methods. Figure 10 is a graph showing the relationship between the number of data T for learning and the determination coefficient E for models constructed by multiple methods. AsFigure 9 As shown, for example, when the number of data T is 40 or less, the value of the correlation coefficient r increases in the order of the method of the present disclosure, the comparative example, LASSO, and no constraint. In addition, as Figure 10 shown, the value of the determination coefficient E approaches 1 in the order of LASSO, the method of the present disclosure, the comparative example, and no constraint. Thus, it can be said that the method of the present disclosure is effective when the training data is small. That is, it is also useful in cases where data cannot be collected sufficiently, or where the prediction model changes over time but only recent data can be used due to reasons such as changes in states that cannot be observed from the data alone.

[0106] <Effect>

[0107] According to the method of the present disclosure, a regression equation can be generated that satisfies a constraint such that there is a certain correspondence between the positive or negative direction of the change in the explanatory variable and the positive or negative direction of the change in the target variable. Therefore, when the user uses the regression equation, it can be known which direction, positive or negative, the value of the input x k should be changed in order to make the predicted value μ close to the desired value. In addition, as described using Figure 7A and Figure 7B , there is also an advantage that there is no need to adjust the parameter α representing the strength of the constraint. In addition, as described using Figure 9 and Figure 10 , the method of the present disclosure is particularly effective when the training data is small.

[0108] Hereinafter, the effects will be supplemented. Here, regarding the regularization term of Equation (2), it can be said as follows.

[0109] [Mathematical Formula 11]

[0110]

[0111] And, for example, when the constraint sign is positive (R + (w)), the subdifferential of the cost function E(w) of Equation (2) with respect to w k is obtained as follows.

[0112] [Mathematical Formula 12]

[0113]

[0114] Note that here, it is assumed that the multiple inputs x k are not correlated, and δ kk represents the identity matrix.

[0115] And w k is obtained as follows.

[0116] [Equation 13]

[0117]

[0118] In addition, if it is solved again, it is obtained as follows.

[0119] [Equation 14]

[0120]

[0121] Here, if α is sufficiently large, the following case is not considered, and w k can be represented by the following equation (11).

[0122] [Equation 15]

[0123] If we set α→+∞

[0124]

[0125] In the upper part of equation (11), it is the same solution as the least squares method. On the other hand, in the general least squares method, no sign constraint is imposed. Therefore, for example, when the number of data T is small, in the case corresponding to the lower part of equation (11), sometimes the same solution as the upper part of equation (11) can be obtained. In this case, it is not clear how to change the values of the explanatory variables in order to make the output of the regression equation close to the desired value. On the other hand, in such a case, according to the technology of the present disclosure, as shown in the lower part of equation (11), the coefficient w k is set to 0. That is, for the explanatory variable x k that does not satisfy the constraint, it is not used in the regression equation produced. Therefore, a regression equation can be generated that satisfies the constraint that there is a certain correspondence in the positive or negative direction of the change of the explanatory variable and the positive or negative direction of the change of the target variable. In addition, it can be said that the value of the parameter α can be set to a sufficiently large value and no adjustment is required.

[0126] In addition, in general LASSO, for example, w is obtained as follows k .

[0127] [Equation 16]

[0128] When

[0129]

[0130] That is, it is estimated by biasing in such a way that α is decreased from the value that should originally converge. Such a bias acts in a way that increases the mean square error. On the other hand, it can be said that according to the technology of the present disclosure, such a bias is not generated, and therefore, the accuracy of the regression equation is improved.

[0131] In addition, according to Equation (11), it satisfies the Oracle property (Fan and Li, 2001). That is, when the sample size increases, the probability of correctly selecting the explanatory variables for the model converges to 1 (consistency of variable selection). In addition, the estimator of the explanatory variables has asymptotic normality.

[0132] <Embodiment 2>

[0133] This embodiment can impose the above-mentioned sign constraints on the regression coefficients and can improve the sparsification performance. In addition, the parameter β for controlling the strength of regularization is set as a so-called hyperparameter. That is, in addition to Figure 6 the processing shown, the optimal value of the coefficient is determined by using an existing cross-validation method. In this embodiment, the cost function shown in the following Equation (12) is used instead of the cost function shown in Equation (2). It should be noted that the regression equation is the same as the regression equation shown in Equation (1).

[0134] [Equation 17]

[0135]

[0136] where

[0137]

[0138] β is a parameter for controlling the strength of regularization and takes a value of 0 or more. In addition, the optimal value of β is determined by using an existing method of cross-validation. The regularization term βR SL (w) also imposes sign constraints on the positive or negative single side. Specifically, in Figure 1 the table, when the constraint sign of x k is positive, the value of R SL+ (w) is taken, and when the constraint sign is negative, the value of R SL- (w) is taken. That is, the regularization term increases the cost according to the sum of the absolute values of the coefficients in the interval where the coefficient w k is either positive or negative corresponding to the constraint sign, and sets the cost to infinity in the other interval. In other words, not only is the cost set to infinity when it is inconsistent with the constraint sign (i.e., equivalent to the case where α in Equation (2) is set to infinity), but the cost is also increased according to β and w when it is consistent with the constraint sign.

[0139] Figure 11A and Figure 11B are schematic diagrams for explaining the constraints imposed on the regression coefficient w. In Figure 11A the graph, the vertical axis represents βR SL+(w), the horizontal axis represents w. The above formula (12) is for the case where the corresponding constraint symbol is positive for the input x k When establishing the corresponding constraint symbol as positive for the input x k and the coefficient w of k is 0 or more, R SL+ (w) = w and E(w) increases correspondingly with the increase of w. On the other hand, when the coefficient w of the input x k is k less than 0, R SL+ (w) = +∞ and the cost diverges to positive infinity. This is an infinite value based on the case of maximizing the prediction performance when the α shown in Figure 2A is a sufficiently large value. That is, for the regularization term of this embodiment, the cost is set to infinity in the interval where it does not match the constraint symbol, and the cost also increases according to the magnitudes of the regression coefficient w and the parameter β in the interval where it matches the constraint symbol. Here, when the coefficient w k is 0 or more, as the input x of the regression formula shown in formula (1) k increases, the predicted value μ based on the regression formula also increases. That is, in the case where the corresponding constraint symbol established with x k is positive, the regularization term becomes smaller when the value of the input x k increases and the value of the predicted value μ also increases, and the regularization term becomes larger when the value of the input x k increases and the value of the predicted value μ decreases. The cost function is defined in this way.

[0140] Figure 11B The vertical axis of the chart of SL- represents βR k (w), and the horizontal axis represents w. The above formula (12) is for the case where the corresponding constraint symbol established with the input x k is negative. When the coefficient w of the input x k is 0 or more, R SL- (w) = +∞ and the cost diverges to positive infinity. This is an infinite value set to maximize the prediction performance when the α shown in Figure 2B is a sufficiently large value, which is intended to be a sufficiently large value. On the other hand, when the coefficient w of the input x k is k less than 0, R SL- (w) = -w and E(w) increases correspondingly with the decrease of w. Here, when the coefficient w k is less than 0, as the input x of the regression formula shown in formula (1) k increases, the predicted value μ based on the regression formula decreases. That is, in the case where the corresponding constraint symbol established with x k is negative, the regularization term becomes smaller when the value of the input x k increases and the value of the predicted value μ decreases, and in the input xk The way in which the regularization term becomes larger as the value increases and the predicted value μ increases is defined as the cost function.

[0141] <Effect>

[0142] By cross-validation based on the Leave-one-out method, the method of this embodiment and the performance of the existing L1 regularization (LASSO) are evaluated. The number of learning data N is 10, and the number of features K is set to 11. Figure 12 It is a graph showing the relationship between the parameter β and the correlation coefficient r. Figure 12 The horizontal axis of the curve graph represents the parameter β, and the vertical axis represents the correlation coefficient r. In addition, the solid line represents the result of the method based on this embodiment, and the dashed line represents the result of the existing L1 regularization (LASSO). For the correlation coefficient r, especially in the range where β is less than 0.001, the result of the method of this embodiment is higher than the result of the existing LASSO. Figure 13 It is a graph showing the relationship between the parameter β and the coefficient of determination R 2 of. Figure 13 The horizontal axis of the curve graph represents the parameter β, and the vertical axis represents the coefficient of determination R 2 . In addition, the solid line represents the result of the method based on this embodiment, and the dashed line represents the result of the existing L1 regularization (LASSO). For the coefficient of determination R 2 Similarly, especially in the range where β is less than 0.001, the result of the method of this embodiment is higher than the result of the existing LASSO. Figure 14 It is a graph showing the relationship between the parameter β and the RMSE (Root Mean Square Error). Figure 14 The horizontal axis of the curve graph represents the parameter β, and the vertical axis represents the RMSE. In addition, the solid line represents the result of the method based on this embodiment, and the dashed line represents the result of the existing L1 regularization (LASSO). Similarly for the RMSE, especially in the range where β is less than 0.001, the result of the method of this embodiment is lower than the result of the existing LASSO. Generally speaking, when the number of explanatory variables is greater than the number of learning data, the number of equations is less than the number of variables to be solved. Therefore, if no regularization is applied, the regression coefficients cannot be uniquely determined. As Figures 12 to 14 shown, if regularization is performed by the method of this embodiment, the regression coefficients can be determined even when the number of explanatory variables is greater than the number of learning data, and moreover, compared with the existing LASSO, the prediction performance (generalization performance) can be improved.

[0143] <Modification example>

[0144] Each component and their combinations in each embodiment are examples, and additional components, omissions, substitutions, and other changes can be appropriately made without departing from the gist of the present invention. The present disclosure is not limited by the embodiments but only by the claims. In addition, various aspects disclosed in this specification can be combined with any other features disclosed in this specification.

[0145] Figure 5 The configuration of the computer shown is an example and is not limited to such an example. For example, at least a part of the functions of the regression analysis device 1 can be realized by dispersing them among multiple devices, or the same function can be provided in parallel by multiple devices. In addition, at least a part of the functions of the regression analysis device 1 can be provided on what is called the cloud. Further, for example, the regression analysis device 1 may not include a part of the configuration such as the verification processing unit 144.

[0146] In addition, the cost function shown in formula (2) is L1-regularized on the positive or negative single side, but it also operates according to the L2 norm or other convex functions. That is, the sum of the squares of the coefficients or other penalty terms with coefficients applied on the positive or negative single side can be used instead of the sum of the absolute values of the coefficients.

[0147] In addition, the content of the data analyzed by the regression analysis device 1 is not particularly limited. In addition to the prediction of characteristic values such as quality in manufacturing described in the embodiments, it can also be applied to non-manufacturing and various other fields.

[0148] In addition, the present disclosure includes a method for executing the above processing, a computer program, and a computer-readable recording medium recording the program. By causing a computer to execute the program, the recording medium recording the program can perform the above processing.

[0149] Here, a computer-readable recording medium refers to a recording medium that can store information such as data and programs through electrical, magnetic, optical, mechanical, or chemical actions and can be read by a computer. As a medium that can be removed from a computer among such recording media, there are floppy disks, magneto-optical disks, optical disks, magnetic tapes, memory cards, etc. In addition, as a recording medium fixed to a computer, there are HDDs, SSDs (Solid State Drives), ROMs, etc.

[0150] Explanation of Reference Numerals

[0151] 1: Regression analysis device

[0152] 11: Communication I / F

[0153] 12: Storage device

[0154] 13: Input / output device

[0155] 14: Processor

[0156] 141: Data acquisition unit

[0157] 142: Coefficient update unit

[0158] 143: Convergence determination unit

[0159] 144: Verification processing unit

[0160] 145: Application processing unit

Claims

1. A regression analysis device for quality prediction in a production device, comprising: A data acquisition unit that reads training data and constraint conditions from a storage device, where the storage device stores the training data, which is sensed data obtained from the production device and serves as the target variable and explanatory variables of the regression model, and the constraint conditions that are predefined to indicate whether the explanatory variable should change in the positive or negative direction to cause the target variable to change in the positive or negative direction; and A coefficient update unit that repeatedly updates the coefficients of the explanatory variables in the regression model using the training data in such a way as to minimize a cost function including a regularization term, where the regularization term increases the cost when the constraint conditions are violated. Information representing the constraint conditions is preset in advance based on knowledge related to the production device.

2. The regression analysis device according to claim 1, wherein the regularization term increases the cost according to the sum of the absolute values of the coefficients in a positive or negative interval corresponding to the constraint conditions.

3. The regression analysis device according to claim 1 or 2, wherein when the coefficients do not converge to values that satisfy the constraint conditions, the coefficient update unit sets the coefficients to 0.

4. The regression analysis device according to claim 1, wherein the coefficient update unit updates the coefficients by the proximal gradient method.

5. A regression analysis method for quality prediction in a production device, characterized in that a computer reads training data and constraint conditions from a storage device, where the storage device stores the training data, which is sensed data obtained from the production device and serves as the target variable and explanatory variables of the regression model, and the constraint conditions that are predefined to indicate whether the explanatory variable should change in the positive or negative direction to cause the target variable to change in the positive or negative direction, the computer repeatedly updates the coefficients of the explanatory variables in the regression model using the training data in such a way as to minimize a cost function including a regularization term, where the regularization term increases the cost when the constraint conditions are violated, and information representing the constraint conditions is preset in advance based on knowledge related to the production device.

6. A computer-readable recording medium that stores a program for quality prediction in a production device, characterized in that the program causes a computer to read training data and constraint conditions from a storage device, where the storage device stores the training data, which is sensed data obtained from the production device and serves as the target variable and explanatory variables of the regression model, and the constraint conditions that are predefined to indicate whether the explanatory variable should change in the positive or negative direction to cause the target variable to change in the positive or negative direction, the program causes the computer to repeatedly update the coefficients of the explanatory variables in the regression model using the training data in such a way as to minimize a cost function including a regularization term, where the regularization term increases the cost when the constraint conditions are violated, and information representing the constraint conditions is preset in advance based on knowledge related to the production device.

Citation Information

Patent Citations

  • Sample component determination method based on optimizing partial least squares regression model

    CN104949936A

  • CNN algorithm and Lasso regression model-based hot-rolling product quality prediction method

    CN110264079A