Data processing method and device

By detecting and replacing outliers and generating outlier-replaced learning data, the problems of decreased prediction accuracy and narrowed numerical range in machine learning are solved, achieving higher prediction accuracy and a wider predictable range.

CN120705749APending Publication Date: 2025-09-26PROTERIAL LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510140472.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2025-02-08
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In machine learning, using learning data containing erroneous data or outliers will lead to a decrease in the prediction accuracy of the regression model, and the amount of data will decrease after removing outliers, resulting in a narrower range of predictable values.

Method used

By detecting outliers, generating predicted values ​​and replacing them with values ​​in the learning data, outlier-substituted learning data is generated and used to train a regression model to improve prediction accuracy and suppress the narrowing of the numerical range.

Benefits of technology

The prediction accuracy of machine learning is improved while maintaining the amount of training data and avoiding the narrowing of the numerical range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705749A_ABST
    Figure CN120705749A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and a data processing device, which can improve prediction accuracy and restrain a condition that a predictable numerical range is narrowed. This data processing method is provided with: an outlier detection step for detecting outliers from a plurality of pieces of data included in learning data (31); a predicted value calculation step for generating a first regression model (36) using, as training data, first learning data (35) from which the outliers have been removed from the learning data (31), and obtaining, using the first regression model (36), a first predicted value, which is a predicted value of the target variable corresponding to the value of the explanation variable for each of the removed outliers; and a data replacement step for replacing the removed outliers from the learning data (31) with values based on the first predicted values to generate outlier-replaced learning data (38).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data processing method and device in machine learning. Background Art

[0002] Various prediction methods using machine learning are known. For example, when predicting the physical properties of a material with an unknown formula, machine learning is performed using data already acquired through experimental production as learning data (training data). The correlation between the material formula and the physical properties is learned, and the prediction is performed using the regression model obtained as a result of the learning.

[0003] Here, if the learning data contains erroneous data or data with large errors, i.e., outliers, the prediction accuracy of the regression model obtained using the learning data will decrease. Therefore, before machine learning, outliers are removed from the learning data (for example, see Patent Document 1).

[0004] However, if outliers are removed from the training data, the amount of training data decreases, which in turn narrows the range of values ​​that can be accurately predicted by a regression model generated using the training model.

[0005] Patent Document 1: Japanese Patent Application Laid-Open No. 2021-33544 Summary of the Invention

[0006] Therefore, an object of the present invention is to provide a data processing method that improves prediction accuracy and can suppress a narrowing of a predictable numerical range.

[0007] The present invention aims to solve the above-mentioned problems and provides a data processing method for processing learning data including multiple data consisting of explanatory variables and target variables. The data processing method comprises the following steps: an outlier detection step for detecting outliers from the learning data; a predicted value calculation step for generating first learning data by removing the outliers detected in the outlier detection step from the learning data, using the first learning data as training data to generate a first regression model, and using the first regression model to obtain a predicted value of the target variable corresponding to the value of the explanatory variable for the removed outlier, i.e., a first predicted value; and a data replacement step for replacing the removed outliers from the learning data with a value based on the first predicted value to generate outlier-replaced learning data.

[0008] The present invention aims to solve the above-mentioned problems and provides a data processing device, which performs data processing on learning data including multiple data consisting of explanatory variables and target variables, the data processing device comprising: an outlier detection processing unit, which detects outliers from the learning data; a predicted value calculation processing unit, which generates first learning data by removing the outliers detected by the outlier detection processing unit from the learning data, uses the first learning data as training data to generate a first regression model, and uses the first regression model to obtain a predicted value of the target variable corresponding to the value of the explanatory variable of the removed outlier, i.e., a first predicted value; and a data replacement processing unit, which replaces the removed outliers from the learning data with values ​​based on the first predicted values ​​to generate outlier-replaced learning data.

[0009] According to the present invention, it is possible to provide a data processing method and apparatus that improves prediction accuracy and suppresses a narrowing of a predictable numerical value range. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a schematic diagram of the structure of a data processing device according to one embodiment of the present invention.

[0011] Figure 2 This is a diagram showing an example of learning data.

[0012] Figure 3 (a) is a diagram showing an example of the calculation results of the coefficient of variation, and (b) is a diagram showing the calculation results of (a) as a histogram.

[0013] Figure 4 This is an explanatory diagram showing training data used in the outlier determination process.

[0014] Figure 5 It means in Figure 3 Graph showing the calculation results of the error rate of each outlier candidate data obtained in (a).

[0015] Figure 6 This is an explanatory diagram illustrating the first learning data.

[0016] Figure 7 (a) is a flowchart of a data processing method according to one embodiment of the present invention, and (b) is a flowchart of the outlier detection process.

[0017] Figure 8 It is a flow chart of data classification processing.

[0018] Figure 9 This is a flowchart of the outlier determination process.

[0019] Figure 10This is a flowchart of the predicted value calculation process.

[0020] Figure 11 This is a flowchart of the data replacement process.

[0021] Figure 12 This is a flowchart of the prediction process.

[0022] Figure 13 This is an explanatory diagram for explaining first learning data according to a modified example of the present invention.

[0023] Figure 14 This is a flowchart of the predicted value calculation process according to a modified example of the present invention.

[0024] Figure 15 It is a graph showing the calculation results of MAPE of Examples and Comparative Examples. DETAILED DESCRIPTION

[0025] [Implementation Method]

[0026] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.

[0027] Figure 1 : This is a schematic diagram of the structure of the data processing device 1 of this embodiment. The data processing device 1 has the function of detecting outliers from the learning data 31 used for machine learning and replacing the value of the target variable of the detected outlier with a predicted value. In addition, the data processing device 1 has the function of performing predictions using a regression model 39, which is generated using the learning data obtained by replacing the outliers with the predicted values ​​(the outlier-replaced learning data 38 described later). In addition, for example, an outlier refers to a value that deviates significantly from other data due to measurement errors, reading errors of measuring instruments, input errors, or human errors, or the influence of noise.

[0028] Data processing device 1 includes control unit 2 and storage unit 3. Data processing device 1 is a computer such as a personal computer or server, and includes a CPU and other computing elements, RAM and ROM and other memories, a hard disk and other storage devices, and a LAN card and other communication interface as a communication device.

[0029] The control unit 2 includes a data acquisition processing unit 21, an outlier detection processing unit 22, a predicted value calculation processing unit 23, a data replacement processing unit 24, a prediction processing unit 25, and a prediction result presentation processing unit 26. Details of each unit will be described later. The storage unit 3 is implemented by a predetermined storage area of ​​a memory or storage device.

[0030] Data processing device 1 also includes a display 4 and an input device 5. Display 4 is, for example, a liquid crystal display, and input device 5 is, for example, a keyboard or mouse. Alternatively, display 4 may be formed of a touch panel, with display 4 also serving as input device 5. Alternatively, display 4 and input device 5 may be separate from data processing device 1 and capable of communicating with data processing device 1 via wireless communication or the like. In this case, display 4 and input device 5 may be formed of portable devices such as tablet computers and smartphones.

[0031] (Data Acquisition Processing Unit 21)

[0032] The data acquisition processing unit 21 performs data acquisition processing to obtain learning data 31 from an external device. In the data acquisition processing, data is acquired via a network, for example, from a server device storing manufacturing performance data, and the acquired data is stored in the storage unit 3 as learning data 31. Furthermore, the learning data 31 may be input to the data processing device 1 via a medium such as a USB memory, and the method for acquiring the learning data 31 is not particularly limited.

[0033] (Learning data 31)

[0034] Here, the learning data 31 will be described. Figure 2 31 is a diagram showing an example of learning data 31. The learning data 31 is a database used as training data in machine learning, and includes a plurality of data consisting of explanatory variables and target variables used for machine learning. Figure 2 The example shows a case where the formulation amounts of materials such as polymers and fillers are used as explanatory variables, and the physical properties of composite materials produced using these materials (here, tensile strength) are used as target variables. By performing machine learning using this learning data 31, a regression model is generated that represents the correlation between the explanatory variables (the formulation amounts of each material) and the target variable (physical properties). This allows prediction of the physical properties obtained when a composite material is produced using an unknown material formulation.

[0035] In this embodiment, each data item included in the learning data 31 includes values ​​for multiple target variables. Here, each data item includes values ​​for five target variables (tensile strength in the example shown). These target variable values ​​are obtained, for example, by producing a composite material using the same recipe to form multiple samples (here, five samples No. 1 to 5) and measuring the physical properties of each sample, such as tensile strength. Here, the case where the learning data 31 initially includes 1,316 data items will be described.

[0036] (Outlier Detection Processing Unit 22)

[0037] The outlier detection processing unit 22 performs outlier detection processing to detect outliers from the multiple data included in the learning data 31. The outlier detection processing corresponds to the outlier detection step of the present invention. The outlier detection processing unit 22 includes a data classification processing unit 221, an outlier determination processing unit 222, and a detection result presentation processing unit 223. The specific processing details of the outlier detection processing described below are merely examples, and outliers may also be detected using other methods. Specifically, the specific method for detecting outliers can be appropriately selected and is not limited to the method described below.

[0038] (Data Classification Processing Unit 221)

[0039] The data classification processing unit 221 performs data classification processing to extract data that are candidates for outliers from the learning data 31. The data classification processing corresponds to the data classification step of the present invention.

[0040] In the data classification process, first, the coefficient of variation of the target variable value is obtained for each data included in the learning data 31. The coefficient of variation can be obtained by the following formula (1).

[0041] (Coefficient of variation) = {(Standard deviation) / (Average value)} × 100 (1)

[0042] Then, among the data included in the learning data 31, data for which the coefficient of variation is greater than a predetermined reference value is classified as outlier candidate data 32, and the remaining data is classified as normal data 33 and stored in the storage unit 3. A large coefficient of variation indicates a large deviation in the value of the target variable, and the possibility of containing an outlier is considered to be increased. In addition, here, the outlier candidate data 32 and normal data 33 are stored in the storage unit 3 as data different from the learning data 31, but this is not limited to this. For example, it is also possible to distinguish the outlier candidate data 32 from the normal data 33 by adding a mark or tag to each data item in the learning data 31. In other words, a portion of the learning data 31 can also be set as the outlier candidate data 32 and the normal data 33.

[0043] The reference value for the coefficient of variation used to determine outlier candidates can be set appropriately. For example, the reference value can be set based on the target variable or the deviation of the entire data set. It can be set to the average, median, mode, average ± σ (σ is the standard deviation), average ± 2σ, average ± 3σ, etc., of the coefficient of variation of all data included in the learning data 31. Here, the coefficient of variation of data for a representative recipe, including actual manufacturing performance, is used as the reference value.

[0044] Figure 3(a) is a diagram showing an example of the calculation results of the coefficient of variation. Figure 3 In (a), the coefficient of variation is calculated for 1316 data points and plotted in ascending order. In the example, the coefficient of variation in a representative recipe is used as the reference value, set to 17.7. Figure 3 As shown in (a) of FIG. 1 , in the example shown in the figure, there are 87 data items exceeding the reference value. In the data classification process, these 87 data items are regarded as outlier candidate data items 32 , and the remaining 1229 data items are regarded as normal data items 33 .

[0045] Figure 3 (b) is to Figure 3 (a) The calculation results of the coefficient of variation are expressed as a histogram. Figure 3 In (b), the average value, median value, average value + σ, average value + 2σ, and average value + 3σ of the coefficient of variation of all data included in the learning data 31 are also shown. Figure 3 As shown in (b), the value 17.7 currently used as the reference value is larger than the average value + σ and smaller than the average value + 2σ.

[0046] (Outlier Determination Processing Unit 222)

[0047] The outlier determination processing unit 222 performs an outlier determination process for determining whether each data item included in the outlier candidate data 32 picked up in the data classification process is an outlier. The outlier determination process corresponds to the outlier determination step of the present invention.

[0048] like Figure 4 As shown, in the present embodiment, in the outlier determination process, normal data 33 (number of data: 1229) is used as training data to generate a regression model for representing the correlation between the explanatory variable and the target variable, namely a second regression model (regression model for outlier determination). In addition, in the present embodiment, each data contains multiple target variable values, but here, the central value of the values ​​of the multiple target variables is used for learning. Then, for each data of the outlier candidate data 32 (number of data: 87), the generated second regression model is used to calculate the second predicted value as the predicted value, and the error rate between the obtained second predicted value and the actual value of the target variable (actual value) is calculated. In addition, in the present embodiment, each data contains multiple target variable values, but here, the central value of the values ​​of the multiple target variables is used as the actual value. The error rate is calculated by the following formula:

[0049] (Error rate) = 100 × {(Predicted value) - (Actual value)} / (Predicted value)

[0050] The outlier determination processing unit 222 determines whether each data item in the outlier candidate data 32 is an outlier based on the error rate calculated by calculation. In this embodiment, the data item is determined to be an outlier when the absolute value of the error rate is greater than a preset threshold.

[0051] Figure 5 Indicates Figure 3 The calculation result of the error rate of each of the 87 outlier candidate data 32 obtained in (a). In the illustrated example, the threshold is set to 20%, but the threshold can be set appropriately. In this case, when the error rate is greater than +20% or less than -20%, the data is determined to be an outlier. Figure 5 In the example of , 33 data out of 87 outlier candidate data 32 are determined to be outliers. The data determined to be outliers are stored in the storage unit 3 as outlier data 34.

[0052] (Detection Result Prompt Processing Unit 223)

[0053] The detection result presentation processing unit 223 performs a detection result presentation process for presenting the outlier determination result, i.e., the outlier detection result. During the detection result presentation process, data detected as outliers (outlier data 34) is displayed on the display 4, etc., to present the data detected as outliers to the user. The detection result presentation processing unit 223 is not essential and may be omitted.

[0054] (Prediction Value Calculation Processing Unit 23)

[0055] The predicted value calculation processing unit 23 performs predicted value calculation processing on each outlier detected in the outlier detection processing to obtain a predicted value (first predicted value) of the target variable. The predicted value calculation processing is equivalent to the predicted value calculation step of the present invention. More specifically, the predicted value calculation processing unit 23 first performs the following steps: Figure 6 As shown, first learning data 35 (learning data for predictive value calculation) is generated by removing the outlier data 34, which are data of outliers detected in the outlier detection process, from the learning data 31. In this embodiment, first learning data 35 (the number of data is 1283 (=1229+(87-33))) is generated by removing all the outlier data 34 from the learning data 31. Then, using the first learning data 35 as training data, a regression model representing the correlation between the explanatory variable and the target variable, namely, a first regression model 36, is generated. The generated first regression model 36 is stored in the storage unit 3.

[0056] The predicted value calculation processing unit 23 then uses the generated first regression model 36 to calculate, for each piece of outlier data 34 (33 pieces of data), a predicted value of the target variable corresponding to the value of the explanatory variable contained in that piece of data. The first predicted value corresponding to each outlier obtained through the predicted value calculation processing is stored in the storage unit 3 as predicted value data 37 (33 pieces of data).

[0057] (Data Replacement Processing Unit 24)

[0058] The data replacement processing unit 24 performs the following data replacement processing: the value of the target variable of each outlier in the learning data 31 (the outlier removed from the learning data 31 in the prediction value calculation process) is replaced by the value of the first prediction value obtained by the prediction value calculation process, and the outlier replaced learning data 38 is generated. In this embodiment, the value of the target variable of the outlier is replaced by the first prediction value. The data replacement processing is equivalent to the data replacement process of the present invention. The data replacement processing unit 24 replaces the generated outlier replaced learning data 38 (the number of data is 1316, refer to Figure 6 ) is stored in storage unit 3.

[0059] (Prediction Processing Unit 25)

[0060] The prediction processing unit 25 uses the outlier replaced learning data 38 obtained by the data replacement processing unit 24 to perform prediction processing for the value of the target variable of the prediction object (third prediction value). The prediction processing is equivalent to the prediction step of the present invention. In the prediction processing, first, the outlier replaced learning data 38 is used as training data to generate a regression model representing the correlation between the explanatory variable and the target variable, that is, a third regression model 39 (regression model for prediction of physical properties, etc.). In addition, for data containing values ​​of multiple target variables, the central value of the values ​​of multiple target variables is used for learning. The generated third regression model 39 is stored in the storage unit 3. Thereafter, the generated third regression model 39 is used to predict the predicted value of the target variable corresponding to the value of the explanatory variable of the prediction object, that is, the third prediction value. For example, in Figure 2 In the example, the target variable is tensile strength (physical property).

[0061] More specifically, the prediction source data 40, which is the value of each explanatory variable inputted via the input device 5, is applied to the regression model 39 to obtain the corresponding target variable value, which is the third predicted value. The obtained third predicted value is stored in the storage unit 3 as prediction data 41.

[0062] (Prediction Result Presentation Processing Unit 26)

[0063] The prediction result presentation processing unit 26 performs a prediction result presentation process for presenting the prediction result in the prediction process. In the prediction result presentation process, for example, prediction data 41 obtained by the prediction process is displayed on the display 4.

[0064] (Data Processing Method)

[0065] Figure 7 (a) is a flow chart of the data processing method of this embodiment. Figure 7 As shown in (a), first, in step S1, data acquisition processing is performed. In the data acquisition processing, the data acquisition processing unit 21 acquires learning data 31 (in Figure 6 In the case of , the number of data is 1316). The acquired learning data 31 is stored in the storage unit 3.

[0066] Then, in step S2, outlier detection processing is performed. In the outlier detection processing, as shown in FIG. Figure 7 As shown in (b), first, the data classification process of step S7 is performed. In the data classification process, as shown in FIG. Figure 8 As shown, first, in step S71, the variation coefficient of each data included in the learning data 31 is calculated. Then, in step S72, the data whose variation coefficient is larger than the preset reference value is regarded as the outlier candidate data 32 (in Figure 6 In the case of , the number of data is 87) is stored in the storage unit 3. Then, in step S73, the data with a coefficient of variation below the reference value is regarded as normal data 33 (in Figure 6 In the case of , the number of data is 1229) stored in the storage unit 3. Then, return to enter Figure 7 Step S8 of (b).

[0067] In step S8, outlier determination processing is performed. In the outlier determination processing, as shown in FIG. Figure 9 As shown, first, in step S81, the variable n representing the data number is substituted with 1 as the initial value, and the number of data of the outlier candidate data 32 is substituted with n_max. Thereafter, in step S82, the normal data 33 is used as training data to generate a second regression model. Thereafter, in step S83, the value of the explanatory variable of the n-th data in the outlier candidate data 32 is applied to the second regression model to obtain the second predicted value (the value of the target variable), and in step S84, the error rate between the second predicted value and the actual value (the value of the target variable of the n-th data) is obtained. Thereafter, in step S85, it is determined whether the absolute value of the obtained error rate is above a preset threshold value. If it is determined to be "yes" (Y) in step S85, in step S86, the n-th data is determined to be an outlier, and the outlier data 34 (in Figure 6In the case of "No" (N) in step S85, the number of data is 33 (stored in the storage unit 3), and the process proceeds to step S88. In the case of "No" (N) in step S85, the nth data is determined to be not an outlier in step S87, and the process proceeds to step S88. In step S88, it is determined whether the variable n is greater than n_max. In the case of "No" (N) in step S88, the variable n is incremented in step S89, and the process returns to step S82. In the case of "Yes" (Y) in step S88, the process returns and the process proceeds to step S88. Figure 7 Step S9 of (b).

[0068] In step S9, the detection result prompt processing is performed. In the detection result prompt processing, the data determined as outliers in step S8, that is, the outlier data 34, is displayed on the display 4, etc. to prompt the detected outliers. After that, return to enter Figure 7 Step S3 of (a).

[0069] In step S3, the predicted value calculation process is performed. In the predicted value calculation process, Figure 10 As shown, first, in step S31, the variable m representing the data number is substituted with 1 as an initial value, and the number of data of the outlier data 34 is substituted into m_max. Then, in step S32, the predicted value calculation processing unit 23 generates the first learning data 35 (in step S32) by removing the outliers (outlier data 34) from the learning data 31. Figure 6 In the case of , the number of data is 1283 (= 1229 + (87-33))). In step S33, the first learning data 35 is used as training data to generate a first regression model 36. Then, in step S34, the value of the explanatory variable of the m-th data in the outlier data 34 is applied to the first regression model 36 to obtain a first predicted value (the value of the target variable). In step S35, the obtained first predicted value is used as the predicted value data 37 (in Figure 6 In the case of , the number of data is 33) stored in the storage unit 3. Then, in step S36, it is determined whether the variable m is greater than m_max. If the determination in step S36 is "No" (N), in step S37, the variable m is incremented and the process returns to step S34. If the determination in step S36 is "Yes" (Y), the process returns and enters Figure 7 Step S4 of (a).

[0070] In step S4, data replacement processing is performed. In the data replacement processing, as shown in FIG. Figure 11 As shown, in step S41, in the learning data 31, the value of the target variable of the outlier (outlier data 34) is replaced with the first predicted value (predicted value data 37), and the outlier replaced learning data 38 (in Figure 6In the case of , the number of data is 1316) and stored in the storage unit 3. Then, return to enter Figure 7 Step S5 of (a).

[0071] In step S5, prediction processing is performed. In the prediction processing, Figure 12 As shown, first, in step S51, the prediction source data 40 is input using the input device 5 or the like. The input prediction source data 40 is stored in the storage unit 3. Then, in step S52, the outlier replaced learning data 38 is used as training data to generate a third regression model 39, and the model is stored in the storage unit 3. Then, in step S53, the prediction source data 40 is applied to the regression model 39 to obtain a third prediction value (the value of the target variable), and in step S54, the obtained third prediction value is stored in the storage unit 3 as prediction data 41. Then, return to enter Figure 7 Step S6 of (a).

[0072] In step S6, a prediction result presentation process is performed. In the prediction result presentation process, the prediction result (prediction data 41) of the prediction process is displayed on the display 4, etc., to present the prediction result. Alternatively, the prediction source data 40 corresponding to the prediction data 41 may be displayed on the display 4, etc. The process then terminates.

[0073] (Variation)

[0074] In this embodiment, in the prediction value calculation process, the data after removing all outlier data 34 from the learning data 31 is used as the first learning data 35 to obtain the first prediction value, but it is not limited to this. Figure 13 As shown in FIG, when obtaining the first predicted value of a certain outlier, the outlier data 34 other than the outlier can also be included in the first learning data 35. That is, all data other than the outlier for which the first predicted value is obtained can also be used as the first learning data 35 (in Figure 13 In this case, the number of data is 1315. In this case, a first regression model 36 is generated for each outlier separately (in Figure 13 In this case, the number of regression models 36 generated for predictive value calculation is 33).

[0075] That is, in the predicted value calculation process, for each outlier detected in the outlier detection step, the first learning data 35 obtained by removing the target outlier data from the learning data 31 (in this modified example, only the target outlier data is removed) can be used as training data to generate a first regression model 36. The generated first regression model 36 can then be used to determine a predicted value of the target variable corresponding to the value of the explanatory variable for the target outlier data, i.e., a first predicted value.

[0076] Figure 14 The control flow of the predicted value calculation process at this time is shown. Figure 14 The control flow shown in Figure 10 In addition to replacing step S32 with step S32a and changing the return destination from step S37 to step S32a, it becomes the same as Figure 10 Same content. Figure 14 As shown, in step S32 a , only the m-th data of the outlier data 34 is removed from the learning data 31 to generate the first learning data 35 .

[0077] In this embodiment, the data replacement process simply replaces the value of the target variable for the outlier with the first predicted value. However, the data replacement process is not limited to this. The data may be replaced with a value based on the first predicted value. For example, the data may be replaced with the average of the value of the target variable for the outlier (e.g., the median value if there are multiple values) and the first predicted value.

[0078] (Regarding Improvement of Prediction Accuracy in the Present Embodiment)

[0079] The prediction accuracy when the data processing method of this embodiment is used is obtained. Outliers are detected using the learning data 31 with 1197 data. In Example 1, the first prediction value is obtained using the first learning data 35 from which all detected outlier data 34 are removed. In Example 2, as in Figure 13 、 Figure 14 As described in , all of the learning data 31 except for the outliers for which the first predicted value is to be calculated is used as the first learning data 35 to calculate the first predicted value. In Examples 1 and 2, the third regression model is generated using the outlier-substituted learning data 38 in which the values ​​of the target variables of the outliers are replaced with the first predicted values, and the MAPE (mean absolute error) is calculated using 81 data prepared separately from the learning data 31 as test data.

[0080] In addition, for comparison, the MAPE calculation was performed in the same manner as in Examples 1 and 2 for Comparative Example 1 (no processing) in which the third regression model was generated using the learning data 31 containing outliers directly and Comparative Example 2 (all outliers removed) in which the third regression model was generated using the learning data with all outliers removed. Figure 15 The results are summarized in .

[0081] like Figure 15As shown, the MAPE of Comparative Example 2 is lower than that of Comparative Example 1, and prediction accuracy is slightly improved by removing outliers. However, in Comparative Example 2, removing outliers reduces the amount of data used for learning, narrowing the range of values ​​that can be predicted with high accuracy. In contrast, in Examples 1 and 2 of the present invention, the MAPE is lower than in Comparative Examples 1 and 2, confirming improved prediction accuracy. Furthermore, in Examples 1 and 2, the target variable value is replaced with the first predicted value without removing outliers, so the amount of data is not reduced, and the range of values ​​of the explanatory variable that can be predicted with high accuracy is wider than in Comparative Example 2.

[0082] (Functions and Effects of Implementation Methods)

[0083] As described above, the data processing method of this embodiment includes a prediction value calculation process, in which the first learning data 35 after removing the outliers detected in the outlier detection process is used as training data to generate a first regression model 36, and the first regression model 36 is used to obtain the predicted value of the target variable corresponding to the value of the explanatory variable of each outlier, that is, the first predicted value, and the outliers in the learning data 31 are replaced by values ​​based on the first predicted value to generate outlier-replaced learning data 38.

[0084] By using the first regression model 36 from which outliers have been removed in the prediction value calculation process as training data, it is possible to obtain the predicted value of the target variable relative to the value of the explanatory variable for the outliers, i.e., the first predicted value, with high accuracy. As a result, it is possible to improve the prediction accuracy in the prediction process using the outlier-replaced learning data 38 as training data after replacing the outliers with the first predicted values ​​(see Figure 15 ). In addition, in this embodiment, the outliers are not removed and replaced with the first predicted values, so the amount of training data does not decrease. Therefore, according to this embodiment, it is possible to suppress the narrowing of the numerical range of the explanatory variable that can be predicted with high accuracy.

[0085] (Summary of Implementation Methods)

[0086] Next, the technical ideas grasped from the above-described embodiments will be described by citing the reference numerals in the embodiments. However, the reference numerals in the following description do not limit the components in the patent protection scope to the components specifically shown in the embodiments.

[0087] [1] A data processing method for processing learning data (31) comprising a plurality of data consisting of explanatory variables and target variables, the data processing method comprising the following steps: an outlier detection step for detecting outliers from the learning data (31); a predicted value calculation step for generating first learning data (35) by removing the outliers detected in the outlier detection step from the learning data (31), using the first learning data (35) as training data to generate a first regression model (36), and using the first regression model to obtain a predicted value of the target variable corresponding to the value of the explanatory variable for the removed outlier, i.e., a first predicted value; and a data replacement step for replacing the removed outlier from the learning data (31) with a value based on the first predicted value to generate outlier-replaced learning data (38).

[0088] [2] The data processing method according to [1], wherein, in the prediction value calculation step, the first learning data (35) is generated by removing all the outliers detected in the outlier detection step from the learning data (31), the first learning data (35) is used as training data to generate the first regression model (36), and the first regression model (36) is used to obtain the first prediction values ​​corresponding to the removed outliers.

[0089] [3] The data processing method according to [1], wherein, in the prediction value calculation step, for each of the outliers detected in the outlier detection step, the first learning data (35) is generated by removing the respective data from the learning data (31), the first learning data (35) is used as training data to generate the first regression model (36), and the first regression model (36) is used to obtain the first prediction value corresponding to the respective data.

[0090] [4] According to the data processing method described in [1], each data included in the learning data (31) includes multiple values ​​of the target variable, and the outlier detection process has the following processes: a data classification process, calculating the variation coefficient of the value of the target variable for each data, and classifying the data into outlier candidate data (32) and normal data (33) based on the obtained variation coefficient and the reference value; and an outlier determination process, using the normal data (33) as training data to generate a second regression model, using the second regression model to calculate the predicted value of the target variable corresponding to the value of the explanatory variable of each data of the outlier candidate data (32), that is, the second predicted value, and judging whether each data of the outlier candidate data is an outlier based on the second predicted value and the value of the target variable of each data of the outlier candidate data.

[0091] [5] The data processing method according to [1] also includes a prediction step, in which the outlier-replaced learning data (38) is used as training data to generate a third regression model (39), and the third regression model (39) is used to predict the predicted value of the target variable corresponding to the value of the explanatory variable of the prediction object, that is, the third predicted value.

[0092] [6] A data processing device (1) that processes data of learning data (31) including a plurality of data consisting of explanatory variables and target variables, the data processing device comprising: an outlier detection processing unit (22) that detects outliers from the learning data (31); a predicted value calculation processing unit (23) that generates first learning data (35) by removing the outliers detected by the outlier detection processing unit (22) from the learning data (31), uses the first learning data (35) as training data to generate a first regression model (36), and uses the first regression model (36) to obtain a predicted value of the target variable corresponding to the value of the explanatory variable for the removed outlier, i.e., a first predicted value; and a data replacement processing unit (24) that replaces the removed outlier from the learning data (31) with a value based on the first predicted value to generate outlier-replaced learning data (38).

[0093] (Note)

[0094] While the embodiments of the present invention have been described above, the embodiments described above do not limit the invention to which the patent protection scope relates. Furthermore, it should be noted that the combination of features described in the embodiments is not necessarily necessary for all technical means for solving the invention's problems. Furthermore, the present invention can be implemented with appropriate modifications within the scope of its main purpose.

[0095] Description of Reference Signs

[0096] 1Data processing device

[0097] 2. Control Unit

[0098] 21 Data acquisition processing unit

[0099] 22 Outlier detection processing unit

[0100] 221 Data Classification Processing Department

[0101] 222 Outlier determination processing unit

[0102] 23 Prediction value calculation processing unit

[0103] 24Data replacement processing unit

[0104] 25 Prediction Processing Department

[0105] 3 Storage unit

[0106] 31 Learning Data

[0107] 32 outlier candidate data

[0108] 33 Normal data

[0109] 34 Outlier Data

[0110] 35 First Learning Data

[0111] 36 First regression model

[0112] 37 Prediction data

[0113] 38 outliers have been replaced with learning data

[0114] 39The third regression model.

Claims

1. A data processing method for processing learning data including a plurality of data consisting of explanatory variables and target variables, characterized in that: The data processing method comprises the following steps: an outlier detection step of detecting outliers from the learning data; a predicted value calculation step of generating first learning data by removing the outliers detected in the outlier detection step from the learning data, using the first learning data as training data to generate a first regression model, and using the first regression model to determine a predicted value of the target variable corresponding to the value of the explanatory variable for the removed outliers, i.e., a first predicted value; and The data replacement step replaces the removed outlier from the learning data with a value based on the first predicted value to generate outlier-replaced learning data.

2. The data processing method according to claim 1, wherein: In the prediction value calculation process, the first learning data is generated by removing all the outliers detected in the outlier detection process from the learning data, the first learning data is used as training data to generate the first regression model, and the first regression model is used to calculate the first prediction values ​​corresponding to the removed outliers.

3. The data processing method according to claim 1, wherein: In the prediction value calculation process, for each outlier detected in the outlier detection process, the first learning data is generated by removing the each data from the learning data, the first learning data is used as training data to generate the first regression model, and the first regression model is used to obtain the first prediction value corresponding to the each data.

4. The data processing method according to claim 1, wherein: Each data included in the learning data includes a plurality of values ​​of the target variable. The outlier detection process has the following steps: a data classification step of obtaining a coefficient of variation of the target variable for each data item, and classifying each data item into outlier candidate data item and normal data item based on the obtained coefficient of variation and a reference value; as well as The outlier determination process uses the normal data as training data to generate a second regression model, uses the second regression model to calculate the predicted value of the target variable corresponding to the value of the explanatory variable of each data of the outlier candidate data, that is, the second predicted value, and determines whether the each data of the outlier candidate data is an outlier based on the second predicted value and the value of the target variable of the each data of the outlier candidate data.

5. The data processing method according to claim 1, wherein: The data processing method further includes a prediction step, in which the outlier-substituted learning data is used as training data to generate a third regression model, and the third regression model is used to predict a predicted value of the target variable corresponding to the value of the explanatory variable of the prediction object, i.e., a third predicted value.

6. A data processing device for processing data for learning including a plurality of data consisting of explanatory variables and target variables, characterized in that: The data processing device comprises: an outlier detection processing unit configured to detect outliers from the learning data; a predicted value calculation processing unit that generates first learning data by removing the outliers detected by the outlier detection processing unit from the learning data, uses the first learning data as training data to generate a first regression model, and uses the first regression model to determine a predicted value of the target variable corresponding to the value of the explanatory variable with the outliers removed, i.e., a first predicted value; and A data replacement processing unit replaces the removed outlier from the learning data with a value based on the first predicted value to generate outlier-replaced learning data.

Citation Information

Patent Citations

  • Learning data refining method and computer system

    JP2021033544A