Data processing apparatus, data processing system, and data processing method
By selectively adjusting the number of feature amounts for each classification label in the data processing apparatus, the method enhances the accuracy of machine learning models, addressing the challenge of balancing feature amounts in existing technologies.
Patent Information
- Application Number
- JP2022013908
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-01
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-02-01
AI Technical Summary
Existing data processing methods struggle to achieve high accuracy in classification tasks due to the challenge of balancing the number of feature amounts between different classification labels.
A data processing apparatus and method that selectively acquire and adjust the number of feature amounts for each classification label, allowing the number of selected feature amounts for one label to be between 1.1 and 2 times the number of selected feature amounts for another label, to generate machine learning models with improved accuracy.
This approach enables the generation of machine learning models with high true negative and true positive rates, significantly improving the accuracy of data processing and classification tasks.
Smart Images

Figure 0007693574000001 
Figure 0007693574000002 
Figure 0007693574000003
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a data processing apparatus, a data processing system, and a data processing method.
Background Art
[0002] For example, a machine learning model is generated based on processed data. Based on the machine learning model, classification of various events and the like are performed. Improvement in the accuracy of data processing is desired.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Embodiments of the present invention provide a data processing apparatus, a data processing system, and a data processing method capable of improving accuracy.
Means for Solving the Problems
[0005] According to an embodiment of the present invention, a data processing apparatus includes a processing unit. The processing unit can acquire a plurality of first feature amounts corresponding to a first classification label and a plurality of second feature amounts corresponding to a second classification label. The processing unit can select at least a part of the plurality of first feature amounts from the plurality of first feature amounts. It is possible to select at least a part of the plurality of second feature amounts from the plurality of second feature amounts. The processing unit can perform a first operation. In the first operation, the first number of the at least a part of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less the second number of the at least a part of the selected plurality of second feature amounts. The processing unit can generate a first machine learning model based on first training data based on the at least a part of the selected plurality of first feature amounts and the at least a part of the selected plurality of second feature amounts.
Brief Description of Drawings
[0006]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Embodiments for Carrying Out the Invention
[0007] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In this specification and each figure, the same reference numerals are assigned to the same elements as those described above with respect to the previously presented figures, and detailed descriptions thereof will be omitted as appropriate.
[0008] (First Embodiment) FIG. 1 is a schematic cross-sectional view illustrating a data processing apparatus according to the first embodiment. As shown in FIG. 1, the data processing apparatus 110 according to the embodiment includes a processing unit 70. A plurality of elements included in the data processing apparatus 110 may be provided at a plurality of different locations. The data processing apparatus 110 may be a part of a data processing system 210. The data processing system 210 may include, for example, a plurality of processing units 70. A part of the plurality of processing units 70 may be provided at a location different from another part of the plurality of processing units 70.
[0009] The processing unit 70 may include, for example, a CPU (Central Processing Unit) or the like. The processing unit 70 includes, for example, an electronic circuit or the like.
[0010] In this example, the data processing apparatus 110 includes an acquisition unit 78. The acquisition unit 78 can acquire various data, for example. The acquisition unit 78 includes, for example, an I / O port or the like. The acquisition unit 78 is an interface. The acquisition unit 78 may have the function of an output unit. The acquisition unit 78 may have, for example, a communication function.
[0011] In this example, the data processing apparatus 110 includes a storage unit 79a. The storage unit 79a can hold various data. The storage unit 79a may be, for example, a memory. The storage unit 79a may include at least one of a ROM (Read Only Memory) and a RAM (Random Access Memory).
[0012] The data processing device 110 may include a display unit 79b and an input unit 79c, etc. The display unit 79b may include various displays. The input unit 79c includes, for example, a device having an operation function (such as a keyboard, a mouse, a touch input panel, or a voice recognition input device, etc.).
[0013] Among the plurality of elements included in the data processing device 110, they can communicate with each other by at least one of wireless and wired methods. The locations where the plurality of elements included in the data processing device 110 are provided may be different from each other. As the data processing device 110, for example, a general-purpose computer may be used. As the data processing device 110, for example, a plurality of computers connected to each other may be used. As at least a part of the data processing device 110 (such as the processing unit 70, etc.), a dedicated circuit may be used. As the data processing device 110, for example, a plurality of circuits connected to each other may be used.
[0014] Hereinafter, an example of the operation of the processing unit 70 in the data processing device 110 (for example, the data processing system 210) will be described.
[0015] FIG. 2(a) and FIG. 2(b) are flowcharts illustrating the operation of the data processing device according to the first embodiment. These figures illustrate the operation of the processing unit 70. These figures show an example of the learning operation performed by the processing unit 70.
[0016] As shown in FIG. 2(a), the processing unit 70 can acquire data (step S10). For example, data is supplied to the acquisition unit 78 (such as an I / O port, see FIG. 1). The data acquired by the acquisition unit 78 is supplied to the processing unit 70. The data includes, for example, a plurality of first feature quantities corresponding to a first classification label and a plurality of second feature quantities corresponding to a second classification label. The first classification label is, for example, a first class classification label. The second classification label is a second class classification label. The plurality of first feature quantities are, for example, a plurality of first feature quantity vectors. The plurality of second feature quantities are, for example, a plurality of second feature quantity vectors. Each of the plurality of first feature quantities may include a plurality of elements. Each of the plurality of second feature quantities may include a plurality of elements.
[0017] As shown in FIG. 2(a), the processing unit 70 can select at least a part of the plurality of first feature quantities from the plurality of first feature quantities and can select at least a part of the plurality of second feature quantities from the plurality of second feature quantities (step S20).
[0018] When selecting, the processing unit 70 can perform a first operation OP1. In the first operation OP1, the above-mentioned at least a part of the first number of the selected plurality of first feature quantities is 1.1 times or more and 2 times or less the above-mentioned at least a part of the second number of the selected plurality of second feature quantities.
[0019] As shown in FIG. 2(a), the processing unit 70 can generate a first machine learning model based on the first teacher data (step S30). The first teacher data is based on at least a part of the selected plurality of first feature quantities and at least a part of the selected plurality of second feature quantities.
[0020] For example, the data processing device 110 targets data related to a plurality of events. The plurality of target events includes, for example, a plurality of first events corresponding to a first classification label and a plurality of second events corresponding to a second classification label. The plurality of first events correspond to, for example, normal products (good products) of the object. The plurality of second events correspond to, for example, non-normal products (defective products) of the object.
[0021] For example, various data related to objects classified as normal products correspond to a plurality of first feature amounts. For example, various data related to objects classified as abnormal products correspond to a plurality of second feature amounts. Such a plurality of first feature amounts and a plurality of second feature amounts are used as teacher data to generate a machine learning model.
[0022] Generally, when generating teacher data, the number of a plurality of first feature amounts is made the same as the number of a plurality of second feature amounts. Using the same number of a plurality of first feature amounts and a plurality of second feature amounts, for example, by adjusting hyperparameters and the like, a machine learning model is generated. The generation of the machine learning model corresponds to, for example, the derivation of a discrimination function.
[0023] As will be described later, according to the inventors' study, it has been found that when using the same number of a plurality of first feature amounts and a plurality of second feature amounts, it is difficult to generate a machine learning model with high accuracy. For example, it has been found that even if hyperparameters and the like are adjusted, it is difficult to derive a discrimination function with high accuracy.
[0024] In the embodiment, the same number of a plurality of first feature amounts and a plurality of second feature amounts are not used. In the embodiment, the number of a plurality of first feature amounts is made different from the number of a plurality of second feature amounts. At least a part of the plurality of first feature amounts is selected and at least a part of the plurality of second feature amounts is selected so as to have different numbers. In other words, a part of the acquired data (the plurality of first feature amounts before selection and the plurality of second feature amounts before selection) is not used as teacher data.
[0025] In the first operation OP1, the above-mentioned at least a part of the first number of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less the above-mentioned at least a part of the second number of the selected plurality of second feature amounts.
[0026] In this way, by using data with different numbers as teacher data, a machine learning model with high accuracy can be obtained. For example, a discrimination function with high accuracy can be obtained. According to the embodiment, for example, a data processing device and a data processing system capable of improving accuracy can be provided.
[0027] As described above, the plurality of target events include, for example, a plurality of first events corresponding to the first classification label and a plurality of second events corresponding to the second classification label. For example, the first occurrence rate of the plurality of first events among the plurality of events is higher than the second occurrence rate of the plurality of second events among the plurality of events. For example, the first occurrence rate of the first event (normal product) is higher than the second occurrence rate of the second event (abnormal product).
[0028] For example, in such a situation, for example, the first operation OP1 is performed. That is, in the first operation OP1, the first occurrence rate of the plurality of first events is higher than the second occurrence rate of the number of second events among the plurality of events. In the first operation OP1, the first number (the number of the selected plurality of first feature amounts) is larger than the second number (the number of the selected plurality of second feature amounts). Such a first operation OP1 can obtain high accuracy.
[0029] In one example, in the first operation OP1, the first occurrence rate corresponds to the occurrence rate of normal products of the object, and the second occurrence rate corresponds to the occurrence rate of abnormal products of the object.
[0030] According to the first machine learning model by the above first operation OP1, for example, "normal" can be determined as "normal" with high accuracy. For example, a high-accuracy true negative rate (TN: True Negative) can be obtained.
[0031] As shown in FIG. 2(a), the processing unit 70 may perform feature scaling (step S25). For example, the first training data is based on a plurality of amounts obtained by performing feature scaling processing on the plurality of first feature amounts and a plurality of amounts obtained by performing feature scaling processing on the plurality of second feature amounts. The plurality of amounts are, for example, a plurality of vectors. Based on the plurality of amounts obtained by the feature scaling processing, generation of the first machine learning model (step S30) is performed. The feature scaling processing may include, for example, at least one of normalization and standardization.
[0032] As shown in FIG. 2(a), the generation of the first machine learning model (step S30) may include a mapping operation (step S31) to a feature space of at least a part of the selected plurality of first feature amounts (for example, an amount subjected to feature amount scaling processing may also be used) and at least a part of the plurality of second feature amounts (for example, an amount subjected to feature amount scaling processing may also be used).
[0033] The mapping operation may include, for example, an operation of at least any one of a kernel function and a neural network function. The mapping operation may include, for example, at least any one of a kernel function, t-SNE (t-Distributed Stochastic Neighbor Embedding), and UMAP (Uniform Manifold Approximation and Projection).
[0034] The above kernel function may include, for example, at least any one of a linear kernel, a polynomial kernel, a Gaussian kernel, a sigmoid kernel, a Laplace kernel, and a Matern kernel.
[0035] As shown in FIG. 2(a), the generation of the first machine learning model (step S30) may include the derivation of a first discriminant function (step S32) of the amount after the mapping operation. The first discriminant function is a discriminant function regarding the first classification label and the second classification label.
[0036] The derivation of the first discriminant function may be based on, for example, at least any one of SVM (Support Vector Machine), neural network (NN), SDG (Stochastic Gradient Descent) Classifier, kNN (k-Nearest Neighbor) Classifier, and Naive Bayes classifier. For example, the first discriminant function can be derived by at least any one of SVM and NN.
[0037] In the data processing device 110 (and the data processing system 210), an operation different from the above first operation OP1 may be performed.
[0038] As shown in FIG. 2(b), the processing unit 70 can acquire data (step S10A). The data includes, for example, a plurality of third feature amounts corresponding to the first classification label and a plurality of fourth feature amounts corresponding to the second classification label. The plurality of third feature amounts are, for example, a plurality of third feature amount vectors. The plurality of fourth feature amounts are, for example, a plurality of fourth feature amount vectors. Each of the plurality of third feature amounts may include a plurality of elements. Each of the plurality of fourth feature amounts may include a plurality of elements. Thus, the processing unit 70 can further acquire a plurality of third feature amounts corresponding to the first classification label and a plurality of fourth feature amounts corresponding to the second classification label (step S10A).
[0039] As shown in FIG. 2(b), the processing unit 70 can select at least a part of the plurality of third feature amounts from the plurality of third feature amounts and can select at least a part of the plurality of fourth feature amounts from the plurality of fourth feature amounts (step S20A). At this time, the processing unit 70 can perform the second operation OP2. In the second operation OP2, the above-mentioned at least a part of the third number of the selected plurality of third feature amounts is 0.1 times or more and 0.9 times or less of the above-mentioned at least a part of the fourth number of the selected plurality of fourth feature amounts. Thus, in the second operation OP2, the number (third number) of the selected plurality of third feature amounts corresponding to the first classification label is smaller than the number (fourth number) of the selected plurality of fourth feature amounts corresponding to the second classification label.
[0040] The processing unit 70 can further generate a second machine learning model based on the second teacher data (step S30A). The second teacher data is based on at least a part of the selected plurality of third feature amounts and at least a part of the selected plurality of fourth feature amounts.
[0041] For example, the plurality of target events includes a plurality of third events corresponding to the first classification label and a plurality of fourth events corresponding to the second classification label. The plurality of third events correspond to, for example, normal products. The plurality of fourth events correspond to, for example, non-normal products.
[0042] In the second operation OP2, the occurrence rate (third occurrence rate) of a plurality of third events in a plurality of events is lower than, for example, the occurrence rate (fourth occurrence rate) of a plurality of fourth events in the plurality of events. For example, in the second operation OP2, the third occurrence rate corresponds to the occurrence rate of normal products of the object. The fourth occurrence rate corresponds to the occurrence rate of non-normal products of the object.
[0043] For example, in the initial stage of production, the occurrence rate of normal products may be lower than the occurrence rate of non-normal products. In such a case, the number (third number) of a plurality of third feature amounts corresponding to normal products (third events) with a low occurrence rate is made smaller than the number (fourth number) of a plurality of fourth feature amounts corresponding to non-normal products (fourth events) with a high occurrence rate. Thereby, a machine learning model with higher accuracy can be generated.
[0044] According to the second machine learning model by the second operation OP2 above, for example, "abnormality" can be determined as "abnormality" with high accuracy. For example, a high-accuracy true positive rate (TP: True Positive) can be obtained.
[0045] As shown in FIG. 2(b), the processing unit 70 may perform feature scaling (step S25A). For example, the second training data is based on a plurality of amounts obtained by performing feature scaling processing on the plurality of third feature amounts and a plurality of amounts obtained by performing feature scaling processing on the plurality of fourth feature amounts. The plurality of amounts are, for example, a plurality of vectors. Based on the plurality of amounts obtained by the feature scaling processing, generation of the second machine learning model (step S30A) is performed. The feature scaling processing may include, for example, at least one of normalization and standardization.
[0046] As shown in FIG. 2(b), the generation of the second machine learning model (step S30A) may include a mapping operation (step S31A) to a feature space of at least a part of the selected plurality of third feature amounts (which may be, for example, amounts subjected to feature scaling processing) and at least a part of the plurality of fourth feature amounts (which may be, for example, amounts subjected to feature scaling processing).
[0047] The mapping operation may include, for example, at least one of the operations of a kernel function and a neural network function. The mapping operation may include at least one of a kernel function, t-SNE (t-Distributed Stochastic Neighbor Embedding), and UMAP (Uniform Manifold Approximation and Projection).
[0048] The above kernel function may include, for example, at least one of a linear kernel, a polynomial kernel, a Gaussian kernel, a sigmoid kernel, a Laplace kernel, and a Matern kernel.
[0049] As shown in FIG. 2(b), the generation of the second machine learning model (step S30A) may include the derivation of the second discriminant function of the quantity after the mapping operation (step S32A). The second discriminant function is a discriminant function regarding the first classification label and the second classification label.
[0050] The derivation of the second discriminant function may be based on, for example, at least one of SVM (Support Vector Machine), neural network (NN), SDG (Stochastic Gradient Descent) Classifier, kNN (k-Nearest Neighbor) Classifier, and a naive Bayes classifier. For example, the second discriminant function can be derived by at least one of SVM and NN.
[0051] At least one of the generation of the first machine learning model and the generation of the second machine learning model may include the adjustment of hyperparameters.
[0052] Such first operation OP1 and second operation OP2 may be switched and executed.
[0053] FIG. 3 is a flowchart illustrating the operation of the data processing apparatus according to the first embodiment. FIG. 3 shows an example of another operation performed by the processing unit 70. FIG. 3 illustrates, for example, a classification operation (or prediction operation).
[0054] The processing unit 70 is further capable of performing a classification operation. In the classification operation, it is possible to acquire another data (another feature amount) (step S50). The another feature amount is a new feature amount acquired separately from the learning operation. The another feature amount is, for example, an unknown feature amount. The another feature amount is, for example, another feature vector. The another feature amount may include, for example, a plurality of elements.
[0055] The processing unit 70 classifies the another feature amount into a first classification label or a second classification label based on the first discrimination function derived in the learning operation (step S60). Thus, the processing unit 70 can classify the another feature amount into the first classification label or the second classification label based on the first machine learning model in the classification operation.
[0056] As shown in FIG. 3, the above another feature amount may be obtained by the processing unit 70 performing feature scaling on the newly obtained data (step S65).
[0057] In the embodiment, a new another feature amount is classified by a machine learning model (for example, a discrimination function) based on the teacher data according to the above first operation OP1 or second operation OP2. High-precision classification is possible.
[0058] In an embodiment, the plurality of first feature amounts and the plurality of second feature amounts may be feature amounts related to the characteristics of a magnetic recording device. For example, the "other feature amount" in the classification operation may be a feature amount related to the characteristics of a magnetic recording device. The feature amounts related to the characteristics of a magnetic recording device may include, for example, at least any one of Signal-Noise Ratio (SNR), Bit Error Rate (BER), Fringe BER, Erase Width at AC erase (EWAC), Magnetic write track width (MWW), Overwrite (OW), and Soft Viterbi Algorithm-BER (SOVA-BER), Viterbi Metric Margin (VMM), Repeatable RunOut (RRO), and Non-Repeatable RunOut (NRRO).
[0059] For example, in a magnetic recording device, there is a magnetic head in which recording characteristic defects occur. Based on test data regarding the magnetic head, it is desired to predict the characteristics of the magnetic head with high accuracy. Machine learning is used for such prediction. In a general machine learning prediction model, machine learning is performed using, as teacher data, data regarding normal products and data regarding abnormal products that have the same number as each other. Then, the characteristics (performance) of the prediction model are adjusted by hyperparameter tuning.
[0060] In the embodiment, as described above, the number of data regarding normal products is different from the number of data regarding abnormal products. A machine learning model using such data as teacher data is generated. Thereby, high-accuracy prediction becomes possible.
[0061] Hereinafter, examples of characteristics in a data processing device will be described. FIGS. 4(a) to 4(c) are graphs illustrating the characteristics of a data processing device. The horizontal axes of these figures correspond to the numbers N0 (names) of a plurality of data. The horizontal axes of these figures correspond to, for example, the adjusted values of hyperparameters. The vertical axes of these figures correspond to the evaluation parameter P1. These figures relate to the true negative rate (TN). The fact that the evaluation parameter P1 is 1 corresponds to all normal products being correctly judged as normal. When the evaluation parameter P1 is less than 1, it corresponds to the occurrence of false positives (FP: False Positive, misjudging normal as abnormal).
[0062] In FIG. 4(a), the first number is 0.5 times the second number. As already explained, the first number is the number of at least a part of the selected plurality of first feature amounts. The second number is the number of at least a part of the selected plurality of second feature amounts.
[0063] In FIG. 4(b), the first number is the same as the second number. In FIG. 4(c), the first number is 2 times the second number. FIG. 4(b) corresponds to the true negative rate (TN) obtained by hyperparameter adjustment in general machine learning. FIG. 4(c) corresponds to the case where the evaluation parameter P1 becomes 1 and a prediction model without the occurrence of FP can be constructed.
[0064] As shown in FIG. 4(c), when the first number is 2 times the second number, by increasing the adjusted value of the hyperparameter, the evaluation parameter P1 becomes 1. Normal products are correctly judged as normal without the occurrence of FP.
[0065] On the other hand, as shown in FIG. 4(b), when the first number is the same as the second number, the evaluation parameter P1 is about 0.7. When the first number is the same as the second number, it is difficult to obtain high accuracy even by adjusting the hyperparameter.
[0066] In an embodiment, when the first number is greater than the second number (e.g., twice), it is considered that an evaluation parameter P1 of 1 is obtained based on the following. For example, when the prediction model misjudges data regarding normal products, the loss function is likely to increase. When the first number is greater than the second number, compared with the case where the first number is the same as the second number, the degree to which the correct answer rate of normal products contributes to loss reduction becomes greater. Thus, it is considered that an evaluation parameter P1 of 1 is obtained when the first number is greater than the second number (e.g., twice).
[0067] Figures 5(a) to 5(c) are graphs illustrating the characteristics of the data processing device. The horizontal axes of these figures correspond to the numbers N0 (names) of a plurality of data. The horizontal axes of these figures correspond to, for example, the adjusted values of hyperparameters. The vertical axes of these figures correspond to the evaluation parameter P2. These figures relate to the true positive rate (TP). The fact that the evaluation parameter P2 is 1 corresponds to the non-normal product being correctly judged as non-normal. When the evaluation parameter P2 is less than 1, it corresponds to the occurrence of false negatives.
[0068] In Figure 5(a), the first number is 0.5 times the second number. In Figure 5(c), the first number is the same as the second number. In Figure 5(b), the first number is twice the second number. Figure 5(c) corresponds to the true positive rate (TP) obtained by hyperparameter adjustment in general machine learning.
[0069] As shown in Figure 5(a), when the first number is 0.5 times the second number, by combining with the adjustment of hyperparameters, the evaluation parameter P2 becomes 1. The non-normal product is correctly judged as non-normal.
[0070] On the other hand, as shown in Figure 5(c), when the first number is the same as the second number, the maximum value of the evaluation parameter P2 is about 0.7 to 0.8. When the first number is the same as the second number, it is difficult to obtain high accuracy even by adjusting the hyperparameters.
[0071] For example, when the first number is the same as the second number, in hyperparameter adjustment, for both the true negative rate (TN) and the true positive rate (TP), the parameters P1 and P2 are about 0.6 to 0.8. By making the first number and the second number different from each other, a high-precision true negative rate (TN) or true positive rate (TP) can be obtained.
[0072] In an embodiment, when the first number is smaller than the second number (for example, 0.5 times), it is considered that the evaluation parameter P2 of 1 is obtained based on the following. For example, when the prediction model misjudges data related to abnormal products, the loss function is likely to increase. When the first number is smaller than the second number, it is considered that the contribution of the correct answer rate of abnormal products to loss reduction is greater than when the first number is the same as the second number. Thus, it is considered that the evaluation parameter P2 of 1 is obtained when the first number is smaller than the second number.
[0073] FIG. 6(a) and FIG. 6(b) are graphs illustrating the characteristics of the data processing device. Let the first number be N1. Let the second number be N2. FIG. 6 illustrates the characteristics when the ratio (N1 / N2) of the first number to the second number is changed. The horizontal axis of these figures is the ratio (N1 / N2). The vertical axis of FIG. 6(a) is the parameter CN1. The parameter CN1 is the average value of the true negative rate (TN) within the range of effective hyperparameters that do not cause overfitting. The vertical axis of FIG. 6(b) is the parameter CP1. The parameter CP1 is the average value of the true positive rate (TP) within the range of effective hyperparameters that do not cause overfitting.
[0074] For example, when the occurrence rate of normal products is high, the true negative rate (TN) is preferably 0.9 or more. Thereby, for example, it becomes easier to improve the yield after failure detection by machine learning. As shown in FIG. 6(a), when the ratio (N1 / N2) is 1.1 or more and 2.0 or less, a high parameter CN1 of 0.9 or more is obtained. In one example according to the embodiment, the first number is preferably 1.1 times or more and 2 times or less the second number. A high true negative rate (TN) of 0.9 or more is obtained.
[0075] For example, when the occurrence rate of defective products is high, the true positive rate (TP) is preferably 0.9 or more. As shown in FIG. 6(b), when the ratio (N1 / N2) is 0.1 or more and 0.9 or less, a parameter CP1 of 0.9 or more can be obtained. In one example according to the embodiment, the first number is preferably 0.1 times or more and 0.9 times or less of the second number. A high true positive rate (TP) of 0.9 or more can be obtained.
[0076] In general machine learning (reference example), the first number is the same as the second number. In this case, both the true positive rate (TP) and the true negative rate (TN) are about 0.7 to 0.8. The reference example in which the first number is the same as the second number is considered to be suitably applied when the occurrence rate of normal products is about the same as the occurrence rate of defective products.
[0077] For example, when the occurrence rate of normal products is 1000 times or more the occurrence rate of defective products, it is considered good to apply a ratio (N1 / N2) of 1.1 or more and 2.0 or less. For example, when the occurrence rate of normal products is less than 1000 times the occurrence rate of defective products, it is considered good to apply a ratio (N1 / N2) of 0.1 or more and 0.9 or less.
[0078] In the embodiment, for example, a machine learning model is generated using, as teacher data, a plurality of data including a set including a class classification label and a feature vector with feature scaling. At this time, the number of first feature vectors corresponding to the first class is made different from the number of the plurality of second feature vectors corresponding to the second class. For example, the plurality of feature vectors may be linearly mapped or non-linearly mapped in the feature space. Using the discriminant function generated in the generated machine learning model, the class classification of another data (another feature amount) is predicted. Such an operation is performed in the processing unit 70.
[0079] The data processing apparatus 110 (and the data processing system 210) according to the embodiment can be applied to, for example, a classification problem (failure prediction) by machine learning. In the embodiment, the number of data serving as teacher data is made different among classes. The ratio of the number among classes is not 1:1. The ratio of the number among classes is adjusted. Thereby, the true positive rate and the true negative rate of the prediction model can be adjusted. In the embodiment, the true positive rate and the true negative rate of the prediction model may be adjusted by hyperparameter adjustment. In the embodiment, a high-precision true positive rate and true negative rate that cannot be obtained only by hyperparameter adjustment can be obtained.
[0080] The data processing system 210 (see FIG. 1) according to the embodiment includes one or more processing units 70 (see FIG. 1). The processing unit 70 in the data processing system 210 can perform the above-described operations described with respect to the data processing apparatus 110.
[0081] (Second Embodiment) The second embodiment relates to a program. The program causes the processing unit 70 (computer) to acquire a plurality of first feature amounts corresponding to a first classification label and a plurality of second feature amounts corresponding to a second classification label. The program causes the processing unit 70 to select at least a part of the plurality of first feature amounts from the plurality of first feature amounts and to select at least a part of the plurality of second feature amounts from the plurality of second feature amounts. The program causes the processing unit 70 to perform a first operation OP1. In the first operation, the above-described at least a part of the first number of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less the above-described at least a part of the second number of the selected plurality of second feature amounts. The program causes the processing unit 70 to generate a first machine learning model based on the first teacher data based on at least a part of the selected plurality of first feature amounts and at least a part of the selected plurality of second feature amounts.
[0082] The embodiment may include a storage medium storing the above program.
[0083] (Third Embodiment) The third embodiment relates to a data processing method. The data processing method causes the processing unit 70 to acquire a plurality of first feature amounts corresponding to a first classification label and a plurality of second feature amounts corresponding to a second classification label. The data processing method causes the processing unit 70 to select at least a part of the plurality of first feature amounts from the plurality of first feature amounts and to select at least a part of the plurality of second feature amounts from the plurality of second feature amounts. The data processing method causes the processing unit 70 to perform a first operation OP1. In the first operation OP1, the above-described at least a part of the first number of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less the above-described at least a part of the second number of the selected plurality of second feature amounts. The data processing method causes the processing unit 70 to generate a first machine learning model based on first teacher data based on the above-described at least a part of the selected plurality of first feature amounts and the above-described at least a part of the selected plurality of second feature amounts.
[0084] The embodiment may include the following configuration (for example, a technical solution). (Configuration 1) Comprising a processing unit, The processing unit can acquire a plurality of first feature amounts corresponding to a first classification label and a plurality of second feature amounts corresponding to a second classification label, The processing unit can select at least a part of the plurality of first feature amounts from the plurality of first feature amounts, can select at least a part of the plurality of second feature amounts from the plurality of second feature amounts, the processing unit can perform a first operation, and in the first operation, the above-described at least a part of the first number of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less the above-described at least a part of the second number of the selected plurality of second feature amounts, A data processing device in which the processing unit can generate a first machine learning model based on first teacher data based on the above-described at least a part of the selected plurality of first feature amounts and the above-described at least a part of the selected plurality of second feature amounts.
[0085] (Configuration 2) The processing unit can further obtain a plurality of third feature amounts corresponding to the first classification label and a plurality of fourth feature amounts corresponding to the second classification label. The processing unit can select at least a part of the plurality of third feature amounts from the plurality of third feature amounts, and can select at least a part of the plurality of fourth feature amounts from the plurality of fourth feature amounts. The processing unit can perform a second operation. In the second operation, the third number of the at least a part of the plurality of selected third feature amounts is 0.1 times or more and 0.9 times or less of the fourth number of the at least a part of the plurality of selected fourth feature amounts. The data processing apparatus according to Configuration 1, wherein the processing unit can further generate a second machine learning model based on second teacher data based on the at least a part of the plurality of selected third feature amounts and the at least a part of the plurality of selected fourth feature amounts.
[0086] (Configuration 3) The plurality of target events include a plurality of third events corresponding to the first classification label and a plurality of fourth events corresponding to the second classification label. In the second operation, the third occurrence rate of the plurality of third events in the plurality of events is lower than the fourth occurrence rate of the plurality of fourth events in the plurality of events. The data processing apparatus according to Configuration 2.
[0087] (Configuration 4) In the second operation, the third occurrence rate corresponds to the occurrence rate of normal products of the object, and the fourth occurrence rate corresponds to the occurrence rate of non-normal products of the object. The data processing apparatus according to Configuration 3.
[0088] (Configuration 5) The plurality of target events include a plurality of first events corresponding to the first classification label and a plurality of second events corresponding to the second classification label. In the first operation, the first occurrence rate of the plurality of first events in the plurality of events is higher than the second occurrence rate of the plurality of second events in the plurality of events. The data processing apparatus according to Configuration 1 or 2.
[0089] (Configuration 6) In the first operation, the first occurrence rate corresponds to the occurrence rate of normal products of the object, and the second occurrence rate corresponds to the occurrence rate of abnormal products of the object. The data processing apparatus according to Configuration 5.
[0090] (Configuration 7) The first teacher data is based on a plurality of amounts obtained by performing feature amount scaling processing on the plurality of first feature amounts, and a plurality of amounts obtained by performing feature amount scaling processing on the plurality of second feature amounts. The data processing apparatus according to any one of Configurations 1 to 6.
[0091] (Configuration 8) The feature amount scaling processing includes at least one of normalization and standardization. The data processing apparatus according to Configuration 7.
[0092] (Configuration 9) The generation of the first machine learning model includes a mapping operation of at least a part of the selected plurality of first feature amounts and at least a part of the plurality of second feature amounts into a feature space. The data processing apparatus according to any one of Configurations 1 to 8.
[0093] (Configuration 10) The mapping operation includes at least one of a kernel function, t-SNE (t-Distributed Stochastic Neighbor Embedding), and UMAP (Uniform Manifold Approximation and Projection). The data processing apparatus according to Configuration 9.
[0094] (Configuration 11) The kernel function includes at least one of a linear kernel, a polynomial kernel, a Gaussian kernel, a sigmoid kernel, a Laplace kernel, and a Matern kernel. The data processing apparatus according to Configuration 10.
[0095] (Configuration 12) The generation of the first machine learning model includes deriving a first discriminant function for the first classification label and the second classification label from the quantity after the mapping operation, for the data processing apparatus according to Configuration 9 or 10.
[0096] (Configuration 13) The derivation of the first discriminant function is based on at least any one of SVM (Support Vector Machine), neural network (NN), SDG (Stochastic Gradient Descent) Classifier, kNN (k-Nearest Neighbor) Classifier, and Naive Bayes classifier, for the data processing apparatus according to Configuration 12.
[0097] (Configuration 14) The processing unit is further capable of a classification operation, In the classification operation, the processing unit classifies another feature amount into the first classification label or the second classification label based on the first discriminant function, for the data processing apparatus according to Configuration 12 or 13.
[0098] (Configuration 15) The processing unit is further capable of a classification operation, In the classification operation, the processing unit classifies another feature amount into the first classification label or the second classification label based on the first machine learning model, for the data processing apparatus according to any one of Configurations 1 to 13.
[0099] (Configuration 16) The another feature amount is obtained by the processing unit performing feature scaling on new data obtained, for the data processing apparatus according to Configuration 14 or 15.
[0100] (Configuration 17) The generation of the first machine learning model includes adjusting hyperparameters, for the data processing apparatus according to any one of Configurations 1 to 16.
[0101] (Configuration 18) The plurality of first feature amounts and the plurality of second feature amounts are for a data processing apparatus according to any one of Configurations 1 to 17 regarding the characteristics of a magnetic recording apparatus.
[0102] (Configuration 19) including one or more processing units, the processing unit is capable of acquiring a plurality of first feature amounts corresponding to a first classification label and a plurality of second feature amounts corresponding to a second classification label, the processing unit can select at least a part of the plurality of first feature amounts from the plurality of first feature amounts, can select at least a part of the plurality of second feature amounts from the plurality of second feature amounts, the processing unit can perform a first operation, and in the first operation, the first number of the at least a part of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less the second number of the at least a part of the selected plurality of second feature amounts, the processing unit can generate a first machine learning model based on first teacher data based on the at least a part of the selected plurality of first feature amounts and the at least a part of the selected plurality of second feature amounts, a data processing system.
[0103] (Configuration 20) causing the processing unit to acquire a plurality of first feature amounts corresponding to a first classification label and a plurality of second feature amounts corresponding to a second classification label, causing the processing unit to select at least a part of the plurality of first feature amounts from the plurality of first feature amounts, causing the processing unit to select at least a part of the plurality of second feature amounts from the plurality of second feature amounts, causing the processing unit to perform a first operation, and in the first operation, the first number of the at least a part of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less the second number of the at least a part of the selected plurality of second feature amounts, a storage medium storing a program for causing the processing unit to generate a first machine learning model based on first teacher data based on the at least a part of the selected plurality of first feature amounts and the at least a part of the selected plurality of second feature amounts.
[0104] (Configuration 21) Cause the processing unit to obtain a plurality of first feature amounts corresponding to the first classification label and a plurality of second feature amounts corresponding to the second classification label. Cause the processing unit to select at least a part of the plurality of first feature amounts from the plurality of first feature amounts, cause the processing unit to select at least a part of the plurality of second feature amounts from the plurality of second feature amounts, cause the processing unit to perform a first operation, and in the first operation, the first number of the at least a part of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less the second number of the at least a part of the selected plurality of second feature amounts. A program for causing the processing unit to generate a first machine learning model based on first teacher data based on the at least a part of the selected plurality of first feature amounts and the at least a part of the selected plurality of second feature amounts.
[0105] (Configuration 22) Cause the processing unit to obtain a plurality of first feature amounts corresponding to the first classification label and a plurality of second feature amounts corresponding to the second classification label. Cause the processing unit to select at least a part of the plurality of first feature amounts from the plurality of first feature amounts, cause the processing unit to select at least a part of the plurality of second feature amounts from the plurality of second feature amounts, cause the processing unit to perform a first operation, and in the first operation, the first number of the at least a part of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less the second number of the at least a part of the selected plurality of second feature amounts. A data processing method for causing the processing unit to generate a first machine learning model based on first teacher data based on the at least a part of the selected plurality of first feature amounts and the at least a part of the selected plurality of second feature amounts.
[0106] According to the embodiment, a data processing device, a data processing system, and a data processing method capable of improving accuracy can be provided.
[0107] The embodiments of the present invention have been described above with reference to examples. However, the present invention is not limited to these examples. For example, regarding the configuration such as the processing unit included in the data processing apparatus, the present invention can be similarly implemented by appropriately selecting from the range known to those skilled in the art, and as long as the same effects can be obtained, it is included in the scope of the present invention.
[0108] Combinations of any two or more elements of each example within the technically possible range are also included in the scope of the present invention as long as they encompass the gist of the present invention.
[0109] Based on the data processing apparatus, data processing system, and data processing method described above as embodiments of the present invention, all data processing apparatuses, data processing systems, and data processing methods that can be appropriately designed and modified by those skilled in the art are also within the scope of the present invention as long as they encompass the gist of the present invention.
[0110] Those skilled in the art can conceive of various modification examples and correction examples within the scope of the idea of the present invention, and it is understood that those modification examples and correction examples also belong to the scope of the present invention.
[0111] Although some embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, and are also included in the invention described in the claims and its equivalent scope.
Explanation of Reference Numerals
[0112] 70... Processing unit, 78... Acquisition unit, 79a... Storage unit, 79b... Display unit, 79c... Input unit, 110... Data processing apparatus, 210... Data processing system, CN1, CP1... Parameters, N0... Number, OP1, OP2... First, second operations, P1, P2... First, second evaluation parameters
Claims
1. comprising a processing unit, the processing unit is capable of acquiring a plurality of first feature amounts corresponding to a first classification label and a plurality of second feature amounts corresponding to a second classification label, the processing unit can select at least a part of the plurality of first feature amounts from the plurality of first feature amounts, and can select at least a part of the plurality of second feature amounts from the plurality of second feature amounts, the processing unit is capable of performing a first operation, and in the first operation, a first number of a plurality of elements included in at least a part of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less of a second number of a plurality of elements included in at least a part of the selected plurality of second feature amounts, the processing unit can generate a first machine learning model based on first teacher data based on at least a part of the selected plurality of first feature amounts and at least a part of the selected plurality of second feature amounts, the processing unit can further acquire a plurality of third feature amounts corresponding to the first classification label and a plurality of fourth feature amounts corresponding to the second classification label, the processing unit can select at least a part of the plurality of third feature amounts from the plurality of third feature amounts, and can select at least a part of the plurality of fourth feature amounts from the plurality of fourth feature amounts, the processing unit is capable of performing a second operation, and in the second operation, a third number of a plurality of elements included in at least a part of the selected plurality of third feature amounts is 0.1 times or more and 0.9 times or less of a fourth number of a plurality of elements included in at least a part of the selected plurality of fourth feature amounts, the processing unit can further generate a second machine learning model based on second teacher data based on at least a part of the selected plurality of third feature amounts and at least a part of the selected plurality of fourth feature amounts, a data processing device.
2. A plurality of target events include a plurality of third events corresponding to the first classification label and a plurality of fourth events corresponding to the second classification label, In the second operation, a third occurrence rate of the plurality of third events in the plurality of events is lower than a fourth occurrence rate of the plurality of fourth events in the plurality of events. The data processing apparatus according to claim 1.
3. In the second operation, the third occurrence rate corresponds to an occurrence rate of normal products of the object, and the fourth occurrence rate corresponds to an occurrence rate of non-normal products of the object. The data processing apparatus according to claim 2.
4. A plurality of target events include a plurality of first events corresponding to the first classification label and a plurality of second events corresponding to the second classification label. In the first operation, a first occurrence rate of the plurality of first events in the plurality of events is higher than a second occurrence rate of the plurality of second events in the plurality of events. The data processing apparatus according to claim 1 or 2.
5. It includes a processing unit. The processing unit can acquire a plurality of first feature amounts corresponding to a first classification label and a plurality of second feature amounts corresponding to a second classification label. The processing unit can select at least a part of the plurality of first feature amounts from the plurality of first feature amounts, and can select at least a part of the plurality of second feature amounts from the plurality of second feature amounts. The processing unit can perform a first operation. In the first operation, a first number of a plurality of elements included in at least a part of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less of a second number of a plurality of elements included in at least a part of the selected plurality of second feature amounts. The processing unit can generate a first machine learning model based on first teacher data based on at least a part of the selected plurality of first feature amounts and at least a part of the selected plurality of second feature amounts. A plurality of target events include a plurality of first events corresponding to the first classification label and a plurality of second events corresponding to the second classification label. In the first operation, a data processing apparatus in which a first occurrence rate of the plurality of first events among the plurality of events is higher than a second occurrence rate of the plurality of second events among the plurality of events.
6. The data processing apparatus according to claim 4 or 5, wherein in the first operation, the first occurrence rate corresponds to an occurrence rate of normal products of the object, and the second occurrence rate corresponds to an occurrence rate of non-normal products of the object.
7. The first teacher data is based on a plurality of amounts obtained by performing feature amount scaling processing on the plurality of first feature amounts and a plurality of amounts obtained by performing feature amount scaling processing on the plurality of second feature amounts, and the data processing apparatus according to any one of claims 1 to 6.
8. Comprising a processing unit, The processing unit can acquire a plurality of first feature amounts corresponding to a first classification label and a plurality of second feature amounts corresponding to a second classification label, The processing unit can select at least a part of the plurality of first feature amounts from the plurality of first feature amounts, and can select at least a part of the plurality of second feature amounts from the plurality of second feature amounts. The processing unit can perform a first operation. In the first operation, a first number of a plurality of elements included in at least a part of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less of a second number of a plurality of elements included in at least a part of the selected plurality of second feature amounts. The processing unit can generate a first machine learning model based on first teacher data based on at least a part of the selected plurality of first feature amounts and at least a part of the selected plurality of second feature amounts. The first teacher data is based on a plurality of amounts obtained by performing feature amount scaling processing on the plurality of first feature amounts and a plurality of amounts obtained by performing feature amount scaling processing on the plurality of second feature amounts, and a data processing apparatus.
9. The processing unit is further capable of a classification operation, In the classification operation, the processing unit classifies another feature amount into the first classification label or the second classification label based on the first machine learning model. The data processing apparatus according to any one of claims 1 to 8.
10. Comprising one or more processing units, The processing unit can acquire a plurality of first feature amounts corresponding to a first classification label and a plurality of second feature amounts corresponding to a second classification label. The processing unit can select at least a part of the plurality of first feature amounts from the plurality of first feature amounts, and can select at least a part of the plurality of second feature amounts from the plurality of second feature amounts. The processing unit can perform a first operation. In the first operation, a first number of a plurality of elements included in at least a part of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less of a second number of a plurality of elements included in at least a part of the selected plurality of second feature amounts. The processing unit can generate a first machine learning model based on first teacher data based on at least a part of the selected plurality of first feature amounts and at least a part of the selected plurality of second feature amounts. A plurality of target events include a plurality of first events corresponding to the first classification label and a plurality of second events corresponding to the second classification label. In the first operation, a first occurrence rate of the plurality of first events in the plurality of events is higher than a second occurrence rate of the plurality of second events in the plurality of events. A data processing system.
11. Cause the processing unit to acquire a plurality of first feature amounts corresponding to a first classification label and a plurality of second feature amounts corresponding to a second classification label. Cause the processing unit to select at least a part of the plurality of first feature amounts from the plurality of first feature amounts, and to select at least a part of the plurality of second feature amounts from the plurality of second feature amounts, cause the processing unit to perform a first operation, and in the first operation, a first number of a plurality of elements included in the at least a part of the selected plurality of first feature amounts is 1.1 times or more and 2 times or less of a plurality of elements included in a second number of the at least a part of the selected plurality of second feature amounts. Cause the processing unit to generate a first machine learning model based on first training data based on the at least a part of the selected plurality of first feature amounts and the at least a part of the selected plurality of second feature amounts. A plurality of target events includes a plurality of first events corresponding to the first classification label and a plurality of second events corresponding to the second classification label. In the first operation, a first occurrence rate of the plurality of first events in the plurality of events is higher than a second occurrence rate of the plurality of second events in the plurality of events. A data processing method.
Citation Information
Patent Citations
Information processing device, information processing method and program
JP2017102865A
Abnormality detection device, abnormality detection method and program
JP2018147172A
Device and method of generating teacher data for machine learning
JP2019028876A
Abnormality detection device and abnormality detection method
JP2020064367A
Learning device for block noise detection and computer program
JP2021149719A