Information Processing Apparatus, Computer-Readable Recording Medium, and Information Processing Method

By storing and processing feature vectors, quality labels and non-quality labels in the information processing device, and calculating clustering accuracy or their variance, the problem of difficulty in determining the cause when the data quality is not high is solved, and the improvement of data quality and performance improvement of deep learning methods are achieved.

CN114424236BActive Publication Date: 2025-06-10MITSUBISHI ELECTRIC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980100361.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-09-30
Publication Date
2025-06-10
Estimated Expiration
2039-09-30

AI Technical Summary

Technical Problem

In deep learning methods, when the data quality is not high, it is difficult to determine the reasons for the deterioration of data quality, and the prior art lacks effective methods to improve data quality.

Method used

An information processing device is designed to store feature vector sets, quality label sets and non-quality label sets, and calculate the variance of the average clustering accuracy or clustering accuracy using the non-quality label clustering unit to determine the types of non-quality labels that have adverse effects on data quality.

Benefits of technology

It can effectively determine the reasons for deterioration in data quality, improve data quality, and thus improve the performance of deep learning methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114424236B_ABST
    Figure CN114424236B_ABST
Patent Text Reader

Abstract

comprising: a storage unit (102) that stores a set of feature vectors, a set of quality labels, and a plurality of sets of non-quality labels; a non-quality label clustering unit (107) that calculates an average clustering accuracy for each of the plurality of sets of non-quality labels, thereby calculating a plurality of average clustering accuracies respectively corresponding to the plurality of sets of non-quality labels, the average clustering accuracy being an average value of clustering accuracies when clustering subsets obtained by dividing a plurality of feature vectors by each of a plurality of elements represented by the plurality of non-quality labels using the set of quality labels; and a processing unit (108) that generates a screen image capable of determining, using the plurality of average clustering accuracies, a type of at least one non-quality label that adversely affects the quality of a plurality of digital data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, a computer-readable recording medium, and an information processing method. Background Art

[0002] Due to the progress of deep learning and related technologies, systems capable of performing complex recognition tasks related to images or sounds have become common systems. In such systems, the underlying structure can be automatically discovered from a large amount of learning data, thereby achieving high generalization performance that cannot be achieved by conventional methods prior to deep learning.

[0003] However, such systems do not function in the situation where rich labeled data that can be used during learning is not available. On the other hand, in various real-world tasks, the situation of obtaining rich learning data is very rare. Therefore, in the vast majority of cases, non-conventional methods such as deep learning do not work.

[0004] For example, methods for automatically diagnosing the soundness of a device based on the sound or vibration generated by the device have been studied for a long time, and various methods have been developed so far. For example, the MT (Mahalanobis-Taguchi) method described in Non-Patent Document 1 is one of the most representative methods. In the MT method, the characteristic space of the normal sample distribution is learned in advance as a reference space, and a determination of normal or abnormal is made based on how much the feature vector observed at the time of diagnosis deviates from the reference space.

[0005] In conventional methods such as the MT method, empirical knowledge and insights are incorporated into the extraction of features, and assumptions are made regarding the distribution of feature vectors, whereby appropriate constraints can be easily imposed on the learned model. Therefore, in such methods, a large amount of data required for deep learning is not needed.

[0006] Prior Art Documents

[0007] Non-Patent Documents

[0008] Non-Patent Document 1: Kazuo Tachibana, "Introduction to Taguchi Methods", Nikkan Kogyo Shimbun, Ltd., 2004, P. 167-185 Summary of the Invention

[0009] Problems to be Solved by the Invention

[0010] However, in the conventional method, only a small amount of data is required for learning. Correspondingly, there is a problem that if the quality is not high, it does not function. However, in this field, there are very few techniques from the perspective of improving the quality of the data to be measured. In particular, there is almost no general method that does not require the knowledge inherent in the task of interest. Moreover, when the quality of the measured data is poor, the cause of the poor data quality cannot be determined.

[0011] Therefore, an object of one or more aspects of the present invention is to be able to determine the cause of the deterioration in the quality of the data set being used.

[0012] Means for solving the problem

[0013] The information processing apparatus according to the first aspect of the present invention is characterized in that the information processing apparatus includes: a storage unit that stores a set of feature vectors, a set of quality labels, and a plurality of sets of non-quality labels, the set of feature vectors including a plurality of feature vectors generated by respectively extracting predetermined features from a plurality of digital data representing measurement values measured from an object, the set of quality labels including a plurality of quality labels respectively corresponding to each of the plurality of digital data and indicating the quality of the object as good or bad, the plurality of sets of non-quality labels each including a plurality of non-quality labels of a type expected to be irrelevant to the quality of the object as good or bad, the plurality of non-quality labels respectively corresponding to each of the plurality of digital data; a non-quality label clustering unit that calculates an average clustering accuracy for each of the plurality of sets of non-quality labels, thereby calculating a plurality of the average clustering accuracies respectively corresponding to each of the plurality of sets of non-quality labels, the average clustering accuracy being an average value of clustering accuracies when clustering subsets obtained by dividing the plurality of feature vectors by each of the plurality of elements represented by the plurality of non-quality labels using the set of quality labels; and a processing unit that generates a screen image capable of using the plurality of average clustering accuracies to determine at least one type of non-quality label that has an adverse effect on the quality of the plurality of digital data.

[0014] The information processing apparatus according to the second aspect of the present invention is characterized in that the information processing apparatus includes: a storage unit that stores a feature vector set, a quality label set, and a plurality of non-quality label sets, the feature vector set including a plurality of feature vectors generated by respectively extracting predetermined features from a plurality of digital data representing measurement values measured from an object, the quality label set including a plurality of quality labels respectively corresponding to the respective digital data in the plurality of digital data and indicating the quality of the object, the plurality of non-quality label sets respectively including a plurality of non-quality labels of a type expected to be irrelevant to the quality of the object, the plurality of non-quality labels respectively corresponding to the respective digital data in the plurality of digital data; a non-quality label clustering unit that calculates a clustering accuracy for a non-quality label set corresponding to one type of non-quality label selected from the plurality of non-quality labels, thereby calculating a plurality of the clustering accuracies, the clustering accuracy being the clustering accuracy when clustering subsets obtained by dividing the plurality of feature vectors by respective elements represented by the plurality of non-quality labels using the quality label set; and a processing unit that generates a screen image capable of determining at least one element that adversely affects the quality of the plurality of digital data using the plurality of clustering accuracies.

[0015] The information processing apparatus according to the third aspect of the present invention is characterized in that the information processing apparatus includes: a storage unit that stores a feature vector set, a quality label set, and a plurality of non-quality label sets, the feature vector set including a plurality of feature vectors generated by respectively extracting predetermined features from a plurality of digital data representing measurement values measured from an object, the quality label set including a plurality of quality labels respectively corresponding to the respective digital data in the plurality of digital data and indicating the quality of the object, the plurality of non-quality label sets respectively including a plurality of non-quality labels of a type expected to be irrelevant to the quality of the object, the plurality of non-quality labels respectively corresponding to the respective digital data in the plurality of digital data; a non-quality label clustering unit that calculates a variance of the clustering accuracy for each non-quality label set in the plurality of non-quality label sets, thereby calculating a plurality of the variances respectively corresponding to the respective non-quality label sets in the plurality of non-quality label sets, the clustering accuracy being the clustering accuracy when clustering subsets obtained by dividing the plurality of feature vectors by respective elements represented by the plurality of non-quality labels using the quality label set; and a processing unit that generates a screen image capable of determining at least one type of non-quality label that adversely affects the quality of the plurality of digital data using the plurality of variances.

[0016] The program according to the first aspect of the present invention is characterized in that the program causes a computer to function as the following parts: a storage unit that stores a set of feature vectors, a set of quality labels, and a plurality of sets of non-quality labels, the set of feature vectors including a plurality of feature vectors generated by respectively extracting predetermined features from a plurality of digital data representing measurement values measured from an object, the set of quality labels including a plurality of quality labels respectively corresponding to the respective digital data in the plurality of digital data and indicating the quality of the object, the plurality of sets of non-quality labels respectively including a plurality of non-quality labels of a type expected to be irrelevant to the quality of the object, the plurality of non-quality labels respectively corresponding to the respective digital data in the plurality of digital data; a non-quality label clustering unit that calculates an average clustering accuracy for each of the plurality of sets of non-quality labels, thereby calculating a plurality of the average clustering accuracies respectively corresponding to the plurality of sets of non-quality labels, the average clustering accuracy being the average of the clustering accuracies when clustering subsets obtained by dividing the plurality of feature vectors by each of the plurality of elements represented by the plurality of non-quality labels using the set of quality labels; and a processing unit that generates a screen image capable of determining at least one type of non-quality label that has an adverse effect on the quality of the plurality of digital data using the plurality of average clustering accuracies.

[0017] The program according to the second aspect of the present invention is characterized in that the program causes a computer to function as the following parts: a storage unit that stores a set of feature vectors, a set of quality labels, and a plurality of sets of non-quality labels, the set of feature vectors including a plurality of feature vectors generated by respectively extracting predetermined features from a plurality of digital data representing measurement values measured from an object, the set of quality labels including a plurality of quality labels respectively corresponding to the respective digital data in the plurality of digital data and indicating the quality of the object, the plurality of sets of non-quality labels respectively including a plurality of non-quality labels of a type expected to be irrelevant to the quality of the object, the plurality of non-quality labels respectively corresponding to the respective digital data in the plurality of digital data; a non-quality label clustering unit that calculates a clustering accuracy for a set of non-quality labels corresponding to one type of non-quality label selected from the plurality of non-quality labels, thereby calculating a plurality of the clustering accuracies, the clustering accuracy being the clustering accuracy when clustering subsets obtained by dividing the plurality of feature vectors by each of the plurality of elements represented by the plurality of non-quality labels using the set of quality labels; and a processing unit that generates a screen image capable of determining at least one element that has an adverse effect on the quality of the plurality of digital data using the plurality of clustering accuracies.

[0018] The program according to the third aspect of the present invention is characterized in that the program causes a computer to function as the following parts: a storage unit that stores a set of feature vectors, a set of quality labels, and a plurality of sets of non-quality labels, the set of feature vectors including a plurality of feature vectors generated by respectively extracting predetermined features from a plurality of digital data representing measurement values measured from an object, the set of quality labels including a plurality of quality labels respectively corresponding to each of the plurality of digital data and indicating the quality of the object, the plurality of sets of non-quality labels each including a plurality of non-quality labels of a type expected to be irrelevant to the quality of the object, and the plurality of non-quality labels respectively corresponding to each of the plurality of digital data; a non-quality label clustering unit that calculates the variance of the clustering accuracy for each of the plurality of sets of non-quality labels, thereby calculating a plurality of the variances respectively corresponding to each of the plurality of sets of non-quality labels, the clustering accuracy being the clustering accuracy when clustering subsets obtained by dividing the plurality of feature vectors by each of the plurality of elements represented by the plurality of non-quality labels using the set of quality labels; and a processing unit that generates a screen image capable of determining at least one type of non-quality label that adversely affects the quality of the plurality of digital data using the plurality of variances.

[0019] The information processing method according to the first aspect of the present invention is characterized in that a set of feature vectors, a set of quality labels, and a plurality of sets of non-quality labels are stored, the set of feature vectors including a plurality of feature vectors generated by respectively extracting predetermined features from a plurality of digital data representing measurement values measured from an object, the set of quality labels including a plurality of quality labels respectively corresponding to each of the plurality of digital data and indicating the quality of the object, the plurality of sets of non-quality labels each including a plurality of non-quality labels of a type expected to be irrelevant to the quality of the object, and the plurality of non-quality labels respectively corresponding to each of the plurality of digital data, the average clustering accuracy is calculated for each of the plurality of sets of non-quality labels, thereby calculating a plurality of the average clustering accuracies respectively corresponding to each of the plurality of sets of non-quality labels, the average clustering accuracy being the average value of the clustering accuracies when clustering subsets obtained by dividing the plurality of feature vectors by each of the plurality of elements represented by the plurality of non-quality labels using the set of quality labels, and a screen image capable of determining at least one type of non-quality label that adversely affects the quality of the plurality of digital data is generated.

[0020] The information processing method according to the second aspect of the present invention is characterized in that a set of feature vectors, a set of quality labels, and a plurality of sets of non-quality labels are stored. The set of feature vectors includes a plurality of feature vectors generated by respectively extracting predetermined features from a plurality of digital data representing measurement values measured from an object. The set of quality labels includes a plurality of quality labels respectively corresponding to each of the plurality of digital data and indicating the quality of the object. Each of the plurality of sets of non-quality labels includes a plurality of non-quality labels of a type expected to be irrelevant to the quality of the object. The plurality of non-quality labels respectively correspond to each of the plurality of digital data. A clustering accuracy is calculated for the set of non-quality labels corresponding to one type of non-quality label selected from the plurality of non-quality labels, and thus a plurality of the clustering accuracies are calculated. The clustering accuracy is the clustering accuracy when clustering subsets obtained by dividing the plurality of feature vectors by each of the plurality of elements represented by the plurality of non-quality labels using the set of quality labels. A screen image is generated that can use the plurality of clustering accuracies to determine at least one element that has an adverse effect on the quality of the plurality of digital data.

[0021] The information processing method according to the third aspect of the present invention is characterized in that a set of feature vectors, a set of quality labels, and a plurality of sets of non-quality labels are stored. The set of feature vectors includes a plurality of feature vectors generated by respectively extracting predetermined features from a plurality of digital data representing measurement values measured from an object. The set of quality labels includes a plurality of quality labels respectively corresponding to each of the plurality of digital data and indicating the quality of the object. Each of the plurality of sets of non-quality labels includes a plurality of non-quality labels of a type expected to be irrelevant to the quality of the object. The plurality of non-quality labels respectively correspond to each of the plurality of digital data. A variance of the clustering accuracy is calculated for each of the plurality of sets of non-quality labels, and thus a plurality of the variances respectively corresponding to each of the plurality of sets of non-quality labels are calculated. The clustering accuracy is the clustering accuracy when clustering subsets obtained by dividing the plurality of feature vectors by each of the plurality of elements represented by the plurality of non-quality labels using the set of quality labels. A screen image is generated that can use the plurality of variances to determine at least one type of non-quality label that has an adverse effect on the quality of the plurality of digital data.

[0022] Advantages of the Invention

[0023] According to one or more aspects of the present invention, it is possible to determine the cause of the deterioration of the quality of the data set being used. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1is a block diagram schematically showing the configuration of the information processing apparatus according to Embodiment 1.

[0025] Figure 2 is a block diagram schematically showing an example of use of the information processing apparatus according to Embodiment 1.

[0026] Figure 3 (A) to (C) are graphs for explaining the clustering accuracy of each subset and the overall clustering in the non-quality labels of the inspector.

[0027] Figure 4 is a graph for explaining the clustering accuracy for the entire data in the case where the non-uniformity caused by the difference between inspectors is eliminated by a certain method.

[0028] Figure 5 (A) and (B) are block diagrams showing an example of the hardware configuration.

[0029] Figure 6 is a flowchart showing the process of the information processing apparatus for displaying an image of a label type evaluation screen.

[0030] Figure 7 is a flowchart showing the process of the information processing apparatus for displaying an image of an accuracy improvement amount screen.

[0031] Figure 8 is a flowchart showing the process of the information processing apparatus for displaying an image of an accuracy influencing factor evaluation screen. Detailed Embodiment

[0032] Hereinafter, as an embodiment, a case where the soundness of a motor is determined based on the vibration of the motor as an object will be described as an example.

[0033] Figure 1 is a block diagram schematically showing the configuration of the information processing apparatus 100 according to Embodiment 1.

[0034] Figure 2 is a block diagram schematically showing an example of use of the information processing apparatus 100 according to Embodiment 1.

[0035] As Figure 2 shown, for example, the information processing apparatus 100 is connected to bases located in different places such as the first factory 200A and the second factory 200B via a network 201 such as the Internet.

[0036] Factories such as the first factory 200A and the second factory 200B manufacture the motor as an object using the same machinery and equipment, and the connection content with the information processing apparatus 100 is also the same. Therefore, the first factory 200A will be described below.

[0037] In Factory 1 200A, there are multiple manufacturing lines 203A, 203B, 203C, … for manufacturing the motor 202.

[0038] Inspectors assigned to each of the manufacturing lines 203A, 203B, 203C, … use inspection devices 204A, 204B, 204C, … configured for each of the manufacturing lines 203A, 203B, 203C, … to inspect the motors 202 manufactured by each of the manufacturing lines 203A, 203B, 203C, ….

[0039] For example, each of the inspection devices 204A, 204B, 204C, … measures the amplitude of vibration when driving the motor 202 and generates digital data DD. This digital data DD includes a motor number as motor identification information for identifying the inspected motor 202 and inspection data representing the amplitude as its measured value.

[0040] In addition, each of the inspection devices 204A, 204B, 204C, … generates non-quality label data ND. This non-quality label data ND shows the motor number of the inspected motor 202, the data number of the digital data DD obtained in this inspection, and non-quality labels of a type that is expected to be independent of the quality of the motor 202. Also, in this embodiment, each of the inspection devices 204A, 204B, 204C, … generates non-quality label data ND that includes multiple types of non-quality labels.

[0041] Here, as types of non-quality labels, there are inspectors, date and time, manufacturing line, location, and inspection device.

[0042] In addition, the non-quality label of the inspector includes the inspector number as inspector identification information for identifying the inspector as its element.

[0043] The non-quality label of the date and time includes the measurement date and time, which is the date and time when the inspection was performed, as its element.

[0044] The non-quality label of the manufacturing line includes the line number as line identification information for identifying the manufacturing line as its element.

[0045] The non-quality label of the location includes the location ID as factory identification information for identifying the factory as its element.

[0046] The non-quality label of the inspection device includes the device number as the inspection device identification number for identifying the inspection device as its element.

[0047] Specifically, the first non-quality label data ND#1 to the fifth non-quality label data ND#5 are generated. The first non-quality label data ND#1 shows the motor number of the inspected motor 202, the data number of the digital data DD obtained in this inspection, and the inspector number of the inspector who conducted this inspection. The second non-quality label data ND#2 shows the motor number of the inspected motor 202, the data number of the digital data DD obtained in this inspection, and the measurement date and time of this inspection. The third non-quality label data ND#3 shows the motor number of the inspected motor 202, the data number of the digital data DD obtained in this inspection, and the line number of the production line that manufactured this motor 202. The fourth non-quality label data ND#4 shows the motor number of the inspected motor 202, the data number of the digital data DD obtained in this inspection, and the site ID of the factory where this motor 202 was manufactured. The fifth non-quality label data ND5 shows the motor number of the inspected motor 202, the data number of the digital data DD obtained in this inspection, and the device number of the inspection device that inspected this motor 202.

[0048] In addition, each non-quality label data ND contains information indicating the type of the corresponding non-quality label.

[0049] Moreover, each inspection device 204A, 204B, 204C,... sends the digital data DD and the non-quality label data ND generated as described above to the information processing device 100 via the network 201.

[0050] In addition, the non-quality label is a type of label that is expected to be independent of the quality. In other words, the non-quality label is a type of label that those conducting quality management do not want to represent the quality. Here, according to the inspector, date and time, production line, site, and inspection device, the quality of the motor 202 is not desired to be represented as good or bad. Therefore, these types are used for labeling.

[0051] In addition, a quality label assigning device 205 is provided in the first factory 200A.

[0052] For example, the motor 202 manufactured by the first factory 200A is subjected to a final inspection by an experienced inspector or the like, and the normal or abnormal result of this inspection and the motor number of the inspected motor 202 are input into the quality label assigning device 205.

[0053] The quality label assigning device 205 generates quality label data CD indicating the input motor number and normal or abnormal, and sends the generated quality label data CD to the information processing device 100 via the network 201. Here, the quality label is a label indicating the quality (here, normal or abnormal).

[0054] Receiving the digital data DD, quality label data CD, and non-quality label data ND sent as described above, the information processing apparatus 100 performs processing.

[0055] As Figure 1 shown, the information processing apparatus 100 includes a communication unit 101, a storage unit 102, a feature extraction unit 103, an input unit 104, a selection unit 105, a quality label clustering unit 106, a non-quality label clustering unit 107, a processing unit 108, and a display unit 109.

[0056] The communication unit 101 communicates with the network 201. For example, the communication unit 101 receives a plurality of digital data DD, a plurality of quality label data CD, and a plurality of non-quality label data ND from a plurality of factories via the network 201.

[0057] The storage unit 102 stores data and programs required for processing in the information processing apparatus 100. For example, the storage unit 102 stores the plurality of digital data DD, the plurality of quality label data CD, and the plurality of non-quality label data ND received by the communication unit 101 as a digital data set DG, a quality label set CG, and a non-quality label set NG, respectively.

[0058] In addition, as described later, the storage unit 102 stores a feature vector set BG generated by the feature extraction unit 103.

[0059] Further, in the present embodiment, as the non-quality label data ND, for example, the first non-quality label data ND#1 to the fifth non-quality label data ND#5 are stored corresponding to the types of non-quality labels.

[0060] The feature extraction unit 103 reads out the digital data set DG stored in the storage unit 102, extracts predetermined features from the inspection data included in the digital data DD in the read digital data set DG, and generates feature vector data BD representing the extracted features and the motor numbers included in the digital data DD. Then, the feature extraction unit 103 stores the plurality of feature vector data BD as a feature vector set BG in the storage unit 102. As a method for feature extraction from inspection data, for example, there are filter bank analysis, wavelet analysis, LPC (Linear Predictive Coding) analysis, cepstrum analysis, etc. In addition, the features extracted here are represented by feature vectors.

[0061] The input unit 104 accepts the input of an instruction from an operator of the information processing apparatus 100.

[0062] For example, the input unit 104 accepts the input of selecting a processing mode. In the present embodiment, the processing modes are a label type evaluation mode, an accuracy improvement amount calculation mode, and an accuracy influencing factor evaluation mode.

[0063] In addition, when the input unit 104 selects the accuracy influence factor evaluation mode, it also accepts the input of the types of non-quality labels for evaluating the factors that affect the accuracy.

[0064] Then, the input unit 104 notifies the selected processing mode and the types of non-quality labels selected when the accuracy influence factor evaluation mode is selected to the selection unit 105 and the processing unit 108.

[0065] The selection unit 105 selects and reads out the data stored in the storage unit 102 according to the selection input to the input unit 104.

[0066] For example, when the selection unit 105 selects the label type evaluation mode, it reads out the feature vector set BG, the quality label set CG, and the non-quality label set NG of all types from the storage unit 102, and provides the read data to the non-quality label clustering unit 107.

[0067] In addition, when the selection unit 105 selects the accuracy improvement amount calculation mode, it reads out the feature vector set BG and the quality label set CG from the storage unit 102, provides the read data to the quality label clustering unit 106, and reads out the feature vector set BG, the quality label set CG, and the non-quality label set NG of all types from the storage unit 102, and provides the read data to the non-quality label clustering unit 107.

[0068] Furthermore, when the selection unit 105 selects the accuracy influence factor evaluation mode, it reads out the feature vector set BG, the quality label set CG, and the non-quality label set NG corresponding to the types of non-quality labels selected by the input unit 104 from the storage unit 102, and provides the read data to the non-quality label clustering unit 107.

[0069] The quality label clustering unit 106 performs clustering based on the feature vector set BG provided by the selection unit 105, compares the quality determination result based on this clustering (such as normal or abnormal) with the inspection result (such as normal or abnormal) represented by the quality label set CG, and calculates the clustering accuracy. The clustering accuracy calculated here is also called the reference clustering accuracy.

[0070] Let the clustering accuracy be the ratio of clustering success or the ratio of clustering failure.

[0071] In the present embodiment, the clustering accuracy is set as the correct solution rate of the quality determination result based on clustering with respect to the inspection result represented by the quality label set CG. However, the present embodiment is not limited to such an example.

[0072] For example, the clustering accuracy can also be the error rate, F value, true positive rate (TPR), or true negative rate (TNR) of the determination result based on the clustering quality relative to the inspection result represented by the quality label set CG.

[0073] When the non-quality label clustering unit 107 receives the non-quality label set NG of all types of non-quality labels from the selection unit 105, it divides the feature vector data BD included in the feature vector set BG provided by the selection unit 105 into subsets of each element of the non-quality labels in each type of the non-quality label set NG. For example, when the type of the non-quality label set NG is the inspector number, the feature vector data BD included in the feature vector set BG is divided according to each inspector number.

[0074] Next, the non-quality label clustering unit 107 performs clustering based on the divided feature vector data BD, compares the determination result based on the clustering quality with the inspection result represented by the quality label set CG, and calculates the clustering accuracy of each subset (in other words, each element). Then, the non-quality label clustering unit 107 calculates the average value of the clustering accuracies of the calculated each subset as the average clustering accuracy according to each type of non-quality label.

[0075] In other words, in the label type evaluation mode and the accuracy improvement amount calculation mode, the non-quality label clustering unit 107 calculates the average clustering accuracy of each type of non-quality label, and provides the calculated average clustering accuracy to the processing unit 108.

[0076] On the other hand, when the non-quality label clustering unit 107 receives the non-quality label set NG of one type of non-quality label from the selection unit 105, it divides the feature vector data BD included in the feature vector set BG provided by the selection unit 105 into subsets of each element of the non-quality label represented by the non-quality label set NG.

[0077] Next, the non-quality label clustering unit 107 performs clustering based on the divided feature vector data BD, compares the determination result based on the clustering quality with the inspection result represented by the quality label set CG, and calculates the clustering accuracy of each subset (in other words, each element).

[0078] In other words, in the accuracy influencing factor evaluation mode, the non-quality label clustering unit 107 calculates the clustering accuracy of each subset in the selected type of non-quality label, and provides the calculated clustering accuracy of each subset to the processing unit 108.

[0079] The processing unit 108 processes according to the processing mode input by the input unit 104, using at least any one of the clustering accuracy calculated by the quality label clustering unit 106 and the average clustering accuracy calculated by the non-quality label clustering unit 107.

[0080] Here, the processing unit 108 generates a screen image that can use multiple average clustering accuracies to determine at least one type of non-quality label that has an adverse effect on the quality of multiple digital data DD, or generates a screen image that can use multiple clustering accuracies to determine at least one factor that has an adverse effect on the quality of multiple digital data DD.

[0081] For example, in the label type evaluation mode, the processing unit 108 generates a label type evaluation screen image that displays at least a part of the types of multiple non-quality labels together with the average clustering accuracy in descending order of the average clustering accuracy.

[0082] In the accuracy improvement amount calculation mode, the processing unit 108 subtracts the clustering accuracy calculated by the quality label clustering unit 106 from each of the multiple average clustering accuracies calculated by the non-quality label clustering unit 107, thereby calculating the improvement amount of the clustering accuracy for each type of non-quality label. Then, the processing unit 108 generates an accuracy improvement amount screen image representing at least a part of the types of multiple non-quality labels and the corresponding calculated improvement amounts.

[0083] In the accuracy influencing factor evaluation mode, the processing unit 108 generates an accuracy influencing factor evaluation screen image that shows at least a part of the corresponding factors together with the clustering accuracy in ascending order of the clustering accuracy of each subset in one type of non-quality label calculated by the non-quality label clustering unit 107.

[0084] The display unit 109 displays various screen images. For example, the display unit 109 displays the label type evaluation screen image, the accuracy improvement amount screen image, or the accuracy influencing factor evaluation screen image generated by the processing unit 108.

[0085] Next, the basic idea of the processing in the information processing apparatus 100 will be described.

[0086] When dividing the feature vector by non-quality labels regardless of quality, when clustering according to each divided subset, it is expected that the average clustering accuracy is higher than when clustering the entire data set in the same way.

[0087] Figure 3 (A) to (C) are graphs for explaining the clustering accuracy of each subset and the overall clustering in the non-quality labels of the inspector.

[0088] For example, Figure 3 (A) is a curve graph formed by depicting a histogram of the normality or abnormality of the motor 202 based on the inspection data measured by inspector A.

[0089] Similarly, Figure 3 (B) is a curve graph formed by depicting a histogram of the normality or abnormality of the motor 202 based on the inspection data measured by inspector B.

[0090] Figure 3 (C) overlays and displays Figure 3 the histogram shown in (A) of Figure 3 and the histogram shown in (B) of

[0091] As Figure 3 shown in (C) of

[0092] However, as Figure 3 shown in (A) of Figure 3 when only considering the data of inspector A, by setting the boundary 300 that determines normality and abnormality, clustering of normality and abnormality can be performed. Similarly, as

[0093] shown in (B) of Figure 4 clustering of normality and abnormality can also be performed for the data of inspector B by setting the boundary 301 that determines normality and abnormality.

[0094] At this time, as shown in Figure 4 it is possible to expect that the average clustering accuracy for the individual subsets of inspectors is consistent with the clustering accuracy for the entire data set in the case where the non-uniformity caused by differences between inspectors is eliminated by a certain method. Therefore, the average clustering accuracy for the individual subsets of inspectors can be used as an expected value of the accuracy obtained in the case where the non-uniformity caused by differences between measurers is eliminated.

[0094] As described above, in the label type evaluation screen image, by arranging the types of non-quality labels in descending order of average clustering accuracy, the deviation of the acquisition method when acquiring inspection data is improved. As a result, it is possible to identify the factors that can improve the clustering accuracy. In other words, it is possible to identify the reasons for the deterioration of the clustering accuracy of the entire data set. That is, it can be understood that the types of non-quality labels with higher average clustering accuracy have a greater impact on the quality of the inspection data and are more likely to be the cause of adverse effects on the quality of the inspection data.

[0095] In addition, in the accuracy improvement amount screen image, by displaying the improvement amount of the clustering accuracy together with the types of non-quality labels, in these types of non-quality labels, a certain method is used to improve the acquisition method when acquiring inspection data. Thus, it is possible to grasp to what extent the overall clustering accuracy can be improved. Along with this, the larger the improvement amount of the clustering accuracy, the more it can be estimated that it is the cause of deteriorating the clustering accuracy of the entire data. That is, it is possible to grasp that the types of non-quality labels with a larger improvement amount of the clustering accuracy have a greater impact on the quality of the inspection data, and the possibility of being the cause of having an adverse effect on the quality of the inspection data is higher.

[0096] Furthermore, in the accuracy influence factor evaluation screen image, by showing the corresponding factors together with this clustering accuracy, when acquiring inspection data, it is possible to grasp which factor's acquisition method needs to be improved. Along with this, it is possible to determine the factor that deteriorates the clustering accuracy of the entire data. That is, it is possible to grasp that the factor with a lower clustering accuracy has a greater impact on the quality of the inspection data, and the possibility of being the cause of having an adverse effect on the quality of the inspection data is higher.

[0097] For example, as Figure 5 shown in (A) above, a part or all of the above-described feature extraction unit 103, selection unit 105, quality label clustering unit 106, non-quality label clustering unit 107, and processing unit 108 can be constituted by a processor 11 such as a memory 10 and a CPU (Central Processing Unit) that executes a program stored in the memory 10. Such a program can be provided via a network, and in addition, it can also be provided by being recorded on a recording medium. That is, such a program can be provided as a program product, for example.

[0098] In addition, for example, as Figure 5 shown in (B) above, a part or all of the feature extraction unit 103, selection unit 105, quality label clustering unit 106, non-quality label clustering unit 107, and processing unit 108 can also be constituted by a processing circuit 12 such as a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array).

[0099] In addition, the communication unit 101 can be implemented by a communication device such as a NIC (Network Interface Card).

[0100] In addition, the storage unit 102 can be implemented by a storage device such as an HDD (Hard Disk Drive).

[0101] The input unit 104 can be implemented by an input device such as a mouse or a keyboard.

[0102] The display unit 109 can be implemented by a display device such as a liquid crystal display.

[0103] As described above, the information processing apparatus 100 can be implemented by a so-called computer.

[0104] Figure 6 It is a flowchart showing the process of the information processing apparatus 100 displaying an image of a label type evaluation screen.

[0105] For example, an operator of the information processing apparatus 100 inputs an instruction to select the label type evaluation mode at the input unit 104, whereby Figure 6 the flowchart shown starts. In this case, the input unit 104 notifies the selection unit 105 and the processing unit 108 that the label type evaluation mode has been selected.

[0106] First, the selection unit 105 reads out the feature vector set BG, the quality label set CG, and the non-quality label set NG corresponding to all types of non-quality labels stored in the storage unit 102, and provides the read data to the non-quality label clustering unit 107 (S10).

[0107] Next, the non-quality label clustering unit 107 selects the non-quality label set NG corresponding to one type of non-quality label for which clustering has not yet been performed from the non-quality label set NG received from the selection unit 105 (S11).

[0108] Next, the non-quality label clustering unit 107 divides the feature vector set BG provided by the selection unit 105 into subsets for each element of the non-quality label represented by the selected non-quality label set NG, and performs clustering for each divided subset (S12).

[0109] Next, the non-quality label clustering unit 107 compares the quality determination result based on the clustering performed in step S12 with the inspection result represented by the quality label set CG, calculates the clustering accuracy for each subset, and calculates the average value thereof, i.e., the average clustering accuracy (S13). The calculated average clustering accuracy is notified to the processing unit 108 together with the type of the non-quality label.

[0110] Next, the non-quality label clustering unit 107 determines whether clustering has been performed in the non-quality label set NG corresponding to all types of non-quality labels (S14). If clustering has been performed in the non-quality label set NG of all types (S14: Yes), the process proceeds to step S15. If there remains a non-quality label set NG of a type for which clustering has not yet been performed (S14: No), the process returns to step S11.

[0111] In step S15, the processing unit 108 generates a label type evaluation screen image, and this label type evaluation screen image displays at least a part of the types of non-quality labels together with the average clustering accuracy in descending order of the average clustering accuracy calculated by the non-quality label clustering unit 107 (S15).

[0112] Next, the display unit 109 displays the label type evaluation screen image generated by the processing unit 108 (S16).

[0113] Figure 7 It is a flowchart showing the process of the information processing apparatus 100 displaying an accuracy improvement amount screen image.

[0114] For example, an operator of the information processing apparatus 100 inputs an instruction to select an accuracy improvement amount calculation mode at the input unit 104, whereby Figure 7 the flowchart shown starts. In this case, the input unit 104 notifies the selection unit 105 and the processing unit 108 that the accuracy improvement amount calculation mode has been selected.

[0115] First, the selection unit 105 reads out the feature vector set BG and the quality label set CG from the storage unit 102, and provides the read data to the quality label clustering unit 106 (S20).

[0116] Next, the quality label clustering unit 106 performs clustering based on the feature vector set BG provided from the selection unit 105 (S21).

[0117] Next, the quality label clustering unit 106 compares the quality determination result based on the clustering performed in step S21 with the inspection result represented by the quality label set CG, and calculates the clustering accuracy (S22). The clustering accuracy calculated here is provided to the processing unit 108.

[0118] Next, the selection unit 105 reads out the feature vector set BG, the quality label set CG, and the non-quality label set NG corresponding to all types of non-quality labels stored in the storage unit 102, and provides the read data to the non-quality label clustering unit 107 (S23).

[0119] Next, the non-quality label clustering unit 107 selects the non-quality label set NG corresponding to one type of non-quality label for which clustering has not yet been performed from the non-quality label set NG received from the selection unit 105 (S24).

[0120] Next, the non-quality label clustering unit 107 divides the feature vector set BG provided from the selection unit 105 into subsets for each element of the non-quality label represented by the selected non-quality label set NG, and performs clustering for each of the divided subsets (S25).

[0121] Next, the non-quality label clustering unit 107 compares the quality determination result based on the clustering performed in step S12 with the inspection result represented by the quality label set CG, calculates the clustering accuracy for each subset, and calculates the average value thereof, i.e., the average clustering accuracy (S26). The calculated average clustering accuracy is notified to the processing unit 108 together with the type of the non-quality label.

[0122] Next, the non-quality label clustering unit 107 determines whether clustering has been performed in the non-quality label set NG corresponding to all types of non-quality labels (S27). If clustering has been performed in the non-quality label set NG of all types (S27: Yes), the process proceeds to step S28. If there remains a non-quality label set NG of a type for which clustering has not been performed (S27: No), the process returns to step S24.

[0123] Next, the processing unit 108 subtracts the clustering accuracy calculated by the quality label clustering unit 106 from the average clustering accuracy of all types of non-quality labels calculated by the non-quality label clustering unit 107, thereby calculating the improvement amount of the clustering accuracy for each such type.

[0124] Next, the processing unit 108 generates an improvement amount of accuracy screen image that shows at least one type of the non-quality label type and the corresponding calculated improvement amount of accuracy.

[0125] Next, the display unit 109 displays the improvement amount of accuracy screen image generated by the processing unit 108 (S30).

[0126] In addition, in Figure 7 the processing of steps S20 to S22 and the processing of steps S23 to S27 can also be performed in parallel.

[0127] Figure 8 is a flowchart showing the processing in which the information processing apparatus 100 displays an accuracy influence factor evaluation screen image.

[0128] For example, an operator of the information processing apparatus 100 inputs an instruction to select the accuracy influence factor evaluation mode at the input unit 104, whereby Figure 8 the flowchart shown starts. In this case, the input unit 104 notifies the selection unit 105 and the processing unit 108 that the accuracy influence factor evaluation mode has been selected.

[0129] First, the selection unit 105 reads out the feature vector set BG, the quality label set CG, and the non-quality label set NG corresponding to the type selected by the input unit 104 from the storage unit 102, and provides the read data to the non-quality label clustering unit 107 (S40).

[0130] Next, the non-quality label clustering unit 107 divides the set of feature vectors BG provided by the selection unit 105 into subsets for each element of the non-quality labels represented by the non-quality label set NG, and performs clustering on each of the divided subsets (S41).

[0131] Next, the non-quality label clustering unit 107 compares the quality determination results based on the clustering performed in step S41 with the inspection results represented by the quality label set CG, and calculates the clustering accuracy for each subset (S42). The clustering accuracy for each subset calculated here is provided to the processing unit 108.

[0132] Next, the processing unit 108 generates an accuracy influencing factor evaluation screen image, which shows at least one of the corresponding elements together with the clustering accuracy in ascending order of the clustering accuracy for each subset in one type of non-quality label calculated by the non-quality label clustering unit 107 (S43).

[0133] Next, the display unit 109 displays the accuracy influencing factor evaluation screen image generated by the processing unit 108 (S44).

[0134] According to the above embodiment, it is possible to generate and display a screen image showing at least one type or element of non-quality label that has an adverse effect on the quality of the digital data DD.

[0135] In the above-described embodiment, as a screen image that can use multiple average clustering accuracies to determine at least one type of non-quality label that has an adverse effect on the quality of multiple digital data DD, the processing unit 108 generates a label type evaluation screen image in the label type evaluation mode. This label type evaluation screen image displays at least a part of the types of multiple non-quality labels together with the average clustering accuracy in descending order of the average clustering accuracy. However, the embodiment is not limited to this example.

[0136] For example, the processing unit 108 may also generate a label type evaluation screen image that shows at least one of the multiple types in descending order of multiple variances.

[0137] In this case, the non-quality label clustering unit 107 only needs to calculate the variance of the clustering accuracy for each subset calculated as described above for each type of non-quality label.

[0138] By displaying the variance of the clustering accuracy of each non-quality label, it is possible to determine a non-quality label with a large deviation in the clustering accuracy of each element. Moreover, by correcting the inspection method of the non-quality label with a large deviation, the quality of the digital data DD can be improved.

[0139] Description of Reference Numerals

[0140] 100: Information processing device; 101: Communication unit; 102: Storage unit; 103: Feature extraction unit; 104: Input unit; 105: Selection unit; 106: Quality label clustering unit; 107: Non-quality label clustering unit; 108: Processing unit; 109: Display unit.

Claims

1. An information processing apparatus, characterized in that, the information processing apparatus includes: a storage unit that stores a set of feature vectors, a set of quality labels, and a plurality of sets of non-quality labels. The set of feature vectors includes a plurality of feature vectors generated by respectively extracting predetermined features from a plurality of digital data representing measurement values measured from an object. The set of quality labels includes a plurality of quality labels corresponding to each of the plurality of digital data and indicating the quality of the object. The plurality of sets of non-quality labels each include a plurality of non-quality labels of a type expected to be irrelevant to the quality of the object, and the plurality of non-quality labels respectively correspond to each of the plurality of digital data; a non-quality label clustering unit that calculates an average clustering accuracy for each of the plurality of sets of non-quality labels, thereby calculating a plurality of the average clustering accuracies corresponding to each of the plurality of sets of non-quality labels. The average clustering accuracy is the average value of the clustering accuracies when clustering subsets obtained by dividing the plurality of feature vectors by each of the plurality of elements represented by the plurality of non-quality labels using the set of quality labels; and a processing unit that generates a screen image capable of determining at least one type of non-quality label that adversely affects the quality of the plurality of digital data using the plurality of average clustering accuracies.

2. The information processing apparatus according to claim 1, characterized in that, the processing unit generates a label type evaluation screen image as the screen image, and the label type evaluation screen image shows at least one of the plurality of types in descending order of the plurality of average clustering accuracies.

3. The information processing apparatus according to claim 1, characterized in that, the information processing apparatus further includes a quality label clustering unit that calculates a reference clustering accuracy when clustering the plurality of feature vectors using the set of quality labels, the processing unit subtracts the reference clustering accuracy from each of the plurality of average clustering accuracies, thereby calculating a plurality of improvement amounts, and generates an accuracy improvement amount screen image as the screen image. The accuracy improvement amount screen image shows at least one of the plurality of types in descending order of the plurality of improvement amounts, together with the corresponding improvement amounts.

4. The information processing apparatus according to any one of claims 1 to 3, characterized in that, the clustering accuracy is the ratio of successful clustering or the ratio of failed clustering.

5. The information processing apparatus according to any one of claims 1 to 3, characterized in that, the information processing apparatus further includes a display unit that displays the screen image.

6. The information processing apparatus according to claim 4, characterized in that, the information processing apparatus further includes a display unit that displays the screen image.

7. An information processing apparatus, characterized in that, the information processing apparatus includes: A storage unit that stores a set of feature vectors, a set of quality labels, and a plurality of sets of non-quality labels. The set of feature vectors includes a plurality of feature vectors generated by respectively extracting predetermined features from a plurality of digital data representing measurement values measured from an object. The set of quality labels includes a plurality of quality labels respectively corresponding to each of the plurality of digital data and indicating the quality of the object. The plurality of sets of non-quality labels each include a plurality of non-quality labels of a type expected to be unrelated to the quality of the object, and the plurality of non-quality labels respectively correspond to each of the plurality of digital data; A non-quality label clustering unit that calculates a clustering accuracy for a set of non-quality labels corresponding to a type of non-quality label selected from the plurality of non-quality labels, thereby calculating a plurality of the clustering accuracies. The clustering accuracy is the clustering accuracy when clustering subsets obtained by dividing the plurality of feature vectors by each of the plurality of elements represented by the plurality of non-quality labels using the set of quality labels; and A processing unit that generates a screen image capable of determining at least one element that has an adverse effect on the quality of the plurality of digital data using the plurality of clustering accuracies.

8. The information processing apparatus according to claim 7, wherein the processing unit generates a precision-affecting element evaluation screen image as the screen image, and the precision-affecting element evaluation screen image shows at least one of the plurality of elements in ascending order of the plurality of clustering accuracies.

9. The information processing apparatus according to claim 7, wherein the clustering accuracy is the ratio of successful clustering or the ratio of failed clustering.

10. The information processing apparatus according to claim 8, wherein the clustering accuracy is the ratio of successful clustering or the ratio of failed clustering.

11. The information processing apparatus according to any one of claims 7 to 10, wherein the information processing apparatus further includes a display unit that displays the screen image.

12. An information processing apparatus, wherein the information processing apparatus includes: A storage unit that stores a set of feature vectors, a set of quality labels, and a plurality of sets of non-quality labels. The set of feature vectors includes a plurality of feature vectors generated by respectively extracting predetermined features from a plurality of digital data representing measurement values measured from an object. The set of quality labels includes a plurality of quality labels respectively corresponding to each of the plurality of digital data and indicating the quality of the object. The plurality of sets of non-quality labels each include a plurality of non-quality labels of a type expected to be unrelated to the quality of the object, and the plurality of non-quality labels respectively correspond to each of the plurality of digital data; A non-quality label clustering unit that calculates the variance of the clustering accuracy for each of the multiple non-quality label sets, thereby calculating multiple variances corresponding to each of the multiple non-quality label sets, where the clustering accuracy is the clustering accuracy when clustering subsets formed by dividing the multiple feature vectors by each of the multiple elements represented by the multiple non-quality labels using the quality label set; and A processing unit that generates a screen image capable of determining at least one type of non-quality label that adversely affects the quality of the multiple digital data using the multiple variances.

13. The information processing apparatus according to claim 12, wherein the processing unit generates a label type evaluation screen image as the screen image, and the label type evaluation screen image shows at least one of the multiple types in descending order of the multiple variances.

14. The information processing apparatus according to claim 12, wherein the clustering accuracy is the proportion of successful clustering or the proportion of failed clustering.

15. The information processing apparatus according to claim 13, wherein the clustering accuracy is the proportion of successful clustering or the proportion of failed clustering.

16. The information processing apparatus according to any one of claims 12 to 15, wherein the information processing apparatus further includes a display unit that displays the screen image.

17. A computer-readable recording medium storing a computer program, when the computer program is executed by a processor, stores a feature vector set, a quality label set, and multiple non-quality label sets, where the feature vector set includes multiple feature vectors generated by respectively extracting predetermined features from multiple digital data representing measurement values measured from an object, the quality label set includes multiple quality labels respectively corresponding to each of the multiple digital data and indicating the quality of the object, the multiple non-quality label sets respectively include multiple non-quality labels of a type expected to be independent of the quality of the object, and the multiple non-quality labels respectively correspond to each of the multiple digital data, calculates the average clustering accuracy for each of the multiple non-quality label sets, thereby calculating multiple average clustering accuracies corresponding to each of the multiple non-quality label sets, where the average clustering accuracy is the average of the clustering accuracies when clustering subsets formed by dividing the multiple feature vectors by each of the multiple elements represented by the multiple non-quality labels using the quality label set, generates a screen image capable of determining at least one type of non-quality label that adversely affects the quality of the multiple digital data using the multiple average clustering accuracies.

18. A computer-readable recording medium storing a computer program, when the computer program is executed by a processor, Store a set of feature vectors, a set of quality labels, and multiple sets of non-quality labels. The set of feature vectors includes multiple feature vectors generated by respectively extracting predetermined features from multiple digital data representing measurement values measured from an object. The set of quality labels includes multiple quality labels respectively corresponding to each of the multiple digital data and indicating the quality of the object. The multiple sets of non-quality labels respectively include multiple non-quality labels of a type expected to be irrelevant to the quality of the object, and the multiple non-quality labels respectively correspond to each of the multiple digital data. Calculate the clustering accuracy for the set of non-quality labels corresponding to one type of non-quality label selected from the multiple non-quality labels, thereby calculating multiple such clustering accuracies. The clustering accuracy is the clustering accuracy when using the set of quality labels to cluster subsets obtained by dividing the multiple feature vectors by each of the multiple elements represented by the multiple non-quality labels. Generate a screen image capable of using the multiple clustering accuracies to determine at least one element that has an adverse effect on the quality of the multiple digital data.

19. A computer-readable recording medium storing a computer program, when the computer program is executed by a processor, Store a set of feature vectors, a set of quality labels, and multiple sets of non-quality labels. The set of feature vectors includes multiple feature vectors generated by respectively extracting predetermined features from multiple digital data representing measurement values measured from an object. The set of quality labels includes multiple quality labels respectively corresponding to each of the multiple digital data and indicating the quality of the object. The multiple sets of non-quality labels respectively include multiple non-quality labels of a type expected to be irrelevant to the quality of the object, and the multiple non-quality labels respectively correspond to each of the multiple digital data. Calculate the variance of the clustering accuracy for each set of non-quality labels in the multiple sets of non-quality labels, thereby calculating multiple such variances respectively corresponding to each set of non-quality labels in the multiple sets of non-quality labels. The clustering accuracy is the clustering accuracy when using the set of quality labels to cluster subsets obtained by dividing the multiple feature vectors by each of the multiple elements represented by the multiple non-quality labels. Generate a screen image capable of using the multiple variances to determine at least one type of non-quality label that has an adverse effect on the quality of the multiple digital data.

20. An information processing method, Characterized in that, Store a set of feature vectors, a set of quality labels, and multiple sets of non-quality labels. The set of feature vectors contains multiple feature vectors generated by respectively extracting predetermined features from multiple digital data representing measurement values measured from an object. The set of quality labels contains multiple quality labels respectively corresponding to each of the multiple digital data and indicating the quality of the object. The multiple sets of non-quality labels each contain multiple non-quality labels of a type expected to be irrelevant to the quality of the object, and the multiple non-quality labels respectively correspond to each of the multiple digital data. Calculate the average clustering accuracy for each of the multiple sets of non-quality labels, thereby calculating multiple average clustering accuracies respectively corresponding to each of the multiple sets of non-quality labels. The average clustering accuracy is the average of the clustering accuracies when clustering subsets obtained by dividing the multiple feature vectors by each of the multiple elements represented by the multiple non-quality labels using the set of quality labels. Generate a screen image capable of using the multiple average clustering accuracies to determine at least one type of non-quality label that has an adverse effect on the quality of the multiple digital data.

21. An information processing method characterized in that Store a set of feature vectors, a set of quality labels, and multiple sets of non-quality labels. The set of feature vectors contains multiple feature vectors generated by respectively extracting predetermined features from multiple digital data representing measurement values measured from an object. The set of quality labels contains multiple quality labels respectively corresponding to each of the multiple digital data and indicating the quality of the object. The multiple sets of non-quality labels each contain multiple non-quality labels of a type expected to be irrelevant to the quality of the object, and the multiple non-quality labels respectively correspond to each of the multiple digital data. Calculate the clustering accuracy for the set of non-quality labels corresponding to one type of non-quality label selected from the multiple non-quality labels, thereby calculating multiple clustering accuracies. The clustering accuracy is the clustering accuracy when clustering subsets obtained by dividing the multiple feature vectors by each of the multiple elements represented by the multiple non-quality labels using the set of quality labels. Generate a screen image capable of using the multiple clustering accuracies to determine at least one element that has an adverse effect on the quality of the multiple digital data.

22. An information processing method characterized in that Store a set of feature vectors, a set of quality labels, and multiple sets of non-quality labels. The set of feature vectors contains multiple feature vectors generated by respectively extracting predetermined features from multiple digital data representing measurement values measured from an object. The set of quality labels contains multiple quality labels respectively corresponding to each of the multiple digital data and indicating the quality of the object. Each of the multiple sets of non-quality labels contains multiple non-quality labels of a type expected to be irrelevant to the quality of the object, and the multiple non-quality labels respectively correspond to each of the multiple digital data. Calculate the variance of the clustering accuracy for each of the multiple sets of non-quality labels, thereby calculating multiple variances respectively corresponding to each of the multiple sets of non-quality labels. The clustering accuracy is the clustering accuracy when clustering subsets obtained by dividing the multiple feature vectors by each of the multiple elements respectively represented by the multiple non-quality labels using the set of quality labels. Generate a screen image capable of determining, using the multiple variances, at least one type of non-quality label that has an adverse effect on the quality of the multiple digital data.

Citation Information

Patent Citations

  • Information processing device, method and program

    CN102163208A

  • Fuzzy C-means clustering method of minimum variance optimization initial cluster center

    CN107330458A