Program and training methods

By integrating low-precision classes and training a modified classification model, the technique addresses classification accuracy issues, enhancing discriminability and overall performance in predicting multiple classes.

JP2026064335APending Publication Date: 2026-04-14BROTHER KOGYO KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
BROTHER KOGYO KK
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing techniques for generating prediction models struggle with enhancing discriminability between classes, leading to suboptimal classification accuracy, particularly when integrating classes with low-precision relationships.

Method used

Integrate classes with low-precision relationships into a single class, reducing the total number of classes, and train a modified classification model to improve classification accuracy by integrating multiple Type 1 classes associated with low-accuracy groups, using a combination of convolutional neural networks and support vector machines.

Benefits of technology

Enhances classification accuracy by reducing the impact of low-precision classes, allowing the model to effectively distinguish between integrated classes, thereby improving overall classification performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064335000001_ABST
    Figure 2026064335000001_ABST
Patent Text Reader

Abstract

Generate a predictive model that performs classification of multiple classes. [Solution] The first classification model is trained to classify multiple Type 1 training images into multiple Type 1 classes. The second classification model is trained to classify multiple features into multiple Type 2 classes using multiple features generated by the trained first classification model using multiple Type 2 training images. The trained second classification model is used to classify multiple features generated by the trained first classification model using multiple test images into multiple Type 2 classes. The first classification model is then trained to have a reduced total number of classes by merging multiple Type 1 classes based on the classification results of multiple Type 2 classes associated with multiple test images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to a technique for generating a prediction model.

Background Art

[0002] Patent Document 1 discloses a technique for classifying a sample having a plurality of trace amounts into any one of a plurality of classes. In this technique, a group of feature quantities necessary for class determination of an unknown sample whose belonging class is unknown is selected from the group of feature quantities. This selection process includes a quantification process of quantifying the discriminability between two classes by each feature quantity of the selected group of feature quantities by pairwise coupling that combines two out of N classes using a learning dataset, and an optimization process of aggregating the quantified discriminabilities for all of the pairwise couplings and selecting a combination of a group of feature quantities that optimizes the aggregation result.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When a combination of a group of feature quantities that optimizes the aggregation result is selected, there are cases where the selected group of feature quantities cannot enhance the discriminability between two classes. Thus, there has been room for improvement in generating a prediction model for classifying a plurality of classes.

[0005] This specification discloses a technique for generating a prediction model for classifying a plurality of classes.

Means for Solving the Problems

[0006] The technique disclosed in this specification can be realized as the following application examples.

[0007] [Application Example 1] A program comprising: a function to train a first classification model, which is a prediction model including one or more convolutional layers, to classify each of the multiple first-type learning images into a first-type class associated with the corresponding label information from among multiple first-type classes representing different types of images, using a plurality of first-type learning images and a plurality of label information indicating the type of each of the plurality of first-type learning images; a function to train a second classification model, which is a prediction model different from the first classification model, to classify the plurality of features into multiple second-type classes representing different types of images, using a plurality of features generated by the trained first classification model using a plurality of second-type learning images; and a plurality of functions generated by the trained first classification model using a plurality of second-type learning images. A program that enables a computer to implement the following functions: a function to classify multiple features generated using test images into multiple Type II classes using a trained Type II classification model; and a function to train a Type I classification model modified to have a total number of classes reduced by the integration of multiple Type I classes, wherein the multiple Type I classes to be integrated are Type I classes associated with a low-precision group which is a group of multiple Type II classes having a predetermined low-precision relationship, and the predetermined low-precision relationship is a relationship indicated by the classification results of the multiple Type II classes associated with the multiple test images, which indicates that the classification precision of one or more Type II classes in the low-precision group does not meet a criterion.

[0008] With this configuration, by integrating multiple Type 1 classes that are associated with low-accuracy groups including Type 2 classes whose classification accuracy does not meet the criteria, Type 1 classes that may reduce classification accuracy are integrated into other classes, and a modified Type 1 classification model is trained to have a total number of classes reduced by the integration. Thus, a modified Type 1 classification model that has been trained to appropriately classify multiple classes can be generated.

[0009] Furthermore, the technologies disclosed herein can be implemented in various forms, for example, in the form of a training method and a data processing device for performing training, a computer program for realizing the functions of such a method or device, a recording medium (e.g., a non-temporary recording medium) on which the computer program is recorded, and so on. [Brief explanation of the drawing]

[0010] [Figure 1] This is an explanatory diagram showing a data processing device as one embodiment. [Figure 2] This is a perspective view of the digital camera 110 and the multifunction printer 900. [Figure 3] (A) is a schematic diagram showing an example of a label sheet. (B) is a diagram showing an example of a target image used for testing. (C)-(G) are diagrams showing examples of training images associated with classes CL0-CL4. [Figure 4] This is a block diagram representing an example of the first classification model 310. [Figure 5] This flowchart shows an example of the training process for classification models 310 and 320. [Figure 6] This is a block diagram representing the structure of dataset 330. [Figure 7] This is a flowchart illustrating an example of the training process for the first classification model 310. [Figure 8] This is a flowchart illustrating an example of the training process for the second classification model 320. [Figure 9] This diagram illustrates the overview of how classification accuracy is calculated. [Figure 10] This is a flowchart illustrating an example of the class integration process. [Figure 11] This figure shows an example of a modified dataset. [Figure 12] This is a block diagram showing an example of the modified first classification model. [Figure 13] This diagram shows how the classification results of multiple training images change. [Figure 14] This is a flowchart illustrating an example of the inspection process. [Figure 15] It is a flowchart showing another embodiment of the process of class integration.

Embodiments for Carrying Out the Invention

[0011] A. First Embodiment: A1. Device Configuration: FIG. 1 is an explanatory diagram showing a data processing device as an example. The data processing device 200 is, for example, a personal computer. The data processing device 200 performs various data processes for inspecting the appearance of an article (for example, a label sheet provided on a product such as a multifunction device). Hereinafter, it is assumed that the appearance of the label sheet 800 provided on the multifunction device 900 is inspected.

[0012] The data processing device 200 includes a processor 210, a storage device 215, a display unit 240, an operation unit 250, a graphics processing unit 260 (referred to as GPU 260), and a communication interface 270. These elements are connected to each other via a bus. The storage device 215 includes a volatile storage device 220 and a non-volatile storage device 230.

[0013] The processor 210 is a device configured to perform data processing, and is, for example, a Central Processing Unit (CPU) or a System on a chip (SoC). The volatile storage device 220 is, for example, a Dynamic Random Access Memory (DRAM), and the non-volatile storage device 230 is, for example, a flash memory. The non-volatile storage device 230 stores the data of the first program 231, the second program 232, the first classification model 310, the second classification model 320, and the data set 330, respectively. In this embodiment, the classification models 310 and 320 are program modules that form prediction models, respectively. Details of the programs 231 and 232, the classification models 310 and 320, and the data set 330 will be described later.

[0014] The display unit 240 is a device configured to display images, such as a liquid crystal display or an organic EL display. The operation unit 250 is a device configured to receive operations by a user, such as buttons, levers, or a touch panel arranged on top of the display unit 240. The display unit 240 and the operation unit 250 may form a so-called touch screen. The user can input various requests and instructions into the data processing device 200 by operating the operation unit 250. The display unit 240 may display operation elements (e.g., buttons, sliders, etc.), and the displayed elements may be operated through the operation of the operation unit 250.

[0015] The GPU 260 is an arithmetic device configured to execute various numerical operations such as image processing and machine learning. The GPU 260 executes various operations according to the instructions of the processor 210. Note that a driver program (not shown) for controlling the GPU 260 may be provided by the manufacturer of the GPU 260.

[0016] The communication interface 270 is an interface for communicating with other devices (e.g., including one or more of a USB interface, a wired LAN interface, a wireless interface of IEEE 802.11, an interface of an industrial camera (e.g., CameraLink, CoaXPress, etc.)). The digital camera 110 is connected to the communication interface 270. The digital camera 110 is used for photographing the label sheet 800.

[0017] FIG. 2 is a perspective view of the digital camera 110 and the multifunction device 900. In this embodiment, the label sheet 800 is attached to the first side surface 901 of the multifunction device 900. The digital camera 110 photographs a portion of the multifunction device 900 including the label sheet 800 for inspection of the label sheet 800. The relative arrangement between the multifunction device 900 and the digital camera 110 is adjusted to a predetermined reference arrangement suitable for photographing so that the distortion of the label sheet 800 in the photographed image is small.

[0018] A2. Label Sheet: Figure 3(A) is a schematic diagram showing an example of a label sheet. In this embodiment, the label sheet 800 is a rectangular sheet. The label sheet 800 can represent various elements (e.g., text, marks, logotypes, etc.). In this embodiment, the label sheet 800 is affixed to the first side surface 901 of the multifunction printer 900 (Figure 2) during its manufacture.

[0019] Figure 3(B) shows an example of a target image used for inspection. Target image I10 represents the portion of the image captured by the digital camera 110 of the multifunction printer 900 that represents the label sheet 800. In this embodiment, target image I10 is a rectangular image having two sides parallel to the first direction Dx and two sides parallel to the second direction Dy which is perpendicular to the first direction Dx.

[0020] A3. Classification Model: Figure 4 is a block diagram representing an example of the first classification model 310. In this embodiment, the first classification model 310 is a classifier that uses a convolutional neural network. The first classification model 310 is trained to classify an image IM representing a captured label sheet 800 into one of several classes (in the example in Figure 4, five classes CL0-CL4). Various defects may occur in the label sheet 800 during the manufacturing of the multifunction printer 900. For example, the label sheet 800 may have various defects such as scratches, stains, missing parts of elements, or peeling of parts of the label sheet 800. In this embodiment, the multiple classes include one class that represents an image of a normal label sheet 800 and multiple classes that represent images of a defective label sheet 800. In the example in Figure 4, class CL0 represents normal, and classes CL1-CL4 represent defective. Hereinafter, class CL0 will also be referred to as normal class CL0. Classes CL1-CL4 that represent defective images will also be referred to as defective classes CL1-CL4. Multiple anomaly classes CL1-CL4 correspond to different types of anomalies (details below).

[0021] In this embodiment, the first classification model 310 has an architecture similar to that of a convolutional neural network called VGG16. The first classification model 310 has a pre-processing unit CN and a post-processing unit FCN that follows the pre-processing unit CN. The pre-processing unit CN has multiple convolutional layers and multiple pooling layers (e.g., max pooling or average pooling). The convolutional layers perform a so-called convolution process on the input data (image or feature map output from the previous layer) using filters (also called kernels) associated with the convolutional layer. As a result, the convolutional layers generate feature maps that represent local features. The pooling layers perform a so-called downsampling on the feature maps output from the previous layer. As a result, the pooling layers generate feature maps with reduced spatial resolution. Spatial resolution is represented by the number of pixels in the first direction Dx and the number of pixels in the second direction Dy. In the example in Figure 4, an RGB bitmap image with a spatial resolution of "224*224" is used as the image IM (the spatial resolution is expressed as "number of pixels in the first direction Dx * number of pixels in the second direction Dy"). The pre-processing unit CN reduces the spatial resolution in the following order: "224*224", "112*112", "56*56", "28*28", "14*14", and "7*7". Although not shown in the diagram, the feature quantities of each pixel position are represented by two or more components. Hereafter, the total number of components representing the feature quantity of a single pixel position will also be called the feature dimension. For example, the feature dimension of the image IM is 3 (RGB). The pre-processing unit CN may increase the feature dimension. For example, the pre-processing unit CN may increase the feature dimension in the following order: 64, 128, 256, and 512.

[0022] The post-processing unit FCN has multiple fully connected layers. The post-processing unit FCN generates confidence data OD by linear transformation of the feature map output from the pre-processing unit CN. The confidence data OD represents the confidence level of each of the multiple classes CL0-CL4. The class associated with the highest confidence level indicates the type of image IM. The post-processing unit FCN gradually reduces the dimensionality of the data through multiple fully connected layers. In the example in Figure 4, the post-processing unit FCN has three fully connected layers fc1-fc3. The dimensionality of the data output from the fully connected layers fc1-fc3 is 4096, 256, and 5.

[0023] The second classification model 320 (Figure 1), although not shown in the illustration, is a so-called support vector machine. As will be described later, the second classification model 320 is trained to classify the feature vector F2 output from the fully connected layer fc2, which is before the last fully connected layer fc3 of the first classification model 310, into one of five classes CL0-CL4.

[0024] A4. Training process: Figure 5 is a flowchart illustrating an example of the training process for classification models 310 and 320. The operator inputs a training start instruction to the data processing device 200 (Figure 1) by operating the control unit 250 of the data processing device 200. The processor 210 starts the training process in accordance with the start instruction. The processor 210 proceeds with the training process according to the first program 231.

[0025] In S110, the processor 210 retrieves the training set data from the dataset 330 from the storage device 215 (non-volatile storage device 230 in this embodiment). The dataset 330 contains multiple training sets. A training set is a set of training images and training labels. A training image is an image representing a label sheet. A training label indicates the class that corresponds to the training image from among multiple classes CL0-CL4. Training labels are also called training data.

[0026] Figures 3(C)-3(G) show examples of training images associated with classes CL0-CL4, respectively. Training image I80a in Figure 3(C) is an example of a training image associated with the 0th class CL0. Training image I80a represents a normal label sheet 800.

[0027] The training image I81a in Figure 3(D) is an example of a training image associated with the first class CL1. Training image I81a represents a label sheet 800 with a scratch 710. The first class CL1 is the class that indicates a scratch.

[0028] The training image I82a in Figure 3(E) is an example of a training image that corresponds to the second class CL2. Training image I82a represents a label sheet 800 with stains 720. The second class CL2 is the class that indicates stains.

[0029] The training image I83a in Figure 3(F) is an example of a training image that corresponds to the third class CL3. Training image I83a represents a label sheet 800 with missing elements (strings, marks, etc.) 730. In the example in Figure 3(F), part of the string 735 is missing. The third class CL3 is a class that indicates missing elements.

[0030] The training image I84a in Figure 3(G) is an example of a training image that corresponds to the fourth class CL4. Training image I84a represents a label sheet 800 with a curl 740. In the example in Figure 3(G), the upper right corner of the label sheet 800 is curled. The fourth class CL4 is a class that indicates a partial curl on the label sheet 800.

[0031] Multiple target images are used to examine multiple label sheets 800 (for example, target image I10 (Figure 3(B))). The position, orientation, and size of the label sheets 800 within the target images can vary among the multiple target images. To train classification models 310, 320 to appropriately classify such target images, dataset 330 includes multiple training images in which one or more of the position, orientation, and size of the label sheets 800 within the training images differ from one another.

[0032] Furthermore, among multiple label sheets 800 having the same type of defect, the configuration of the defect may differ from one another. For example, the size of the scratch, the shape of the scratch, and the location of the scratch on the label sheet 800 may differ among multiple label sheets 800 having the defect. In order to train classification models 310 and 320 to appropriately classify such defects in label sheets 800, the dataset 330 includes multiple training images that correspond to the same type of defect and show defects with different configurations from one another. For example, multiple training images that correspond to the first class CL1 include multiple training images in which one or more of the following configurations differ from one another: the size of the scratch, the shape of the scratch, and the location of the scratch on the label sheet 800. Multiple training images that correspond to the second class CL2 include multiple training images in which one or more of the following configurations differ from one another: the size of the stain, the shape of the stain, the location of the stain on the label sheet 800, and the color of the stain. Multiple training images that correspond to the third class CL3 include multiple training images in which one or more of the following configurations differ from one another: the size of the chip, the shape of the chip, and the location of the chip on the label sheet 800. The multiple training images associated with the fourth class CL4 include multiple training images in which one or more of the following characteristics differ from one another: the size of the peel, the shape of the peel, and the position of the peeled portion on the label sheet 800.

[0033] The data for dataset 330, which includes multiple training images, is generated in advance. There are various methods for generating the data for dataset 330. For example, training image data associated with the 0th class CL0 may be generated by photographing a label sheet 800 without defects with a digital camera 110. Training image data associated with an abnormal class may be generated by photographing a label sheet 800 with defects with a digital camera 110. Training image data for multiple abnormal classes CL1-CL4 may be generated using multiple label sheets 800 having different types of defects. Multiple training image data may be generated by photographing a single label sheet 800 multiple times. New training image data may be generated by processing (also called data augmentation) the images representing the photographed label sheets 800.

[0034] Figure 6 is a block diagram representing the structure of dataset 330. Dataset 330 is divided into training set 330a and test set 330b. Training set 330a and test set 330b each contain multiple training sets for each class CL0-CL4. The class associated with a training set is indicated by the training label of the training set. The figure shows the training images for each class CL0-CL4 indicated by the training label. For example, training images I80a-I84a in training set 330a are associated with classes CL0-CL4, respectively. Training images I80b-I84b in test set 330b are associated with classes CL0-CL4, respectively. Training set 330a is used to tune the parameters of classification models 310 and 320. Test set 330b is used to evaluate the tuned classification models 310 and 320. In S110 (Figure 5), the processor 210 acquires the data for the training set 330a. The training set 330a and the test set 330b may be predetermined. Alternatively, the processor 210 may use random numbers to divide the dataset 330 into the training set 330a and the test set 330b.

[0035] In S115 (Figure 5), the processor 210 trains the first classification model 310 using the training set 330a. At this stage, the total number of classes in the first classification model 310 is P, and the total number of classes that show abnormalities is N. In the example in Figure 4, P=5, N=4, and P=1+N.

[0036] Figure 7 is a flowchart showing an example of the training process for the first classification model 310. In S210, the processor 210 acquires the data for the training set 330a (Figure 6). If the data for the training set 330a has already been acquired in S110 (Figure 5), S210 may be omitted.

[0037] In S220, the processor 210 initializes multiple computational parameters of the first classification model 310. In this embodiment, the multiple computational parameters include multiple weights and biases for multiple filters of multiple convolutional layers and multiple weights and biases for multiple fully connected layers. Each computational parameter is set to, for example, a random value.

[0038] In S230, the processor 210 generates confidence data OD by performing calculations on the first classification model 310 (Figure 4) using the training image data. The processor 210 selects multiple training sets from the training set 330a that were not used in the process shown in Figure 7, and generates multiple confidence data ODs using multiple training images from multiple training sets. Here, the processor 210 converts the spatial resolution of the training images to a spatial resolution suitable for the first classification model 310 by performing a training image resolution conversion process in order to input the training image data into the first classification model 310.

[0039] In S240, the processor 210 calculates an error value that represents the difference between the confidence data OD and the target data associated with the training images used to generate the confidence data OD. The target data represents the target value (i.e., the correct answer) of the confidence data OD. Specifically, the confidence level of the class indicated by the training label associated with the training image is 1, and the confidence level of the other classes is 0. The error value is calculated based on a predetermined loss function (such an error value is also called a loss value). In this embodiment, so-called cross-entropy is used as the loss function. Cross-entropy can be used as an index value that shows the amount of deviation between two probability distributions (e.g., confidence data OD and target data). The loss function may be any other function (e.g., sum of squared errors). In S240, the processor 210 calculates multiple error values ​​using multiple confidence data ODs.

[0040] In S250, the processor 210 adjusts several computational parameters of the first classification model 310 using multiple error values. For example, the processor 210 adjusts the computational parameters so that the index value calculated using the multiple error values ​​becomes small (the index value may be various values ​​that correlate with the magnitude of the difference, such as the mean, maximum, median, and sum). As an algorithm for adjusting the computational parameters, for example, an algorithm using backpropagation and gradient descent may be employed.

[0041] Processor 210 may have GPU 260 perform some or all of the calculations in S230-S250.

[0042] In S260, the processor 210 determines whether training is complete. The conditions for training completion can be various. In this embodiment, the processor 210 selects several unused training sets from the training set 330a in the process shown in Figure 7, and uses the selected training sets to perform calculations on the first classification model 310. This generates several confidence data ODs for evaluation. The processor 210 then determines that training is complete if the conditions indicating that several error values ​​obtained from the multiple confidence data ODs for evaluation are small are met (for example, the average error value is less than or equal to a predetermined threshold value).

[0043] If training is not complete (S260: No), the processor 210 proceeds to S230 and continues training. If training is complete (S260: Yes), in S270, the processor 210 stores the data representing the trained first classification model 310 in the storage device 215 (for example, the non-volatile storage device 230). Then, the processor 210 terminates the process shown in Figure 7, i.e., the process shown in S115 of Figure 5.

[0044] In S120, the processor 210 acquires multiple features output from the layer immediately preceding the final layer of the first classification model 310. In this embodiment, the final layer is the fully connected layer fc3 (Figure 4), the layer immediately preceding the final layer is the fully connected layer fc2, and the feature output from the layer immediately preceding the final layer is feature F2. Feature F2 is the feature used by the final fully connected layer fc3 to calculate the confidence data OD. Such feature F2 may represent different image features between classes CL0-CL4. The processor 210 acquires data for multiple feature F2 by performing calculations on the trained first classification model 310 using data from multiple training images. In this embodiment, multiple training images used are those used to train the first classification model 310. Specifically, the processor 210 uses the training image data used to calculate the confidence data OD in S230 (Figure 7).

[0045] In S125 (Figure 5), the processor 210 trains the second classification model 320 using multiple sets of features F2 and training labels associated with features F2. The training labels associated with features F2 are training labels associated with the training images used to calculate features F2. Features F20a-F24a in Figure 6 are examples of features F2 associated with classes CL0-CL4.

[0046] Figure 8 is a flowchart illustrating an example of the training process for the second classification model 320. In S310, the processor 210 acquires training data. The training data includes multiple sets of feature F2 data and training label data associated with feature F2. The processor 210 refers to the training set 330a (Figure 6) and acquires multiple training labels associated with the multiple feature F2 acquired in S120 (Figure 5).

[0047] In S320 (Figure 8), the processor 210 initializes the training parameters. These training parameters are used to train the second classification model 320. In this embodiment, the second classification model 320 is a support vector machine. A support vector machine classifies multiple samples into multiple classes using a boundary (also called a hyperplane). In training a support vector machine, the boundary is calculated by maximizing the margin. A technique called the kernel trick is used for training. The kernel trick is a technique for linearly separating nonlinear data in a high-dimensional space. The kernel trick reduces the computational complexity in high-dimensional space for optimizing the boundary of the support vector machine by using a kernel function. Various functions can be used as kernel functions, such as polynomial kernels, Gaussian kernels, and sigmoid kernels. In this embodiment, a Gaussian kernel is used. The Gaussian kernel is also called a Radial Basis Function (RBF) kernel.

[0048] The training parameters include various parameters that correspond to the kernel function. When a Gaussian kernel is used, the training parameters include a gamma value and a regularization parameter. The gamma value is a parameter that corresponds to the reciprocal of the magnitude of the Gaussian kernel. A larger gamma value reduces the influence range of the sample. The gamma value changes the balance between classification accuracy and margin maximization. A smaller gamma value may decrease classification accuracy, but the margin may increase and the boundary may become simpler. The regularization parameter indicates the magnitude of the penalty for misclassification during training. A smaller regularization parameter may increase the likelihood of misclassification, but decrease the likelihood of overfitting.

[0049] The gamma value and regularization parameters are determined experimentally in advance so that the second-classification model 320 can be properly trained. The processor 210 may initialize the gamma value and regularization parameters to predetermined values. Similarly, if other training parameters are used, their initial values ​​may be determined experimentally in advance.

[0050] One boundary of a support vector machine is the ability to classify multiple samples into two classes. Various configurations can be used for support vector machines that perform three or more classifications. For example, a combination of multiple two-class classifications (one-to-other) or multiple two-class classifications (one-to-one) may be used.

[0051] In S330 (Figure 8), the processor 210 trains the second classification model 320 according to the training parameters. In this embodiment, the processor 210 uses multiple sets of features F2 and training labels to train the second classification model 320 to classify each feature F2 into a corresponding class among P classes (here, five classes CL0-CL4). This calculates the hyperplane of the second classification model 320. The training method for the second classification model 320 can be varied. For example, the library called "scikit-learn" can be used to train a support vector machine using gamma values ​​and regularization parameters. The processor 210 may also have the GPU 260 perform some or all of the calculations for training the second classification model 320.

[0052] In S340, the processor 210 stores the data representing the trained second classification model 320 in the storage device 215 (for example, the non-volatile storage device 230). Then, the processor 210 completes the process shown in Figure 8, i.e., the process shown in S125 of Figure 5.

[0053] In S130, the processor 210 acquires data from the test set 330b of the dataset 330 (Figure 6). Hereafter, the training images included in the test set 330b (for example, training images I80b-I84b) will also be referred to as test images, and the training labels associated with the test images will also be referred to as test labels.

[0054] In S135 (Figure 5), the processor 210 obtains multiple features F2 from the trained first classification model 310 using the test set 330b. The processor 210 generates the feature F2 data by performing calculations on the first classification model 310 (Figure 4) using the test image data. The features F2 generated by the trained first classification model 310 can appropriately represent the different image features between classes CL0-CL4. Feature F20b-F24b in Figure 6 are examples of features F2 that correspond to classes CL0-CL4.

[0055] In S140 (Figure 5), the processor 210 uses multiple sets of features F2 and test labels to calculate the classification accuracy of the trained second classification model 320. Figure 9 is a diagram illustrating the calculation of classification accuracy. In this embodiment, the processor 210 calculates the percentage for each combination of the ground truth class and the class predicted by the second classification model 320. For example, the first column C1 in the figure shows the percentage of prediction results for each class CL0-CL4 when the ground truth is class 0 CL0. For example, the percentage of prediction results for class 1 CL1 (here, 1.1%) indicates that 1.1% of the features F2 associated with class 0 CL0 are incorrectly classified as class 1 CL1. Percentages for other combinations of ground truth classes and predicted classes are calculated similarly.

[0056] As illustrated, the correct classification rate (i.e., the percentage of correctly classified items) may vary depending on the type of defect in the label sheet 800. Furthermore, multiple types of defects that are not easily distinguishable using images may be misclassified into each other's classes. For example, a peeled label sheet 800 (Figure 3(G)) can erase a portion of an element from the image, similar to a missing element (Figure 3(F)). Therefore, distinguishing between a missing element and a peeled label sheet 800 using images may not always be easy. In this case, as shown in Figure 9, the correct classification rate R33 for the third class CL3 (missing element) and the correct classification rate R44 for the fourth class CL4 (peeled label sheet) may be lower than the correct classification rates of other classes. Also, the percentage of features F2 in the third class CL3 that are misclassified into the fourth class CL4 (R43) and the percentage of features F2 in the fourth class CL4 that are misclassified into the third class CL3 (R34) may be higher than the misclassification rates between other classes.

[0057] The first classification model 310 is trained to classify a plurality of classes CL0 - CL4 to be classified. Even when the plurality of classes CL0 - CL4 to be classified includes a plurality of classes (e.g., the third class CL3 and the fourth class CL4) that are not easily distinguishable using images, the first classification model 310 is trained to classify the plurality of classes that are not easily distinguishable. Therefore, the trained first classification model 310 can classify each class CL0 - CL4 with a certain degree of accuracy. On the other hand, the second classification model 320 is a prediction model different from the first classification model 310. As shown in FIG. 9, the second classification model 320 may not be able to successfully classify a plurality of classes that are not easily distinguishable using images.

[0058] A plurality of classes that are not easily distinguishable can affect the classification of other classes. For example, by training the first classification model 310 to distinguish between the third class CL3 (chipped) and the fourth class CL4 (curled), the accuracy of the classification of the other classes CL0 - CL2 by the first classification model 310 may decrease. Such a decrease in accuracy may be undesirable. In this embodiment, a plurality of classes that are not easily distinguishable are integrated.

[0059] In S145 (FIG. 5), the processor 210 forms a first classification model and a dataset for Q-class classification (Q < P) by integrating a plurality of classes that are not easily distinguishable. FIG. 10 is a flowchart showing an example of the process of class integration. In S410, the processor 210 selects an unprocessed combination of two classes from the P classes CL0 - CL4. Hereinafter, the two selected classes are referred to as target classes CLa and CLb. In this embodiment, two classes indicating anomalies are selected.

[0060] In S415, the processor 210 determines whether the correct classification rates R1a and R1b for target classes CLa and CLb, respectively, are less than or equal to the first threshold R1th. The first correct classification rate R1a is the proportion of feature quantities F2 of the first target class CLa that are correctly classified into the first target class CLa. The second correct classification rate R1b is the proportion of feature quantities F2 of the second target class CLb that are correctly classified into the second target class CLb. The first threshold R1th is the lower limit of the acceptable correct classification rate and may be various values. In this embodiment, the first threshold R1th is predetermined. The first threshold R1th may be a value within the range of 40% or more and 70% or less (for example, 50%).

[0061] If both the correct classification rates R1a and R1b are less than or equal to the first threshold R1th (S415: Yes), then in S420, the processor 210 determines whether the mutual misclassification rates R2a-b and R2b-a of the target classes CLa and CLb, respectively, are greater than or equal to the second threshold R2th. The first mutual misclassification rate R2a-b is the proportion of feature quantities F2 of the first target class CLa that are incorrectly classified into the second target class CLb. The second mutual misclassification rate R2b-a is the proportion of feature quantities F2 of the second target class CLb that are incorrectly classified into the first target class CLa. The second threshold R2th is the upper limit of the acceptable mutual misclassification rate and may be various values. In this embodiment, the second threshold R2th is predetermined. The second threshold R2th may be a value within the range of 10% or more and 30% or less (for example, 20%).

[0062] If both the mutual misclassification rates R2a-b and R2b-a are greater than or equal to the second threshold R2th (S420: Yes), in S425, the processor 210 classifies the target classes CLa and CLb into the low-precision group, which is a group with a low-precision relationship. Then, the processor 210 proceeds to S430.

[0063] In S415, if either or both of the correct classification rates R1a and R1b are greater than the first threshold R1th (S415: No), the processor 210 skips S420 and S425 and proceeds to S430.

[0064] If, in S420, one or both of the cross-classification rates R2a-b and R2b-a are less than the second threshold R2th (S420: No), the processor 210 skips S425 and proceeds to S430.

[0065] In S430, processor 210 determines whether all combinations of the two classes indicating an anomaly have been processed. If there are any unprocessed combinations remaining (S430: No), processor 210 proceeds to S410 and processes the new combinations.

[0066] If all combinations have been processed (S430: Yes), in S435, the processor 210 integrates the classes for each low-precision group (in this embodiment, a pair of classes). Specifically, the processor 210 performs the following processing: (First step) The training labels of each low-precision group are combined to generate a modified dataset. (Second step) Calculate the number of classes Q after integration. (Third step) Change the dimension of the output of the first classification model from P to Q.

[0067] Figure 11 shows an example of a modified dataset. This dataset 330z represents the dataset obtained by merging the third class CL3 and the fourth class CL4 of dataset 330 in Figure 6. The codes CL3o and CL4o represent the third class and the fourth class before the merger, respectively (these classes are referred to as the old third class CL3o and the old fourth class CL4o). The processor 210 updates the training labels that represent the old third class CL3o and the old fourth class CL4o to training labels that represent the new third class CL3. As a result, the training image of the old third class CL3o (missing part) (e.g., training image I83a (Figure 3(F))) and the training image of the old fourth class CL4o (peeling part) (e.g., training image I84a (Figure 3(G))) are associated with the same third class CL3. The training set 330az is formed from the training set 330a (Figure 6) by merging classes CL3o and CL4o. Test set 330bz is formed from test set 330b (Figure 6) by integrating classes CL3o and CL4o. The processor 210 stores the data of the modified dataset 330z in the storage device 215 (e.g., non-volatile storage device 230).

[0068] If the total number of low-precision groups is 2 or more, the processor 210 merges the low-precision groups one by one. Merging one low-precision group reduces the total number of classes by 1. The processor 210 calculates the number of classes Q after merging, depending on the total number of low-precision groups being merged. In the example in Figure 11, the number of classes Q after merging is 4. The total number of classes showing abnormalities M is 3 (Q = 1 + M).

[0069] Processor 210 changes the number of dimensions of the output of the first classification model 310 from P to Q. Figure 12 is a block diagram showing an example of the modified first classification model. This first classification model 310z represents the classification model obtained by changing the number of dimensions of the output of the last fully connected layer fc3 of the first classification model 310 in Figure 4 from P to Q (here, from 5 to 4). The configuration of the first classification model 310z is the same as the configuration of the first classification model 310 in Figure 4, except that the number of dimensions of the output of the last fully connected layer fc3z is 4. The configuration of the pre-processing unit CN of the first classification model 310z is the same as the configuration of the pre-processing unit CN of the first classification model 310 (Figure 4). The configuration of the post-processing unit FCNz of the first classification model 310z is the same as the configuration of the post-processing unit FCN of the first classification model 310 (Figure 4), except that the configuration of the last fully connected layer fc3z. The confidence data ODz generated by the first classification model 310z represents the confidence level of each of the Q classes CL0-CL3. The processor 210 stores parameter data representing the configuration of the modified first classification model 310z (e.g., the size of the filters in each convolutional layer, the number of dimensions of the input and output of each fully connected layer, etc.) in the storage device 215 (e.g., the non-volatile storage device 230).

[0070] As described above, the processor 210 forms the first classification model 310z and the dataset 330z for Q-class classification. Then, the processor 210 completes the process shown in Figure 10, that is, the process shown in S145 of Figure 5.

[0071] In S150, the processor 210 trains the modified first classification model 310z using the modified training set 330az. The training method for the first classification model 310z is the same as the training method for the first classification model 310 in S115. The processor 210 trains the first classification model 310z by performing the process shown in Figure 7 using the modified training set 330az. In this embodiment, all computational parameters of the first classification model 310z are adjusted through training.

[0072] In S155 (Figure 5), the processor 210 stores the data of the trained first classification model 310z in the storage device 215 (for example, the non-volatile storage device 230). With this, the processor 210 completes the process shown in Figure 5.

[0073] Figure 13 illustrates the change in classification results for multiple training images. The figure shows feature spaces SP and SPz. The feature space SP on the left represents multiple features (e.g., feature F2) of multiple training images processed by the first classification model 310 (Figure 4) before integration. The feature space SPz on the right represents multiple features (e.g., feature F2) of multiple training images processed by the first classification model 310z (Figure 12) after integration. For illustrative purposes, each feature space SP and SPz are represented in a simplified two-dimensional form in the figure. In reality, feature spaces SP and SPz can have three or more dimensions.

[0074] In each feature space SP and SPz, hatched circles represent features of the training image corresponding to the 0th class CL0 (normal) (e.g., training image I80a (Figure 3(C))). In the left feature space SP, upward-pointing triangles represent features of the training image corresponding to the 3rd class CL3 (missing) (e.g., training image I83a (Figure 3(F))). Downward-pointing triangles represent features of the training image corresponding to the 4th class CL4 (curled) (e.g., training image I84a (Figure 3(G))). In the right feature space SPz, squares represent features of the training image corresponding to the combined 3rd class CL3 (missing + curled) (e.g., training images I83a, I84a).

[0075] The feature space SP on the left shows the regions for classes CL0, CL3, and CL4. These regions represent the classification results of the trained first classification model 310. As shown in the figure, the training images for the third class CL3 and the fourth class CL4 are correctly classified. However, the training image I80j, which corresponds to the zeroth class CL0, is misclassified as the third class CL3. The boundary separating the difficult-to-distinguish third class CL3 and fourth class CL4 can have a complex shape. During the training of the first classification model 310, several computational parameters of the first classification model 310 are adjusted to form such a boundary. Such adjustments can affect the boundaries between other classes. For example, the possibility of misclassification of other classes (e.g., the zeroth class CL0) may increase.

[0076] The feature space SPz on the right shows the regions for classes CL0 and CL3. These regions represent the classification results of the trained first classification model 310z. As shown in the figure, the training images for the 0th class CL0 and the 3rd class CL3 are correctly classified. By merging multiple classes that are difficult to distinguish, the merged classes can be separated by a boundary with a simple shape. Compared to before merging, the difficulty of training the first classification model 310z is reduced after merging. Therefore, the possibility of misclassification by the trained first classification model 310z is reduced.

[0077] A5. Inspection process: Figure 14 is a flowchart illustrating an example of the inspection process. The inspection process may be performed by the data processing device 200 (Figure 1) used to train the first classification model 310z (Figure 12). Alternatively, a data processing device different from the data processing device 200 may perform the inspection process. Hereafter, we will assume that the data processing device 200 performs the inspection process. The operator inputs an instruction to start the inspection process to the data processing device 200 (Figure 1) by operating the operation unit 250 of the data processing device 200. The processor 210 starts the inspection process according to the inspection instruction. The processor 210 proceeds with the inspection process according to the second program 232.

[0078] In S810, the processor 210 supplies a shooting instruction to the digital camera 110 (Figure 2). The digital camera 110, in response to the shooting instruction, photographs the portion of the multifunction device 900 that includes the label sheet 800 and supplies the data of the captured image to the data processing device 200. The processor 210 obtains data for the target image, which is an image for inspection, by extracting the portion representing the label sheet 800 from the captured image (for example, target image I10 (Figure 3(B))).

[0079] There may be various methods for extracting the portion representing the label sheet 800 from the captured image. For example, the processor 210 may extract a predetermined portion of the captured image as the target image. Alternatively, the processor 210 may extract the portion representing the label sheet 800 from the captured image using template matching with a template image representing the label sheet 800, or by using an object detection model (e.g., a YOLO model) that has been trained to detect the label sheet 800.

[0080] In S820, the processor 210 generates confidence data ODz by performing calculations for the first classification model 310z (Figure 12) using the data of the target image. The class indicated by the confidence data ODz, i.e., the type of the target image, indicates the inspection result of the label sheet 800. Furthermore, in order to input the data of the target image into the first classification model 310z, the processor 210 performs a resolution conversion process on the target image to convert the spatial resolution of the target image to a spatial resolution suitable for the first classification model 310z.

[0081] In S830, the processor 210 stores the inspection result data in the storage device 215 (e.g., the non-volatile storage device 230). Then, the processor 210 terminates the process shown in Figure 14. The inspection result data can be used for various processes, such as displaying the inspection results. As shown in Figure 12, in this embodiment, the inspection result indicates either normal or abnormal. If the inspection result indicates an abnormality, a process may be performed to remove the corresponding multifunction printer 900 from the production line. Furthermore, if the label sheet 800 has an abnormality, the inspection result also indicates the type of abnormality. The operator may perform a process corresponding to the type of abnormality. For example, if the type of abnormality is a scratch (first class CL1), the operator may adjust the operation of the equipment used to manufacture the multifunction printer 900 (e.g., a belt conveyor, a robotic arm, etc.). If the type of abnormality is dirt (second class CL2), the operator may clean the equipment used to manufacture the multifunction printer 900. If the type of defect is a chip or peel (Class 3 CL3), the operator may adjust the operation of the equipment used to manufacture the label sheet 800 (e.g., a printing machine) or the equipment used to apply the label sheet 800 (e.g., a robot).

[0082] As described above, in this embodiment, the processor 210 (Figure 1) executes the following processes according to the first program 231. In S115 (Figure 5), the processor 210 trains the first classification model 310 (Figure 4) using multiple training images (e.g., training images I80a-I84a) from the training set 330a (Figure 6) and multiple training labels indicating the type of each of the multiple training images. The processor 210 trains the first classification model 310 to classify each of the multiple training images into a class from among multiple classes CL0-CL4 that corresponds to the corresponding training label. As explained in Figure 4, the first classification model 310 is a predictive model that includes one or more convolutional layers. The multiple training images in the training set 330a are examples of multiple first-type training images used to train the first classification model 310. The training labels are examples of label information indicating the type of first-type training image. The multiple classes CL0-CL4 of the first classification model 310 are examples of multiple first-kind classes that represent different types of images. Hereafter, classes CL0-CL4 of the first classification model 310 will also be referred to as first-kind classes CL0-CL4.

[0083] In S125 (Figure 5), the processor 210 trains the second classification model 320 using multiple features F2. As described in S120, the multiple features F2 are generated by the trained first classification model 310 using multiple training images. In this embodiment, multiple training images (e.g., training images I80a-I84a) from the training set 330a (Figure 6) are used. These multiple training images are examples of multiple Type II training images used to generate the features F2. The second classification model 320 is a different prediction model from the first classification model 310. The processor 210 trains the second classification model 320 to classify the multiple features F2 into multiple classes CL0-CL4. The multiple classes CL0-CL4 of the second classification model 320 are examples of multiple Type II classes that represent different types of images. Hereinafter, the classes CL0-CL4 of the second classification model 320 will also be referred to as Type II classes CL0-CL4.

[0084] In S140 (Figure 5), the processor 210 classifies multiple features F2 into multiple second-kind classes CL0-CL4 using the trained second-classification model 320. The multiple features F2 used are generated by the trained first-classification model 310 using multiple test images, as described in S135.

[0085] In S150 (Figure 5), the processor 210 trains the modified first classification model 310z (Figure 12). As described in S145, the first classification model 310z has been modified to have a reduced total number of classes Q by merging multiple first kind classes (e.g., classes CL3, CL4). The multiple classes that are merged from the multiple first kind classes CL0-CL4 of the first classification model 310 are the classes that correspond to the low-precision group from the multiple second kind classes CL0-CL4 of the second classification model 320. As described in Figure 10, the low-precision group is a group of multiple second kind classes having a predetermined low-precision relationship (S425). In this embodiment, a predetermined low-accuracy relationship indicates that the correct classification rates R1a and R1b of multiple Type 2 classes CLa and CLb included in the low-accuracy group are less than or equal to the first threshold R1th (S415: Yes), and the mutual misclassification rates R2a-b and R2b-a are greater than or equal to the second threshold R2th (S420: Yes). The correct classification rates R1a and R1b and the mutual misclassification rates R2a-b and R2b-a are examples of classification accuracy, while the first threshold R1th and the second threshold R2th are examples of values ​​that indicate the standard for classification accuracy. The lower the correct classification rates R1a and R1b, the lower the classification accuracy. The higher the mutual misclassification rates R2a-b and R2b-a, the lower the classification accuracy. The correct classification rates R1a, R1b and the mutual misclassification rates R2a-b, R2b-a (and consequently, the classification accuracy shown in Figure 9) are classification results associated with multiple test images, as explained in S140 (Figure 5), and are shown by the classification results of multiple Type 2 classes CL0-CL4 of the second classification model 320. The low-accuracy relationship represented by such correct classification rates R1a, R1b and mutual misclassification rates R2a-b, R2b-a is a relationship shown by the classification results associated with multiple test images, and is an example of a relationship that indicates that the classification accuracy of one or more Type 2 classes in the low-accuracy group does not meet the criteria.

[0086] In this configuration, by merging multiple Type I classes that are associated with a low-accuracy group containing Type II classes whose classification accuracy does not meet the criteria, Type I classes that may degrade classification accuracy are merged into other classes. For example, in the examples in Figures 11 and 12, Type I classes CL3 and CL4, which may degrade classification accuracy, are merged. Then, the modified Type I classification model 310z (Figure 12) is trained to have a total number Q of classes reduced by the merger. The modified Type I classification model 310z can appropriately classify multiple classes. In this way, the processor 210 can generate a modified Type I classification model 310z that is a model trained to appropriately classify multiple classes. Furthermore, it is permissible that multiple Type I classes in the original Type I classification model 310 include multiple classes that are not easily distinguishable by the original Type I classification model 310. Therefore, the degree of freedom in designing the original Type I classification model 310 and the dataset 330 is improved.

[0087] Furthermore, in this embodiment, as explained in S120 (Figure 5), the calculation of multiple feature quantities F2 in S125 uses multiple training images used to train the first classification model 310. That is, the multiple second-type training images used to calculate the multiple feature quantities F2 (S120) used for training the second classification model 320 (S125) are included in the multiple first-type training images used for training the first classification model 310 (S115). In this way, the same training images are used for training the first classification model 310 and the second classification model 320. The multiple feature quantities F2 used for training the second classification model 320 are generated by the first classification model 310. Since the second classification model 320 is trained using these multiple feature quantities F2, the training results of the first classification model 310 are reflected in the training of the second classification model 320. As described above, the classification results obtained using the second classification model 320 (Figure 9) can appropriately represent low-precision groups corresponding to multiple classes that are not easily classified by the first classification model 310. The processor 210 can generate a first classification model 310z that appropriately classifies multiple classes by training the first classification model 310z, which has been modified by integrating multiple first-kind classes associated with such low-precision groups.

[0088] Furthermore, in this embodiment, as explained in S130-S145 (Figure 5) and Figure 6, the low-precision relationship is indicated by the classification results of multiple test images included in a test set 330b, which is different from the training set 330a. That is, each of the multiple test images is an image that is not included in either the multiple first-type training images used to train the first classification model 310 or the multiple second-type training images used to train the second classification model 320. As described above, the feature quantity F2 generated by the first classification model 310 is used to train the second classification model 320. Therefore, the second classification model 320 may overfit to the multiple first-type training images used to train the first classification model 310, or to the multiple second-type training images used to train the second classification model 320. Even in such a case, the classification results obtained using multiple test images can appropriately indicate low-precision groups corresponding to multiple classes that are not easily classified by the first classification model 310. The processor 210 can train the first classification model 310z, which has been modified according to the appropriate low-precision groups.

[0089] Furthermore, in this embodiment, as explained in S115 and S125 (Figure 5), the total number P of multiple Type 2 classes CL0-CL4 in the second classification model 320 is the same as the total number P of multiple Type 1 classes CL0-CL4 in the first classification model 310. Therefore, the classification result obtained using the second classification model 320 (Figure 9) can appropriately show low-precision groups corresponding to multiple classes that are not easily classified by the first classification model 310.

[0090] Furthermore, in this embodiment, the first classification model 310 (Figure 4) includes two or more fully connected layers fc1-fc3 that form the final part of the first classification model 310. Feature vectors F2 are features output from a specific fully connected layer (in this embodiment, fully connected layer fc2) that precedes the last fully connected layer fc3 among the two or more fully connected layers fc1-fc3. Such feature vectors F2 can appropriately represent image features for deriving confidence data OD (i.e., classification results) output from the last fully connected layer fc3, i.e., image features that differ among multiple classes CL0-CL4. Since such feature vectors F2 are used for classification by the second classification model 320, the classification results obtained using the second classification model 320 (Figure 9) can appropriately indicate low-precision groups corresponding to multiple classes that are not easily classified by the first classification model 310.

[0091] Furthermore, in this embodiment, as shown in Figure 4, the specific fully connected layer that outputs the feature vector F2 is the second-to-last fully connected layer fc2 among the two or more fully connected layers fc1-fc3. Such a feature vector F2 can appropriately represent the image features necessary to derive the image classification result, that is, the image features that differ among multiple classes CL0-CL4.

[0092] Furthermore, in this embodiment, the second classification model 320 is a support vector machine. Compared to neural networks that include convolutional layers, support vector machines separate multiple classes by boundaries with simpler shapes. The second classification model 320 may misclassify multiple classes that are not easily distinguishable using images. The classification results obtained using the second classification model 320 (Figure 9) can appropriately show low-precision groups corresponding to multiple classes that are not easily classified by the first classification model 310.

[0093] Furthermore, in this embodiment, as explained in Figure 10, a predetermined low-precision relationship indicates that the classification accuracy of two Type 2 classes CLa and CLb forming a low-precision group does not meet the criteria. Here, the fact that the classification accuracy of two Type 2 classes CLa and CLb does not meet the criteria indicates that the correct classification rates R1a and R1b of the two Type 2 classes CLa and CLb are less than or equal to the first threshold R1th, and the misclassification rates R2a-b and R2b-a between the two Type 2 classes CLa and CLb are greater than or equal to the second threshold R2th. Such a low-precision relationship can appropriately represent a group of two classes CLa and CLb that are not easily distinguishable using images. Note that the first threshold R1th is an example of a first value which is the threshold for the correct classification rates R1a and R1b. The second threshold R2th is an example of a second value which is the threshold for the misclassification rates R2a-b and R2b-a between the two classes.

[0094] Furthermore, in this embodiment, as explained in S150 (Figure 5), the processor 210 adjusts all computational parameters of the modified first classification model 310z through training. That is, the processor 210 trains the entire modified first classification model 310z. Therefore, the trained first classification model 310z can appropriately classify images.

[0095] Furthermore, in this embodiment, as explained in Figure 4, the multiple Type 1 classes CL0-CL4 of the first classification model 310 include the normal class CL0 and multiple abnormal classes CL1-CL4. The normal class CL0 represents the image type that represents a normal object, and each of the multiple abnormal classes CL1-CL4 represents the image type that represents a label sheet 800 that has an abnormality. The multiple abnormal classes CL1-CL4 are associated with different types of abnormalities. As explained in S410 (Figure 10), the Type 2 classes CLa and CLb included in the low-precision group are both Type 2 classes that are associated with any of the abnormal classes CL1-CL4. Therefore, compared to the case where any of the abnormal classes CL1-CL4 are merged with the normal class CL0, the possibility of misclassification between normal and abnormal by the modified first classification model 310z is reduced.

[0096] Furthermore, in this embodiment, as described in S115 (Figure 5), the processor 210 uses multiple sets of training images and training labels from the training set 330a to train the first classification model 310 so that each of the multiple training images is classified into a first kind class (here, one of the first kind classes CL0-CL4) associated with the training label. These multiple training images are examples of multiple first kind training images used to train the first classification model 310. The training labels are examples of label information indicating the type of first kind training image. With this configuration, the processor 210 can train the first classification model 310 to classify images into classes indicating the type of image.

[0097] Furthermore, in this embodiment, as explained in S150 (Figure 5), the processor 210 trains the modified first classification model 310z using the modified training set 330az (Figure 11). As explained in S435 (Figure 10), the modified training set 330az is formed from training set 330a (Figure 6) by integrating classes CL3o and CL4o. That is, the training images of training set 330az are the same as the training images of training set 330a. In this way, the processor 210 trains the modified first classification model 310z using at least some of the multiple first-type training images used to train the first classification model 310. Therefore, the burden of preparing training images is reduced compared to when the training images used to train the first classification model 310 are not used to train the modified first classification model 310z.

[0098] B. Second example: Figure 15 is a flowchart illustrating another embodiment of the class integration process. The only difference from the process in Figure 10 is that S415 is omitted. The other parts of the integration process in Figure 15 are the same as the corresponding parts in Figure 10. Steps in Figure 15 that are the same as those in Figure 10 are given the same reference numerals and their explanations are omitted. The process in Figure 15 is executed at S145 (Figure 5) instead of the process in Figure 10.

[0099] In this embodiment, a predetermined low-accuracy relationship indicates that the mutual misclassification rates R2a-b and R2b-a of multiple Type II classes CLa and CLb included in the low-accuracy group are greater than or equal to the second threshold R2th (S420: Yes). The mutual misclassification rates R2a-b and R2b-a are examples of classification accuracy, and the second threshold R2th is an example of a value indicating the criterion for classification accuracy.

[0100] Thus, a given low-precision relationship indicates that the classification accuracy of two Type II classes CLa and CLb forming a low-precision group does not meet the criteria. Here, the fact that the classification accuracy of two Type II classes CLa and CLb does not meet the criteria indicates that the misclassification rates R2a-b and R2b-a between the two Type II classes CLa and CLb are greater than or equal to the second threshold R2th. Such low-precision relationships can appropriately represent groups of two classes CLa and CLb that are not easily distinguishable using images. The processor 210 can generate a first classification model 310z that appropriately classifies multiple classes by training the first classification model 310z modified according to an appropriate low-precision group. Note that the second threshold R2th is an example of a third value which is the threshold for the misclassification rates R2a-b and R2b-a between the two classes.

[0101] C. Variant: (1) The low-precision relationship is not limited to the relationship described in Figures 10 and 15, but may be any relationship that indicates that the classification accuracy of one or more Type II classes in a low-precision group, which is a group of multiple Type II classes having a low-precision relationship, does not meet the criteria. For example, in the process in Figure 10, S420 may be omitted. That is, the low-precision relationship may indicate that the correct classification rates R1a and R1b of multiple Type II classes CLa and CLb included in the low-precision group are less than or equal to the first threshold R1th (S415: Yes).

[0102] Furthermore, if the misclassification rate R2a-b from the first target class CLa to the second target class CLb is greater than or equal to a predetermined third threshold, the processor 210 may determine that the target classes CLa and CLb have a low-accuracy relationship. Here, the mutual misclassification rate R2b-a from the second target class CLb to the first target class CLa may be less than or equal to the third threshold.

[0103] Furthermore, if the correct classification rate R1a of the first target class CLa is below a predetermined fourth threshold, and the misclassification rate R2a-b from the first target class CLa to the second target class CLb is the highest among the misclassification rates from the first target class CLa to other classes, the processor 210 may determine that the target classes CLa and CLb have a low-precision relationship.

[0104] (2) The features used for classification by the second classification model are not limited to feature F2 (Figure 4), but may be various features generated by the first classification model. For example, feature F1 output from the third-to-last fully connected layer fc1 may be used. Alternatively, features output from the convolutional or pooling layers included in the pre-processing unit CN may be used. In either case, the features generated by the first classification model may represent different image features among multiple classes. Therefore, the classification results of the features by the second classification model can appropriately represent groups of multiple classes that are not easily distinguishable using images.

[0105] (3) In training the modified first classification model (e.g., Figure 5: S150), only a portion of the first classification model may be trained. For example, only the last fully connected layer fc3z (Figure 12) may be trained. Alternatively, only the subsequent processing unit FCNz may be trained. Thus, only a portion including the modified last fully connected layer may be trained. In either case, the remaining portion of the modified first classification model (e.g., first classification model 310z) may be the same as the corresponding portion of the pre-trained first classification model (e.g., pre-trained first classification model 310).

[0106] Furthermore, the trained first classification model is trained to extract different features from multiple classes that are not easily distinguishable using images. Therefore, if the entire modified first classification model is trained, the accuracy of the classification by the modified first classification model can be improved compared to if only a part of the modified first classification model is the same as the corresponding part of the original trained first classification model.

[0107] (4) The second type of training image used to calculate the training features (e.g., feature F2) for the second classification model (Figure 5: S120) may be different from the first type of training image (S115) used to train the first classification model. In this case as well, the classification result obtained using the second classification model, as shown in Figure 9, may indicate a low-precision group containing multiple classes that are not easily classified by the first classification model.

[0108] (5) Multiple test images used to acquire low-precision groups having low-precision relationships (Figure 5: S135, S140) may be training images included in the first type of training images (S115) for training the first classification model, or the second type of training images (S120) for calculating features for training the second classification model. In this case as well, the classification result by the second classification model of multiple features obtained using multiple test images can indicate misclassification of multiple classes, i.e., low-precision groups that include multiple classes that are not easily classified by the first classification model.

[0109] (6) Different training images may be used for training the modified Class I model than the Class I training images used for training the Class I model before the modification. For example, the dataset 330 (Figure 6) may be divided into three sets: a training set 330a, a test set 330b, and a training set for the Class I model 310z.

[0110] (7) The configuration of the first classification model is not limited to the configuration of the first classification model 310 in Figure 4, but may be various configurations including one or more convolutional layers. For example, the total number of convolutional layers in the pre-processing unit CN may be one or more of various numbers. The total number of fully connected layers in the post-processing unit FCN may be one or more of various numbers. The post-processing unit FCN may be omitted.

[0111] (8) The second classification model is not limited to support vector machines, but may be any type of predictive model that classifies multiple features into multiple classes. Preferably, the second classification model is configured to be able to misclassify multiple classes that are not easily classified by the first classification model into each other's classes. For example, the second classification model may be a predictive model that does not have convolutional layers. Also, the second classification model may be a predictive model in which the total number of computational parameters adjusted by training is smaller than that of the first classification model. The second classification model may be, for example, a predictive model called a decision tree. Both support vector machines and decision trees can be configured such that the total number of computational parameters adjusted by training is smaller than that of a convolutional neural network called VGG16.

[0112] The second classification model may be a classification model that classifies multiple samples into multiple clusters according to a clustering algorithm that automatically determines the total number of clusters. Examples of clustering methods include Density-based spatial clustering of applications with noise (DBSCAN) or Mean Shift. During training of the second classification model, multiple features are classified into multiple clusters through clustering. These clusters represent multiple classes indicating different types of images. Each cluster has a cluster region, which is the area where the multiple features belonging to the cluster are distributed. The cluster region may be, for example, the region corresponding to the convex hull containing the multiple features. Through training, the cluster regions of each of the multiple clusters are calculated. The trained second classification model may classify new features into classes (i.e., image types) that correspond to the cluster region containing those features.

[0113] The second type classes of the second classification model may be determined independently of the first type classes of the first classification model. The total number of multiple second type classes in the second classification model may differ from the total number of multiple first type classes in the first classification model. In either case, the classification results obtained using the second classification model may identify low-precision groups of multiple second type classes that have a low-precision relationship. Then, multiple first type classes corresponding to multiple second type classes included in the low-precision group may be merged. Here, the correspondence between the first type classes of the first classification model and the second type classes of the second classification model may be determined as follows: a first type class in which an image is classified by the first classification model may be associated with a second type class in which the same image is classified by the second classification model 320.

[0114] (9) The training process for the first classification model and the second classification model is not limited to the process shown in Figure 5, but may be any various process suitable for the first classification model and the second classification model. For example, the acquisition of training features (S120) may be performed by the user or by a program other than the first program 231. The acquisition of evaluation features (S130, S135) may be performed by the user or by a program other than the first program 231. Integration (S145) may be performed by the user or by a program other than the first program 231. The training process for the first classification model is not limited to the process shown in Figure 7, but may be any various method suitable for the first classification model. The training process for the second classification model is not limited to the process shown in Figure 8, but may be any various process suitable for the second classification model.

[0115] (10) Anomaly classes indicating anomalies are not limited to scratches, stains, chips, and peeling (Figures 3(D)-3(G)), but may indicate a variety of anomalies. For example, multiple anomaly classes may include an anomaly class indicating that the color of an element (string, mark, etc.) is different from the intended color.

[0116] The objects classified into multiple classes, including normal and abnormal classes, are not limited to label sheets 800, but can be various objects. The objects can be various objects provided on a product, such as three-dimensional markings or painted patterns. The product is not limited to the multifunction printer 900, but can be any product, such as a sewing machine, cutting machine, machine tool, or mobile terminal.

[0117] In either case, multiple Type I classes in the first classification model may include one normal class and multiple abnormal classes.

[0118] (11) Multiple Class I classes of the Class I model may represent image types that are unrelated to normal or abnormal. The Class I model may be used for various processes that classify image types, not limited to inspection processes. For example, the Class I model may be configured to classify scanned images of ink dots formed on paper into one of multiple Class I classes. Here, multiple Class I classes may represent different types of ink (e.g., ink manufacturer, ink manufacturing date, etc.). It may not be easy for the Class I model to distinguish between multiple scanned images that show specific manufacturers. It may also not be easy for the Class I model to distinguish between multiple scanned images that show similar manufacturing dates. Thus, multiple Class I classes that are not easily distinguishable by the Class I model may be merged according to the classification results of the Class II model.

[0119] Furthermore, the first classification model may be configured to classify captured images of animals into one of several Class I classes. Here, the multiple Class I classes may represent multiple types of animals, such as dogs, cats, jaguars, and leopards. Alternatively, the multiple Class I classes may represent multiple specific breeds of animals, such as golden retrievers, labrador retrievers, and chihuahuas. Distinguishing between a jaguar image and a leopard image using the first classification model may not be easy. Similarly, distinguishing between a golden retriever image and a labrador retriever image using the first classification model may not be easy. In such cases, multiple Class I classes that are not easily distinguishable by the first classification model may be merged according to the classification results of the second classification model.

[0120] (12) The images classified by the first classification model are not limited to images taken by the digital camera 110 (Figure 1), but may be various types of images. For example, images read by a scanner may be processed. Thus, read images representing objects optically read by reading devices such as digital cameras and scanners may be classified by the first classification model. Alternatively, images generated using a computer may be classified by the first classification model.

[0121] (13) The configuration of the device for training the classification model is not limited to the configuration of the data processing device 200 in Figure 1, but may be various configurations. For example, the GPU 260 may be omitted. Also, the data processing device may be a different type of device from a personal computer (e.g., a digital camera, scanner, smartphone). Furthermore, multiple devices (e.g., computers) that can communicate with each other via a network may each share a portion of the data processing function of the data processing device and provide the data processing function as a whole (a system equipped with these devices corresponds to the data processing device).

[0122] In the above embodiments and modifications, some of the configurations implemented by hardware may be replaced with software, and conversely, some or all of the configurations implemented by software may be replaced with hardware. For example, the processing by the first classification model may be performed by a dedicated hardware circuit such as an Application Specific Integrated Circuit (ASIC) instead of a program module. The processing by the second classification model may be performed by a dedicated hardware circuit such as an ASIC instead of a program module.

[0123] Furthermore, if some or all of the functions of this disclosure are implemented by a computer program, that program may be provided in the form of a computer-readable recording medium (e.g., a non-temporary recording medium). The program may be used while stored on the same or a different recording medium (computer-readable recording medium) as it was provided. "Computer-readable recording medium" is not limited to portable recording media such as memory cards and CD-ROMs, but may also include internal storage devices within a computer, such as various ROMs, and external storage devices connected to a computer, such as hard disk drives.

[0124] The above embodiments and modifications can be combined as appropriate. Furthermore, the above embodiments and modifications are provided to facilitate understanding of this disclosure and do not limit the present invention. The present invention can be modified and improved without departing from its spirit, and equivalents thereof are included. [Explanation of symbols]

[0125] 110…Digital camera, 200…Data processing device, 210…Processor, 215…Storage device, 220…Volatile storage device, 230…Non-volatile storage device, 231…First program, 232…Second program, 240…Display unit, 250…Operation unit, 260…Graphics processing unit (GPU), 270…Communication interface, 310, 310z…First classification model, 320…Second classification Model, 330, 330z…Dataset, 330a, 330az…Training set, 330b, 330bz…Test set, 710…Scratch, 720…Dirt, 730…Chipped, 735…Text string, 740…Peeling, 800…Label sheet, 900…Multifunction printer, 901…First side, CN…Pre-processing stage, FCN, FCNz…Post-processing stage, fc1-fc3, fc3z…Fully connected layer, Dx…First direction, Dy…Second direction

Claims

1. It is a program, A function to train a first classification model, which is a prediction model including one or more convolutional layers, to classify each of the multiple first type training images into a first type class that corresponds to the corresponding label information among multiple first type classes that represent different types of images, using multiple first type training images and multiple label information that indicates the type of each of the multiple first type training images. A function to train a second classification model, which is a predictive model different from the first classification model, to classify multiple features into multiple second-type classes that represent different types of images, using multiple features generated by a trained first classification model using multiple second-type learning images, A function that uses a trained second classification model to classify multiple feature quantities generated using multiple test images by a trained first classification model into the multiple second classes, A function for training a first classification model modified to have a reduced total number of classes by integrating multiple first-order classes, wherein the multiple first-order classes to be integrated are first-order classes associated with a low-precision group, which is a group of multiple second-order classes having a predetermined low-precision relationship, the predetermined low-precision relationship is a relationship indicated by the classification results of the multiple second-order classes associated with the multiple test images, and the relationship indicates that the classification accuracy of one or more second-order classes in the low-precision group does not meet a criterion, and the function, A program that enables a computer to realize something.

2. The program according to claim 1, The plurality of Type 2 learning images are included in the plurality of Type 1 learning images. program.

3. A program according to claim 1 or 2, Each of the aforementioned multiple test images is an image that is not included in either the aforementioned multiple Type 1 learning images or the aforementioned multiple Type 2 learning images. program.

4. A program according to claim 1 or 2, The total number of the plurality of Type 2 classes in the second classification model is the same as the total number of the plurality of Type 1 classes in the first classification model. program.

5. A program according to claim 1 or 2, The first classification model includes two or more fully connected layers that form the final part of the first classification model, Each of the aforementioned feature quantities is a feature quantity output from a specific fully connected layer that precedes the last fully connected layer among the two or more fully connected layers. program.

6. L according to claim 5, The aforementioned specific fully connected layer is the second to last fully connected layer among the two or more fully connected layers. program.

7. A program according to claim 1 or 2, The second classification model is a support vector machine. program.

8. A program according to claim 1 or 2, The predetermined low-precision relationship is a relationship that indicates that the classification precision of two Type 2 classes forming the low-precision group does not meet the criteria, and the fact that the classification precision of the two Type 2 classes does not meet the criteria indicates that the correct classification rate of each of the two Type 2 classes is less than or equal to the first value, and the misclassification rate of each of the two Type 2 classes to each other's classes is greater than or equal to the second value. program.

9. A program according to claim 1 or 2, The predetermined low-accuracy relationship is a relationship that indicates that the classification accuracy of the two Type 2 classes forming the low-accuracy group does not meet the criterion, and the fact that the classification accuracy of the two Type 2 classes does not meet the criterion indicates that the misclassification rate of the two Type 2 classes to each other's classes is 3 or higher. program.

10. A program according to claim 1 or 2, The function for training the modified first classification model trains the entire modified first classification model. program.

11. A program according to claim 1 or 2, The plurality of first classes include a normal class and a plurality of abnormal classes, the normal class indicates a type of image representing a normal object, each of the plurality of abnormal classes indicates a type of image representing an object having an abnormality, and the plurality of abnormal classes are associated with different types of abnormalities. The Type II class included in the aforementioned low-precision group is a Type II class that corresponds to the abnormal class. program.

12. A program according to claim 1 or 2, The function for training the modified first classification model trains the modified first classification model using at least a portion of the plurality of first type training images. program.

13. It is a training method, A step of training a first classification model, which is a prediction model including one or more convolutional layers, so as to classify each of the multiple first type training images into a first type class that corresponds to the corresponding label information among a plurality of first type classes that represent different types of images, using a plurality of first type training images and a plurality of label information that indicates the type of each of the plurality of first type training images. A step of training a second classification model, which is a predictive model different from the first classification model, to classify multiple features into multiple second classes that represent different types of images, using multiple features generated by a trained first classification model using multiple second-type learning images, The process involves classifying multiple features generated using multiple test images by a pre-trained first classification model into the multiple second classes using a pre-trained second classification model, A step of training a first classification model modified to have a total number of classes reduced by the integration of multiple first-order classes, wherein the multiple first-order classes to be integrated are first-order classes associated with a low-precision group which is a group of multiple second-order classes having a predetermined low-precision relationship, the predetermined low-precision relationship is a relationship indicated by the classification results of the multiple second-order classes associated with the multiple test images, and the relationship indicates that the classification accuracy of one or more second-order classes among the low-precision group does not meet a criterion, and the step of training a first classification model modified to have a total number of classes reduced by the integration of multiple first-order classes, wherein the multiple first-order classes to be integrated are first-order classes associated with a low-precision group which is a group of multiple second-order classes having a predetermined low-precision relationship the predetermined low-precision relationship is a relationship indicated by the classification results of the multiple second-order classes associated with the multiple test images, and the relationship indicates that the classification accuracy of one or more second-order classes among the low-precision group does not meet a criterion, and the step of training a first classification model modified to have a total number of classes reduced by the integration of multiple first-order classes, wherein the multiple first-order classes to be integrated are first-order classes associated with a low-precision group which is a group of multiple second-order classes having a predetermined low-precision relationship, the predetermined low-precision relationship is a relationship indicated by the classification results of the multiple second-order classes associated with the multiple test images which is a relationship which indicates that the classification accuracy of one or more second-order classes among the low-precision group is not met, and the step of training a first classification model modified to have a total number of classes reduced by the integration of multiple first-order classes, wherein the multiple first-order classes to be integrated are A training method that includes [the following].

Citation Information

Patent Citations

  • Feature quantity selecting method, feature quantity selecting program, feature quantity selecting device, multiclass classification method, multiclass classification program, multiclass classification device, and feature quantity set

    WO2022065216A1