Deep learning model generation device, and deep learning model generation method
The deep learning model generation device addresses the high workload of creating marking images by using evaluation indices and data augmentation, enabling efficient and accurate deep learning model generation for defect inspection.
Patent Information
- Application Number
- JP2025071850
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-03
AI Technical Summary
The high workload associated with preparing large numbers of marking images for deep learning models, particularly in the context of defect inspection, remains a significant challenge, as existing methods like crowdsourcing, automatic marking, and artificial generation require pre-trained models, which do not significantly reduce the labor involved.
A deep learning model generation device that recursively modifies a deep learning model by using evaluation indices to evaluate differences between per-pixel and per-image labels, incorporating pseudo-pixel labels and data augmentation techniques like CutMix, reducing the need for extensive manual labeling.
This approach significantly reduces the workload while maintaining high accuracy by leveraging pseudo-pixel labels and data augmentation, allowing for efficient generation of deep learning models for defect inspection.
Smart Images

Figure 2025100887000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a deep learning model generation device and a deep learning model generation method.
Background Art
[0002] For example, a surface property inspection device for inspecting the surface properties of an inspection object such as a steel plate is an inspection device that estimates the type and shape of defects using a deep learning (deep learning, DL) model for an image obtained by photographing the surface of the inspection object and outputs them as output values. The pre-trained deep learning model used in the above inspection device is a model that is trained in advance by preparing a large number of defective images (marking images) with labels assigned in pixel units so that the error between the estimated result when each defective image is input and the correct marking image is reduced.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] As described above, in order to construct the deep learning model, it is necessary to prepare a large number of marking images. However, preparing a large number of marking images has a very high workload problem. As a specific work image for creating a marking image, it is an operation of projecting a defective image on a tablet and labeling (marking) it in pixel units for each type of defect. However, it takes several minutes to finish marking one defective image, and marking hundreds or thousands of defective images required for constructing a deep learning model is an impossible task. The problem is how to create marking images in a labor-saving manner and construct a DL model.
[0005] Regarding a method of collecting such marking images as efficiently as possible to construct a deep learning model, the following are cited as known techniques / prior arts. (1) Marking using crowdsourcing This is a method of paying an external company or organization to commission the creation of marking images. However, while there is no problem with general images (e.g., animals and buildings) as they have no particular confidentiality, it cannot be used for highly confidential images. (2) Automatic marker (Patent Document 1) This is a method in which unmarked data is inferred using a pre-trained model, data with confidence is directly marked with the output result of the model and added to the training data, and data without confidence is manually marked and added to the training data. With this method, only data without confidence is marked, so a reduction in the workload can be expected. However, since the pre-trained model as the starting point needs to be prepared in advance, it is necessary to create a certain number of marked data in advance anyway, and it can still be said that the workload is high.
[0006] (3) Artificial generation of marking images using an image generation model (Patent Document 2) This is a method of artificially augmenting marking images using a known image generation model (adversarial generative network, GAN). An image generation model is a model that outputs an artificial marking image labeled in pixel units through the operation of a pre-trained model based on any input random number. Similar to (2) the automatic marker, since the pre-trained model as the starting point needs to be prepared in advance, it can still be said that the workload is high. As can be seen from (1), (2), and (3) above, a large number of defect marking images are required to construct a defect estimation model in pixel units, but no method has been proposed at present that can significantly reduce the workload.
[0007] Therefore, an object of the present invention is to provide a deep learning model generation device and a deep learning model generation method that reduce the workload and enable the generation of a highly accurate deep learning model.
Means for Solving the Problems
[0008] [1] A deep learning model generation device that generates a deep learning model by recursively modifying a deep learning model that estimates a per-pixel label, which is a label for each pixel included in the image, when an image is input. A first evaluation index is an evaluation index that evaluates that the difference between the per-pixel label given to the per-pixel label-assigned image, which is an image with the per-pixel label, and the estimation result has become small. A second evaluation index is an evaluation index that evaluates that the difference between the pseudo per-pixel label obtained by replacing the label of the estimation result with the per-image label given to the per-image label-assigned image, which is an image with the per-image label, and the estimation result has become small. When a third evaluation index is defined as an evaluation index that evaluates that the difference calculated from the estimation result has become small for at least the per-pixel label-assigned image among the per-pixel label-assigned image, the per-image label-assigned image, and the image without a label, and the label information of the per-pixel label given to the per-pixel label-assigned image or the per-image label given to the per-image label-assigned image is not essential, a deep learning model is generated based on an evaluation index obtained by adding at least one of the second evaluation index or the third evaluation index to the first evaluation index. Deep learning model generation device. [2] The pseudo per-pixel label is the result of weighted averaging the estimation results obtained by inputting the estimation results obtained by inputting the deep learning model corrected up to the previous time, the estimation results obtained by inputting the current deep learning model, and the estimation results obtained by inputting the deep learning model including the time before the previous time, or the result obtained by inputting the deep learning model obtained by weighted averaging the deep learning models including the time before the previous time. The deep learning model generation device according to [1], wherein the label of the estimation result is replaced with the per-image label given to the per-image label-assigned image. [3] When the pixel ratio is defined as the ratio of the number of pixels of the label in the image to the number of pixels of the image, and the image ratio is defined as the ratio of the number of images per label to the number of images, the third evaluation index is the difference between the distribution of the pixel ratio calculated from the estimation result and the target distribution of the pixel ratio determined in advance, the difference between the distribution of the image ratio calculated from the estimation result and the target distribution of the image ratio determined in advance, or the difference related to CutMix, which is a data augmentation method of the deep learning model, at least one of which is added to the deep learning model generation device according to [1]. [4] The target distribution of the pixel ratio and the target distribution of the image ratio are set based on at least one of the labels assigned to the per-pixel labeled image or the labels assigned to the per-image labeled image, for the deep learning model generation device according to [3]. [5] The image is an image obtained for steel products during and after manufacturing, for the deep learning model generation device according to any one of [1] to [4].
[0009] [6] A deep learning model generation method that generates a deep learning model by recursively modifying a deep learning model that estimates a per-pixel label, which is a label for each pixel included in the image, when an image is input. A first evaluation metric is an evaluation metric that evaluates that the difference between the per-pixel label assigned to the per-pixel label assigned image, which is an image with the per-pixel label, and the estimation result has become smaller. A second evaluation metric is an evaluation metric that evaluates that the difference between the pseudo per-pixel label obtained by replacing the label of the estimation result with the per-image label assigned to the per-image label assigned image, which is an image with the per-image label, and the estimation result has become smaller. A third evaluation metric is defined as an evaluation metric that evaluates that the difference calculated from the estimation result has become smaller, targeting at least the per-pixel label assigned image among the per-pixel label assigned image, the per-image label assigned image, and the image without a label, where the label information of the per-pixel label assigned to the per-pixel label assigned image and the per-image label assigned to the per-image label assigned image is not essential. When generating a deep learning model based on an evaluation metric obtained by adding at least one of the second evaluation metric or the third evaluation metric to the first evaluation metric.
Advantages of the Invention
[0010] According to the present invention, it is possible to reduce the workload and generate a high-precision deep learning model.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Embodiments for Carrying Out the Invention
[0012] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. The present invention relates to so-called deep learning technology (deep learning), and relates to an apparatus (deep learning model generation apparatus) for generating a deep learning model in which learning is performed using deep learning technology. Deep learning is based on a neural network having deep layers and is a technique that has developed machine learning. A neural network is composed of an input layer, an intermediate layer (hidden layer), and an output layer, and a model in which the intermediate layer consists of multiple layers is a deep neural network. Deep learning performs learning using this deep neural network and can automatically extract features and learn using a large amount of data.
[0013] In the present invention, deep learning technology is particularly used for image recognition. In addition to using marking data (referred to as pixel-by-pixel labeled image in this specification) in which a label (information indicating the correct answer) is assigned to each pixel as teacher data of an image, as weak teacher data, labeling data in which a label is assigned to each image and can be created labor-savingly (referred to as image-by-image labeled image in this specification) is used. It is also characterized in that regularization is performed using data including data without a label in order to suppress overfitting. Note that the number (type) of labels per pixel and the number (type) of labels per image are the same.
[0014] First, as an example of an apparatus including the deep learning model generated by the deep learning model generation apparatus, a surface property inspection apparatus will be described. <Surface Defect Inspection Device> The surface defect inspection device is a device for inspecting the surface properties of an object to be inspected, such as steel products. When the surface defect inspection device inputs a surface image of the object to be inspected (hereinafter referred to as the input image), it estimates the type and shape (position) of the defect using a deep learning model and outputs the estimation result. The input image is an image taken by a digital camera or the like and is composed of, for example, 800×600 pixels. Specifically, it includes not only images taken by an optical camera of steel products during or after manufacturing, but also images obtained by a device that measures or estimates various physical quantities such as the surface temperature distribution. In addition, it includes images obtained by an internal flaw detection device using ultrasonic waves or magnetism to confirm that there are no internal defects. Furthermore, it includes images obtained by an appearance inspection device when shipping steel products such as products wound in a coil shape. Defects include, for example, scratches, stains, foreign objects, local temperature drops, and poor coil winding shapes.
[0015] The surface defect inspection device estimates the defects included in the input image from multiple types of defects (for example, three types of defects, Defect A to Defect C), and estimates and outputs the pixels where the defects exist. The output image corresponds to the types of defects that can be output. When corresponding to three types of defects, the presence or absence and location of each defect are shown by three images. Note that the number of types of defects is not limited to three, and at least one type is sufficient and can be set arbitrarily.
[0016] Figure 1 is a conceptual diagram for explaining the processing by the surface defect inspection device 1. When an input image I having Defect B is input to the surface defect inspection device 1, an output image OD corresponding to Defects A to C is output by the deep learning model M1. When the input image I includes Defect B, the deep learning model M1 of the surface defect inspection device 1 shows the shape of Defect B by making the pixels where Defect B exists white pixels in the output image DB corresponding to Defect B among the output images OD, and shows the pixels that are the base (texture) where no defect exists as black pixels. Among the output images OD in Figure 1, the output images DA and DC indicate that there are no such types of defects in the input image. Note that the output method is not limited to this, and an output image may be obtained by, for example, indicating the location estimated to be defective with a rectangular frame and indicating the type of defect with a symbol.
[0017] Next, a deep learning model generation device that generates the above-described deep learning model will be described. <Deep learning model generation device> The deep learning model generation device 100 is a device that generates a deep learning model based on a plurality of evaluation metrics. As shown in FIG. 2, the deep learning model generation device 100 includes a deep learning model initial creation unit 10 that creates an initial deep learning model M0 (see FIG. 6) necessary for generating the deep learning model, a deep learning model modification unit 20 that modifies the initial deep learning model M0 to obtain a deep learning model M1 (see FIG. 1) that can be used in a surface property inspection device or the like, and a storage unit 30 that stores teacher data and the like. The deep learning model M1 is generated by modifying the initial deep learning model M0 created by the deep learning model initial creation unit 10 by the deep learning model modification unit 20. Note that the deep learning model may be generated by a single deep learning model generation unit without modifying the deep learning model initial creation unit 10 by the deep learning model modification unit 20.
[0018] Here, terms and symbols used in this specification will be described. 〔Image and pixel〕 Let the total number of images I used in deep learning be I, and each image is represented by I i (i ∈ 1, 2, …, i, …, I). Let the maximum number of pixels in the vertical and horizontal directions of the image be J and K, respectively. Each pixel of the image I i is represented by I i (j,k) (j, k ∈ J, K).
[0019] 〔Label for each pixel〕 A per-pixel label is a label for each pixel of an image. If a deep learning model is used in the surface property inspection apparatus described above, the per-pixel label is the type of defect for each pixel. There may be cases where the performance of the per-pixel label is assigned to each of the multiple pixels in one image, and it is estimated by the deep learning model.
[0020] 〔Per-Pixel Label-Assigned Image〕 Labeling is performed by an operator. In addition to the operator identifying the type of defect for each image, the operator assigns the performance of the label for each pixel. FIG. 3 is a diagram showing an image to which per-pixel labels are assigned. In the figure, the black part is the pixel identified as having no defect, and the white part is the pixel identified as having defect A. In this specification, an image including information on the type of defect and to which per-pixel labels are assigned is called a per-pixel label-assigned image. Usually, when generating a deep learning model, a large number of per-pixel label-assigned images are prepared as teacher data, but the burden of the assignment work is large. The index set of the per-pixel label-assigned image is represented by I pix as shown. The per-pixel label assigned to the per-pixel label-assigned image I i is represented by T i (j,k) (0 ≦ T i (j,k) ≦ T). Note that the index: i is the same as the index of the image I i as shown.
[0021] 〔Per-Image Label〕 A per-image label is a label for each image. That is, there may be cases where the performance of one per-image label is assigned to one image. 〔Per-Image Label-Assigned Image〕 Labeling is performed by an operator. The operator identifies the type of defect for each image and assigns the performance of the label for each image. In this specification, an image to which per-image labels are assigned is called a per-image label-assigned image. Since only one label needs to be assigned to each image, the burden of the assignment work is small. The index set of the per-image label-assigned image is represented by I im as shown. The per-image label-assigned image Ii The per-pixel label assigned to each image is t i (0 ≤ t i ≤ T). Note that index: i is the same as the index of image I i .
[0022] [Images without labels] Among all the images, the index set of the images without labels, which are other than the per-pixel label-assigned images and the per-image label-assigned images, is I non .
[0023] [Deep learning model initial creation unit] The deep learning model initial creation unit 10 creates an initial deep learning model M0, which serves as the basis for constructing the final deep learning model M1, using a plurality of per-pixel label-assigned images MD (teacher data) stored in the storage unit 30. Specifically, using a known deep learning model creation method such as CNN (Convolutional Neural Network), for the per-pixel label-assigned image MD, an initial deep learning model M0 is created based on a first evaluation index (described later) that evaluates that the difference between the defect estimation result and the per-pixel label-assigned image MD has become small.
[0024] [Deep learning model modification unit] The deep learning model modification unit 20 modifies the initial deep learning model M0 created using the deep learning model initial creation unit 10 to complete the final deep learning model M1. Similar to the deep learning model initial creation unit 10, the deep learning model modification unit 20 modifies the initial deep learning model M0 based on the first evaluation index that evaluates that the difference between the defect estimation result and the per-pixel label-assigned image MD has become small for the per-pixel label-assigned image MD, and also modifies the initial deep learning model M0 based on a second evaluation index and a third evaluation index described later. That is, the deep learning model modification unit 20 adds information from the per-image label-assigned image LD, which has a low creation cost and can be prepared in large quantities, in addition to the per-pixel label-assigned image MD, and modifies the initial deep learning model M0 to suppress overfitting.
[0025] As described above, the configuration including the deep learning model initial creation unit 10 and the deep learning model modification unit 20 is not essential, and the deep learning model generation device 100 may have a configuration including a single deep learning model generation unit.
[0026] The plurality of evaluation metrics used to generate the deep learning model are a first evaluation metric for calculating a supervised learning term, a second evaluation metric for calculating a weak supervised learning term, and a third evaluation metric for calculating a regularization term. The deep learning model generation device 100 performs learning so that the sum of each term (supervised learning term + weak supervised learning term + regularization term) becomes small. Table 1 is a table summarizing the supervised learning term represented by a mathematical formula for the first evaluation metric, the weak supervised learning term represented by a mathematical formula for the second evaluation metric, and the regularization term represented by a mathematical formula for the third evaluation metric. The mathematical formulas for each term will be described later. R1 is a mathematical formula related to the pixel ratio (first regularization structure) and the image ratio (second regularization structure) among the regularization terms, and is a mathematical formula obtained by adding the following formulas (18) and (19). R2 represents a mathematical formula related to CutMix (third regularization structure) among the regularization terms. Also, α, β1, and β2 are weighting coefficients for each term (α ≥ 0, β1 ≥ 0, β2 ≥ 0). As shown in Table 1, the supervised learning term is essential, but at least one of the weak supervised learning term and the regularization term may be added. That is, α = 0 and β1 = 0 can be set, α = 0 and β2 = 0 can be set, β1 = 0 can be set, β2 = 0 can be set, β1 = 0 and β2 = 0, etc. can be set.
[0027]
Table 1
[0028] Hereinafter, the first evaluation metric, the second evaluation metric, and the third evaluation metric used to determine the accuracy of the deep learning model M1 generated by the deep learning model generation device 100 will be described.
[0029] <First Evaluation Metric> The first evaluation index is an index for evaluating that the difference between the teacher data (pixel-by-pixel labeled image) and the estimation result has become small by performing supervised learning using the teacher data. Specifically, the first evaluation index evaluates that the difference between the pixel-by-pixel label given to the pixel-by-pixel labeled image MD and the estimation result has become small for the pixel-by-pixel labeled image MD. The deep learning model initial creation unit 10 calculates a supervised learning term according to the first evaluation index. The supervised learning term calculated based on the first evaluation index can be expressed as the following formula (1).
[0030]
Equation
[0031] L(x, y) is a term for calculating the difference between the actual result x of the label and the estimation result y. The mathematical formula of the supervised learning term may be any formula in which the value of L(x, y) monotonically increases as the difference increases. For example, CategoricalCrossentropy mainly used in multi-class classification, BinaryCrossentropy mainly used in binary classification, and MeanSquaredError generally used can be used. Here, F(I i ) is the estimation result obtained by inputting the image I i into the deep learning model, and F(I i ) (j,k) is the estimation result of the pixel-by-pixel label. That is, the supervised learning term is the sum of the differences between the actual result T i (j,k) of the pixel-by-pixel label and the estimation result F(I i ) (j,k) of the pixel-by-pixel label for all pixels, and further integrated for all pixel-by-pixel labeled images (I pix ).
[0032] <The second evaluation index> The second evaluation index uses the "pseudo pixel-wise label", which is the result of the pixel-wise label created pseudo-using weak teacher data (image-by-image labeled images), and evaluates that the difference between the pseudo pixel-wise label and the pixel-wise label of the estimation result has become smaller. Specifically, the second evaluation index evaluates that the difference between the pseudo pixel-wise label and the estimation result has become smaller for the image-by-image labeled images.
[0033] The pseudo pixel-wise label is obtained by replacing the label of the estimation result by the deep learning model with the image-by-image label given to the image-by-image labeled image. Figure 4 is a diagram for explaining the method of obtaining the pseudo pixel-wise label. Figure 4(a) is the estimation result by the deep learning model, and Figure 4(b) is the image-by-image label t i to which information on the defect of the estimation result shown in Figure 4(a) is added, which is the pseudo pixel-wise label. As shown in Figure 4, the pseudo pixel-wise label is created by leaving the information indicating no defect, which is indicated by 0, as it is, and changing the label corresponding to the defect indicated by 1 to 3 to the image-by-image label given to the image-by-image labeled image. Although the pixel-wise label and the image-by-image label are explained as integers, they may be symbols instead. In that case, if the symbols are replaced with integers separately, they can be processed as integers.
[0034] The deep learning model correction unit 20 calculates a weak teacher learning term according to the second evaluation index, and corrects the initial deep learning model M0 created using the deep learning model initial creation unit 10. Similar to the first evaluation index, using the term L(x,y) that calculates the difference between the actual result x of the label and the estimation result y, the weak teacher learning term calculated based on the second evaluation index can be expressed as in the following formula (2).
[0035]
Equation
[0036] The pseudo pixel-wise label is represented by T’ i (j,k) as shown. Here, the image Ii The estimated result obtained by inputting it into the deep learning model M of this time m is denoted as F(I i ), and the estimated result obtained by inputting the image I i into the deep learning model M m-1 modified up to the previous time is denoted as F -1 (I i ). Also, as shown in formulas (3) to (5), M EMA1 is the deep learning model obtained by weighted averaging (exponential averaging) the deep learning model M m (weights and coefficients) of this time and the deep learning model M m-1 of the previous time. When M EMA2 is the deep learning model obtained by gradually weighted averaging the current deep learning model M m and the previous value of M EMA2 , and M EMA3 is the deep learning model obtained by gradually weighted averaging the previous deep learning model M m-1 and the previous value of M EMA3 , the estimated results obtained by inputting the image I i into M EMA1 , M EMA2 , and M EMA3 are denoted as F EMA1 (I i ), F EMA2 (I i ), and F EMA3 (I i ), respectively. Here, η is the weighted average coefficient (0 ≤ η ≤ 1). The initial values of M EMA2 and M EMA3 and M -1 may be set to zero.
[0037]
Equation
[0038] The function of the label for each pseudo-pixel, that is, the estimated result F(I i ) (j,k) regarding the defect is converted into the label t i for each image, and the function can be any one of the following formulas (6) to (10).
[0039]
Number
[0040] Also, using the estimation results F(I i ) or F -1 (I i ), as shown in formulas (11) to (13), instead of using the exponentially weighted moving averages EMA1(I i ), EMA2(I i ), and EMA3(I i ) calculated by weighted averaging, any one of the following formulas (14) to (16) may be used instead of formulas (6) to (10).
[0041]
Number
[0042] <Third Evaluation Index> The third evaluation index targets at least the per-pixel labeled image among the per-pixel labeled image, the per-image labeled image, and the unlabeled image. It is not essential to have the label information of the per-pixel label assigned to the per-pixel labeled image or the per-image label assigned to the per-image labeled image, and it is an index for evaluating that the difference calculated from the estimation result has become small. Since the third evaluation index does not require the label assigned to the image, the unlabeled image can also be targeted.
[0043] The deep learning model modification unit 20 calculates a regularization term according to the third evaluation index and modifies the initial deep learning model M0 created using the deep learning model initial creation unit 10. The regularization term calculated based on the third evaluation index can be expressed as the following formula (17).
[0044]
Number
[0045] The third evaluation index evaluates that the appropriate regularization structure is such that deep learning does not progress in an inappropriate direction, and the difference calculated by the first regularization structure, the second regularization structure, and the third regularization structure described below is reduced. Hereinafter, each regularization structure will be described.
[0046] <The first regularization structure> The first regularization structure is a regularization structure that utilizes information on the pixel ratio of defects. The pixel ratio of defects is the ratio of the number of pixels constituting the defect to the number of pixels in the image (480,000 for an 800×600 image). The first regularization structure performs regularization based on this pixel ratio of defects. Specifically, the difference between the distribution of the pixel ratio of defects calculated from the estimation results for a plurality of images and the preset target distribution (target distribution) of the pixel ratio of defects is added. The distribution of the pixel ratio of defects is the probability distribution of the pixel ratio calculated for each type of defect.
[0047] The third evaluation index related to the first regularization structure can be expressed as in the following formula (18).
[0048]
Equation
[0049] l(X,Y) is a function that calculates the difference between the probability distributions of distribution X and distribution Y, and may be calculated using the mean squared error or may be calculated using the Kullback-Leibler distance or the like. U i (F(I i )) is a function that inputs the image I i and calculates the distribution of the pixel ratio of defects.
[0050] The target distribution U aim of the pixel ratio of defects is set based on at least one of the labels assigned to the per-pixel labeled image and the labels assigned to the per-image labeled image. When targeting the per-pixel labeled image and the per-image labeled image, the distribution calculated based on the assigned label may be used as the target distribution of the defective pixel ratio. (2) When targeting also images without an assigned label (that is, when targeting per-pixel labeled images, per-image labeled images, and other images), images without an assigned label may be regarded as having the same distribution as the target distribution obtained in (1) above and set accordingly. Alternatively, for example, if there is prior information regarding each target distribution in surface property inspection work by human visual inspection, etc., it may be set using that prior information. For example, if the target distribution of the pixel ratios for defects A, B, and C is U aim then U aim = (0.1, 0.2, 0.3) can be set. β 11 is a weight function (β 11 ≥ 0). That is, the regularization term related to the first regularization structure calculates the difference between the distribution of the target distribution U aim of the defective pixel ratio and the distribution of the defective pixel ratio estimated by inputting the image Ii. When it is particularly desired to improve the accuracy of a predetermined defect, the weight (importance) may be set so that the difference regarding the predetermined defect becomes large.
[0051] <Second Regularization Structure> The second regularization structure is a regularization structure that utilizes information on the defective image ratio. The defective image ratio is the ratio of the number of defective images to the total number of images used in the second regularization structure (for example, when targeting I pix , I im , I non , it is the ratio of the number of defective images to I pix + I im + I non ). The second regularization structure performs regularization based on this defective image ratio. Specifically, the difference between the distribution of the defective image ratio calculated from the estimation result and the preset target distribution of the defective image ratio is added. The third evaluation index related to the second regularization structure can be expressed as in the following formula (19).
[0052]
Number
[0053] V i (F(I i )) is a function that takes the image I i as input and calculates the distribution of the defective image ratio.
[0054] The target distribution V aim of the defective image ratio is set based on at least one or more of the labels assigned to the per-pixel labeled image and the labels assigned to the per-image labeled image. For example, if the target distribution of the image ratios for defects A, B, and C is V aim , then V aim = (0.1, 0.01, 0.05) can be set. β 12 is a weight function (β 12 ≥ 0). That is, the regularization term related to the second regularization structure calculates the difference between the target distribution V aim of the defective image ratio and the distribution of the defective image ratio estimated by inputting the image Ii.
[0055] <The Third Regularization Structure> The third regularization structure is a regularization structure that utilizes the well-known technique "CutMix" in the data augmentation of images in a deep learning model. CutMix is a technique that creates new training data by combining two images and uses this new training data to improve image recognition accuracy. CutMix is described in the following literature. et al. S. Yun. Cutmix: Regularization strategy to train strong classifiers with localizable features. In arXiv:1905.04899.
[0056] The method of CutMix will be described with reference to FIG. 5. In CutMix, image processing (hereinafter referred to as composite image processing) of pasting a second image D2 (mixed-in image) onto a first image D1 (original image) is utilized.
[0057] First, two input images (the first image D1 and the second image D2) are prepared. The first image D1 is a pixel-by-pixel labeled image MD, an image-by-image labeled image LD, or an image without a label. The second image D2 is a randomly determined image, and like the original image, it is a pixel-by-pixel labeled image MD, an image-by-image labeled image LD, or an image without a label.
[0058] Next, the two input images D1 and D2 are synthesized using composite image processing (flow F2). Specifically, a predetermined rectangular region S of the second image D2 is cut out, and this region S is pasted at the same position of the first image D1. Thereby, a composite input image DM is generated as new training data. Note that the rectangular region S of the composite image processing may have a shape other than a rectangle, and the size is not limited either. Next, for each of the three images D1, D2, and DM, estimation is performed using a deep learning model to obtain an estimation result RD1 of the first image D1, an estimation result RD2 of the second image D2, and an estimation result RDM of the composite input image DM. On the other hand, the estimation result RD1 of the first image D1 and the estimation result RD2 of the second image D2 are synthesized using composite image processing (flow F3) to obtain a composite estimation result RD12. The deep learning model modification unit 20 modifies the deep learning model based on an evaluation index that evaluates that the difference between the estimation result RDM of the composite input image DM and the composite estimation result RD12 has become small. Note that the number of composite images for data augmentation depends on the number of combinations of the two images to be targeted and the number of methods of composite image processing, but it is not particularly limited and any number may be used.
[0059] As described above, since the third evaluation index (all regularization term structures) does not require the performance of the labels assigned to the images, images without assigned labels can also be included in the target. That is, the images targeted in the third evaluation index are the index set I of the label-assigned images for each image im and / or the index set I of the label-assigned images for each pixel pix may be used, but the index set I of the images without assigned labels non may also be included.
[0060] <Method for Generating Deep Learning Model> Next, a method for generating the deep learning model M1 using the deep learning model generation device 100 will be described. FIG. 6 is a schematic diagram for explaining the method for generating the deep learning model. Here, a method for generating a deep learning model used in a surface property inspection device capable of discriminating three types of defects (defect A to defect C) will be described.
[0061] First, the operator uses the measuring device 40 to capture a surface image of the inspection object 41 (for example, steel products during and after manufacturing). It is preferable that the amount of images captured is as large as possible, for example, 100,000 images. Next, the operator OP creates a per-pixel label-assigned image MD with a high workload for a part of the images (for example, 10,000 images), and on the other hand, creates a per-image label-assigned image LD with a low workload for, for example, 29,000 images, and stores them in the storage unit 30 (see FIG. 2).
[0062] <Generation of Initial Deep Learning Model> Next, an initial deep learning model M0 is generated using the per-pixel label-assigned image MD created in the above process. Specifically, for 10,000 per-pixel label-assigned images MD, the initial deep learning model M0 is generated by the deep learning model initial creation unit 10 of the deep learning model generation device 100.
[0063] <Modification of Initial Deep Learning Model> Next, using the per-pixel labeled image MD, the per-image labeled image LD, and the unlabeled images (e.g., the remaining 70,000 images), the initial deep learning model M0 is modified. In modifying the initial deep learning model M0, the deep learning model modification unit 20 modifies the initial deep learning model M0 based on an evaluation index obtained by adding at least one of the second evaluation index and the third evaluation index in addition to the first evaluation index.
[0064] In modifying the initial deep learning model M0 based on the second evaluation index, as shown in FIG. 6, the per-pixel labeled image MD and the per-pseudo-pixel labeled image PMD are input into the deep learning model M0 being learned as teacher data, and the deep learning model modification unit 20 evaluates that the difference between the estimation result and the per-pixel labeled image MD and the per-pseudo-pixel labeled image PMD has become small, and recursively modifies the initial deep learning model M0 or the modified deep learning model M1 as shown in flow F1.
[0065] Similarly, in modifying the initial deep learning model M0 based on the third evaluation index, at least the per-pixel labeled image MD among the per-pixel labeled image MD, the per-pseudo-pixel labeled image PMD, and the unlabeled images is input into the deep learning model M0 being learned as teacher data / weak teacher data, and the initial deep learning model M0 or the modified deep learning model M1 is recursively modified as shown in flow F1.
[0066] According to the above embodiment, since the amount of the required per-pixel labeled images (marking data) can be reduced, the work load can be reduced and a high-precision deep learning model can be generated. Further, by adding a regularization structure, it is possible to suppress the deep learning from progressing in an inappropriate direction.
[0067] FIG. 7 is a graph for explaining the effect when the deep learning model generated by the deep learning model generation device 100 of the present invention is used in a surface property inspection device. In FIG. 7, “comparative example” is a deep learning model generated only by a supervised learning term using conventional marking data, and “Example 1” and “Example 2” are both deep learning models using the labeling data of the present invention. “Example 1” uses the supervised learning term and the weak supervised learning term as evaluation indexes, and “Example 2” is the result generated by further adding the image ratio and the regularization term of CutMix to the evaluation indexes. Note that the vertical axis of the ratio in the graph is a value in a certain range of values normalized so that a value greater than 0% is 0 and a value smaller than 100% is 1, so that the degree of good or bad can be understood.
[0068] As shown in FIG. 7, in “Example 1” in which the weak supervised learning term according to the present invention is added as an evaluation index, the detection rate of defects and the coincidence rate of defect types are greatly improved with respect to the “comparative example”, and it can be seen that the over-detection rate of erroneously detecting an image with a defect for an image without a defect (texture) can be significantly reduced. Furthermore, in “Example 2” in which the regularization term is also added to the evaluation index, it can be seen that the over-detection rate can be greatly reduced while maintaining almost the same accuracy in the detection rate and the defect type coincidence rate with respect to “Example 1”.
Explanation of symbols
[0069] 1... Surface property inspection device, 100... Deep learning model generation device, 10... Deep learning model initial creation unit, 20... Deep learning model modification unit, 30... Storage unit, 40... Measuring device, A, B, C... Defects, M0... Initial deep learning model, M1... Deep learning model, I... Input image, OD... Output image, MD... Pixel-by-pixel labeled image, LD... Image-by-image labeled image, PMD... Pseudo-pixel-by-pixel labeled image.
Claims
1. A deep learning model generation device that generates by recursively modifying a deep learning model that estimates a per-pixel label, which is a label for each pixel included in the image, when an image is input, using, as a first evaluation metric, an evaluation metric that evaluates that the difference between the per-pixel label assigned to the per-pixel label-assigned image, which is an image with the per-pixel label assigned thereto, and the estimation result has become smaller, using, as a second evaluation metric, an evaluation metric that evaluates that the difference between the pseudo per-pixel label obtained by replacing the label of the estimation result with the per-image label assigned to the per-image label-assigned image, which is an image with the per-image label assigned thereto, and the estimation result has become smaller, when a third evaluation metric is defined as an evaluation metric that evaluates that the difference calculated from the estimation result has become smaller, targeting at least the per-pixel label-assigned image among the per-pixel label-assigned image, the per-image label-assigned image, and the image without a label assigned, where the label information of the per-pixel label assigned to the per-pixel label-assigned image or the per-image label assigned to the per-image label-assigned image is not essential, A deep learning model generation device that generates a deep learning model based on an evaluation metric obtained by adding at least one of the second evaluation metric or the third evaluation metric to the first evaluation metric.
2. The deep learning model generation device according to claim 1, wherein the pseudo per-pixel label is obtained by replacing the label of the estimation result with the per-image label assigned to the per-image label-assigned image for any one of the estimation result obtained by inputting to the deep learning model corrected up to the previous time, the estimation result obtained by inputting to the current deep learning model, the result obtained by weighted-averaging the estimation results obtained by inputting to the deep learning models including before the previous time, or the result obtained by inputting to the deep learning model obtained by weighted-averaging the deep learning models including before the previous time.
3. When defining the pixel ratio as the ratio of the number of pixels of the label in the image to the number of pixels of the image, and the image ratio as the ratio of the number of images for each label to the number of images, the third evaluation metric is the difference between the distribution of the pixel ratio calculated from the estimation result and the predetermined target distribution of the pixel ratio, and the difference between the distribution of the image ratio calculated from the estimation result and the predetermined target distribution of the image ratio. The deep learning model generation device according to claim 1, wherein at least one of the differences related to CutMix, which is a data augmentation method for the deep learning model, is added.
4. The deep learning model generation device according to claim 3, wherein the target distribution of the pixel ratio and the target distribution of the image ratio are set based on at least one of the labels assigned to the per-pixel labeled image or the labels assigned to the per-image labeled image.
5. The deep learning model generation device according to any one of claims 1 to 4, wherein the image is an image obtained for steel products during and after manufacturing.
6. A deep learning model generation method for generating a deep learning model by recursively modifying a deep learning model that estimates a per-pixel label, which is a label for each pixel included in the image, when an image is input, using, as a first evaluation index, an evaluation index for evaluating that the difference between the per-pixel label assigned to the per-pixel labeled image, which is an image with the per-pixel label assigned thereto, and the estimation result has become smaller, using, as a second evaluation index, an evaluation index for evaluating that the difference between the per-pixel pseudo-label obtained by replacing the label of the estimation result with the per-image label assigned to the per-image labeled image, which is an image with the per-image label assigned thereto, and the estimation result has become smaller, when a third evaluation index is defined as an evaluation index for evaluating that, for at least the per-pixel labeled image among the per-pixel labeled image, the per-image labeled image, and the image without a label assigned thereto, the difference calculated from the estimation result has become smaller, without the label information of the per-pixel label assigned to the per-pixel labeled image or the per-image label assigned to the per-image labeled image being essential, A deep learning model generation method for generating a deep learning model based on an evaluation index obtained by adding at least one of the second evaluation index or the third evaluation index to the first evaluation index.
Citation Information
Patent Citations
Active learning method and active learning device
JP2020154602A
Deep layer learning device, image generation device and deep layer learning method
JP2021149160A