Model generation device, model generation program, model generation method, learned model, and image analysis device
The learning model generation device with a linear discriminator addresses the challenges of high labor and accuracy issues in image analysis by enabling efficient learning with minimal labeled pixels, achieving accurate pixel classification.
Patent Information
- Application Number
- JP2023190901
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-11-08
- Publication Date
- 2025-07-02
- Estimated Expiration
- 2043-11-08
AI Technical Summary
Existing image analysis technologies require large amounts of learning data and manual annotation, leading to high labor costs and insufficient accuracy when using small datasets, especially due to memory constraints and inefficient use of multi-layer perceptrons in the head part of neural networks.
A learning model generation device that utilizes a neural network with a multi-layer structure and a linear discriminator for estimating pixel labels, allowing learning with only a fraction of labeled pixels and reducing the need for extensive manual annotation.
Enables high-accuracy image analysis with reduced labor and time by using a small number of labeled pixels and learning images, improving computational efficiency and speed.
Smart Images

Figure 0007701691000001 
Figure 0007701691000002 
Figure 0007701691000003
Abstract
Description
Technical Field
[0001] The present invention relates to a learning model generation device, a learning model generation program, a learning model generation method, a learned model, and an image analysis device for inferring the label of a pixel of an image.
Background Art
[0002] Conventionally, machine learning techniques have been known and are used in various fields such as text generation and image processing. In particular, image processing is generally said to be a forte of so-called AI (Artificial Intelligence), and machine learning techniques are also used for example in the analysis of images.
[0003] Regarding the technique of performing image recognition using machine learning techniques, for example, Patent Document 1 is known. Patent Document 1 discloses a technique of inputting an image into a multi-layer neural network and estimating pixel classification.
[0004] In such a technique, generally, after extracting high-order features while reducing the resolution by a convolutional neural network (CNN) (the part that performs this process is called a backbone), the features are integrated to increase the resolution (the part that performs this process is called a neck), and finally, pixel classification is estimated using the integrated features (the part that performs this process is called a head).
[0005] In a model with such a structure, in the neck process, it is necessary to retain high-resolution and high-dimensional features, and the memory used tends to increase. Therefore, due to memory constraints, the batch size cannot be increased during learning, and there is a problem that the accuracy of image recognition by the generated model decreases due to bias in the loss during learning.
[0006] As a document describing such a problem, Non-Patent Document 1 is known. In Non-Patent Document 1, instead of performing classification learning for all pixels included in an image, after calculating image features using a backbone, only for randomly selected points among the pixels included in the image, operations of the neck and head are performed, and learning is carried out from the results.
[0007] As a result, it is only necessary to execute operations of the neck and head for the selected points, so the memory efficiency is improved and the batch size can be increased. Therefore, pixels of an image can be classified (labeled) with high precision.
Prior Art Documents
Patent Documents
[0008]
Patent Document 1
Non-Patent Documents
[0009]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0010] By the way, in order to perform highly accurate image analysis in the above-described technology, generally a large amount of learning data is required. In addition, annotations (correct labels) need to be assigned to all pixels of the learning images, and the labor required to prepare one learning image was also extremely large. For this reason, the learning process had to be carried out by experts or the like, and the time, cost, and human cost were high in order to obtain a learned model suitable for the user's application.
[0011] Also, in the technology described in Non-Patent Document 1, learning is performed only on some randomly selected points, but a multi-layer perceptron is used for the head, and many items such as the model structure and learning settings need to be adjusted. In addition, sufficient accuracy could not be obtained in the situation where learning is performed on a small number of randomly selected points or the number of learning images is small.
[0012] Regarding such problems, as a result of intensive research, the present inventor has found that by adopting a model with a specific structure, it is possible to perform highly accurate image analysis (pixel classification) even under disadvantageous conditions such as a small amount of learning data.
[0013] Based on the above research results, an object of the present invention is to provide a novel technology that can achieve accuracy sufficient for practical use according to the application and can easily realize image analysis by machine learning technology.
Means for Solving the Problems
[0014] In order to solve the above problems, the present invention is a learning model generation device that generates a model for applying to a device that estimates the labels of pixels in an image, and includes a feature amount calculation unit, an estimation unit, and a learning unit. The feature amount calculation unit obtains a feature amount corresponding to the labeled pixel for a learning image composed of labeled pixels and unlabeled pixels, using a neural network model with a multi-layer structure. The estimation unit estimates the label of the labeled pixel by a linear discriminator based on the feature amount. The learning unit corrects the calculation parameters of the linear discriminator so that the estimation result by the estimation unit matches the label of the learning image, and performs learning.
[0015] With such a configuration, learning can be performed using a learning image in which only some pixels are labeled, and a learned model can be generated. Here, as a result of research by the present inventor, it has been found that by adopting a linear discriminator for estimation based on feature amounts, it is possible to perform image analysis (pixel classification) with sufficient accuracy even using such a learning image. Therefore, according to the present invention, it is possible to generate a learned model that can perform image analysis with sufficient accuracy using a learning image that can be easily created.
[0016] In a preferred form of the present invention, there are a plurality of types of the labels, and the learning unit performs learning for at least one or more of the labeled pixels for each type of label.
[0017] With such a configuration, it becomes possible to perform learning according to the label to be classified, and an effect of improving analysis accuracy is expected. Here, as a result of research by the present inventor, according to the model of the present invention that adopts a linear discriminator for estimation based on feature amounts, it has been confirmed that it is possible to perform image analysis with sufficient accuracy using a learning image including at least one labeled pixel for each label without giving a large amount of annotation in the learning image.
[0018] In a preferred embodiment of the present invention, the number of labeled pixels is 1 / 2 or less of the number of pixels of the learning image per learning image.
[0019] By adopting such a configuration, the labor for creating the learning image can be greatly reduced.
[0020] In a preferred embodiment of the present invention, the number of labeled pixels is 1 / 100 or less of the number of pixels of the learning image per learning image.
[0021] By adopting such a configuration, the labor for creating the learning image can be extremely greatly reduced. Further, as a result of research by the present inventor, according to the model of the present invention that employs a linear discriminator for estimation based on feature amounts, it has been confirmed that even under the condition that the annotation of the learning image is extremely less compared with the prior art, it is possible to analyze the image with sufficient accuracy.
[0022] In a preferred embodiment of the present invention, the number of learning images used for learning is 100 or less.
[0023] By adopting such a configuration, the labor for creating the learning image can be extremely greatly reduced. Generally, in order to generate a model for performing image annotation, it is necessary to prepare about several tens of thousands of learning images. However, as a result of research by the present inventor, according to the model of the present invention that employs a linear discriminator for estimation based on feature amounts, it has been confirmed that it is possible to analyze the image with sufficient accuracy even with a small number of learning images.
[0024] In a preferred embodiment of the present invention, the learning unit performs learning by correcting the operation parameters of the linear discriminator without changing the parameters of the neural network model.
[0025] By adopting such a configuration, the burden during learning is small, and the time required for learning is significantly shortened.
[0026] In a preferred embodiment of the present invention, the label classifies the presence or absence of an abnormality and / or the type of abnormality.
[0027] In order to solve the above problems, the present invention is a learning model generation program that causes a computer to function as a device for generating a model for applying to a device that estimates labels of pixels in an image. The program causes the computer to function as a feature quantity calculation unit, an estimation unit, and a learning unit. The feature quantity calculation unit acquires a feature quantity corresponding to the labeled pixel for a learning image composed of labeled pixels and unlabeled pixels using a neural network model having a multi-layer structure. The estimation unit estimates the label of the labeled pixel using a linear discriminator based on the feature quantity. The learning unit corrects the calculation parameters of the linear discriminator so that the estimation result by the estimation unit matches the label of the learning image and performs learning.
[0028] In order to solve the above problems, the present invention is a learning model generation method for applying to a device that estimates labels of pixels in an image. The method includes a feature quantity calculation step of acquiring a feature quantity corresponding to the labeled pixel for a learning image composed of labeled pixels and unlabeled pixels using a neural network model having a multi-layer structure, an estimation step of estimating the label of the labeled pixel using a linear discriminator based on the feature quantity, and a learning step of correcting the calculation parameters of the linear discriminator so that the estimation result in the estimation step matches the label of the learning image and performing learning.
[0029] In order to solve the above problems, the present invention provides a learned model for causing a computer to function so as to estimate the label of a pixel in an image, the learned model including a neural network having a multi-layer structure and a linear discriminator coupled so as to receive information based on the output of the neural network and output an estimation result of the label of the pixel. The neural network outputs a feature amount corresponding to a pixel in the input image, and the linear discriminator has learned operation parameters such that the output matches the label of the labeled pixel for a training image composed of labeled pixels and unlabeled pixels. For each pixel in the input image input to the neural network, an operation is performed based on the operation parameters of the neural network and the linear discriminator, and the computer is caused to function so as to output an estimation result of the label for each pixel in the input image.
[0030] In order to solve the above problems, the present invention provides an image analysis apparatus for estimating the label of a pixel in an image, the apparatus including an acquisition unit and a classification unit. The acquisition unit acquires an input image for which the label of a pixel is to be estimated, and the classification unit inputs the input image to a learned model and outputs an estimation result of the label for each pixel in the input image. The learned model includes a neural network having a multi-layer structure and a linear discriminator coupled so as to receive information based on the output of the neural network and output an estimation result of the label of the pixel. The neural network outputs a feature amount corresponding to a pixel, and the linear discriminator has learned operation parameters such that the output matches the label of the labeled pixel for a training image composed of labeled pixels and unlabeled pixels.
Advantages of the Invention
[0031] According to the present invention, it is possible to realize accuracy that can withstand practical use according to the application, and it is possible to provide a novel technique for easily realizing image analysis by machine learning technology.
Brief Description of the Drawings
[0032]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Mode for Carrying Out the Invention
[0033] Hereinafter, with reference to the accompanying drawings, a more detailed description will be given. Although the preferred embodiments are shown in the drawings, it can be implemented in many different forms and is not limited to the embodiments described in this specification.
[0034] For example, in this embodiment, the configuration, operation, etc. of the image analysis system will be described. However, the same effects can also be achieved by an apparatus having the same functions, a method executed by the apparatus, a computer program for causing a computer device to execute the method, etc. The program may be provided as a non-transitory computer-readable recording medium or may be provided so as to be downloadable from an external server.
[0035] In the present invention, "annotation" is an annotation corresponding to the position of an image. In this embodiment, in particular, it refers to a label indicating the classification of pixels. The label may be assigned to each pixel, or may be something that specifies the label by indicating a certain range. In the latter case, the pixels existing in the specified range may be treated as if the specified label is assigned.
[0036] The label indicates what is represented in the part corresponding to the pixels in the image. For example, in an image composed of a person, an animal, and a background, it is information indicating whether each pixel is part of a person, part of an animal, or part of the background. The types of labels can be arbitrarily determined. For example, in order to detect defective parts of a product, labels indicating normal or abnormal may be assigned. Furthermore, there may be multiple types of labels indicating abnormalities, differentiated by further types.
[0037] In another example, for instance, classification of cell types in a microscopic image, classification of ingredients in a bento box, detection or classification of organisms included in an image (such as organisms in an aquarium), detection of screw holes and gripping points of industrial products and parts, detection and counting of parts in a tray, etc. can also be performed depending on how the labels are set. Also, between multiple images or frames of a moving image, by attaching a label indicating the moving location, it can be used for detecting the moving location.
[0038] <1. Functional Configuration> FIG. 1 is a block diagram showing the configuration of the image analysis system of the present embodiment. As shown in FIG. 1, the image analysis system 0 includes a model generation device 1, an analysis device 2, and a storage device DB. The model generation device 1 and the storage device DB are connected by wire or wirelessly and configured to be communicable. Also, the storage device DB and the analysis device 2 are connected by wire or wirelessly and configured to be communicable. There are no restrictions on the type of communication protocol applied between these devices, the type of network, etc.
[0039] The model generation device 1 generates a learned model by performing learning based on learning images and stores it in the storage device DB. The analysis device 2 analyzes the input image using the learned model stored in the storage device DB. Here, the analysis of the input image refers to estimating a label for the pixels of the input image.
[0040] <1.1. Functional Configuration of Model Generation Device 1> The model generation device 1 includes an acquisition unit 11, a feature amount calculation unit 12, an estimation unit 13, and a learning unit 14. This is a specific implementation of information processing by software using hardware.
[0041] The acquisition unit 11 acquires learning images. The learning images are composed of labeled pixels and unlabeled pixels. In this embodiment, the acquisition unit 11 generates learning images by receiving an image input from a user and further receiving a label specification for the pixels of the image. The pixels for which the label is specified become labeled pixels. The generation of learning images will be described later.
[0042] The feature amount calculation unit 12 calculates the feature amount for each pixel of the learning image acquired by the acquisition unit 11 using a neural network model with a multi-layer structure. In this embodiment, the convolutional neural network (CNN) is used to calculate the image feature amount by the feature amount calculation unit 12, and the feature amount corresponding to the labeled pixels in particular is acquired from among them.
[0043] The estimation unit 13 estimates the label of the labeled pixels of the learning image using a linear discriminator based on the feature amount corresponding to the labeled pixels acquired by the feature amount calculation unit 12. As the linear discriminator, for example, a simple perceptron, logistic regression, linear support vector machine, Fisher's linear discriminator, etc. can be used for the estimation unit 13 to perform the estimation. The estimation here is a process of calculating the score for each label based on the feature amount of the pixel. The one with the highest score may be estimated as the label of the pixel.
[0044] Here, in order to estimate the label corresponding to the pixel of the image, the features extracted by the CNN (corresponding to the backbone) are integrated and upsampled to the same resolution as the input image (corresponding to the neck). Note that the resolution does not necessarily have to be upsampled to the same level as the input image. For example, features with a resolution appropriately reduced to 1 / 4, 1 / 8, 1 / 16, 1 / 32, etc. of the input image may be used.
[0045] FIG. 3 is a schematic diagram showing the structure of the model in this embodiment. The model includes a feature extraction layer, a feature integration layer, and a discrimination layer. The feature amount calculation unit 12 of this embodiment acquires the output of the CNN in the feature extraction layer, and further executes an operation of integrating and increasing the resolution of this output in the feature integration layer, and delivers the result to the estimation unit 13. The estimation unit 13 estimates the label corresponding to the pixel of the input image by inputting it to the linear discriminator in the discrimination layer.
[0046] Regarding the operation of the feature integration layer, any method can be used. For example, similar to the technique described in Non-Patent Document 1, a network with a HyperColumn structure may be used as the feature integration layer. The network with a HyperColumn structure is one of the methods for integrating features extracted from a plurality of layers corresponding to the same pixel.
[0047] More specifically, the network with a HyperColumn structure can obtain the features of each pixel by aligning the features with different resolutions extracted by the backbone to the same resolution as the input image or about one-fourth of the resolution of the input image by an enlargement process and combining them by concatenate (combination process).
[0048] The learning unit 14 performs learning by correcting the operation parameters of the linear discriminator so that the estimation result by the estimation unit 13 matches the label of the learning image (the label of the pixel with the target label). Since the linear discriminator has significantly fewer parameters to be adjusted compared to the non-linear discriminator, learning can be performed at high speed.
[0049] <1.2. Functional Configuration of Analysis Device 2> The analysis device 2 includes an acquisition unit 21, a feature amount calculation unit 22, and an estimation unit 23. The analysis device 2 analyzes an unknown image using a learned model. The acquisition unit 21 acquires an input image to be analyzed for pixel labeling. Further, the feature amount calculation unit 22 and the estimation unit 23 function as the classification unit of the present invention.
[0050] Here, the process of estimating the labels of the pixels is similar to the learning process performed by the model generating device 1. That is, the feature calculation unit 22 and the estimation unit 23 execute the same process as the feature calculation unit 12 and the estimation unit 13, respectively, on pixels of an unknown input image using a trained model stored in the storage device DB. Therefore, detailed description of the processes will be omitted here.
[0051] The model generating device 1, the analysis device 2, and the storage device DB may be partially or entirely realized by a plurality of computers operating in cooperation with each other. These devices may also be realized by a single computer. In this case, the acquiring unit 11 and the acquiring unit 21, the feature calculating unit 12 and the feature calculating unit 22, and the estimating unit 13 and the estimating unit 23 may each serve as a single computer.
[0052] <2. Hardware configuration> An information processing device 9 (computer device) such as a general-purpose server or a personal computer can be used as the model generating device 1, the analysis device 2, and the storage device DB. In this embodiment, the model generating device 1 is the information processing device 9 in which a computer program (learning model generating program) that executes the learning model generating method is installed.
[0053] Fig. 2 is a hardware configuration diagram of the information processing device 9. As shown in Fig. 2, the information processing device 9 has a control unit 901, a storage unit 902, a communication unit 903, an input unit 904, and an output unit 905, which are used to perform the functions of each unit and each process.
[0054] The control unit 901 has a processor such as a CPU capable of executing an instruction set, and executes an OS, programs, and the like. The storage unit 902 includes a volatile memory such as a RAM capable of storing an instruction set, and a non-volatile recording medium such as an HDD or SSD capable of recording an OS, a determination program, and the like. The communication unit 903 has an interface for physically connecting to a network, and controls communications with the network NW to input and output information. The input unit 904 includes an operation input device capable of input processing such as a touch panel or a keyboard, a voice input device capable of voice input such as a microphone, and the like. The output unit 905 includes a display device capable of display processing such as a display, and a voice output device such as a speaker.
[0055] <3. Learning Image Generation> FIG. 4 is a flowchart showing the process related to learning image generation. As described above, in the present embodiment, the acquisition unit 11 generates a learning image by receiving an input of an image from the user and receiving a label specification for the pixels of the image.
[0056] Here, in the present embodiment, it is assumed that the user who intends to use the analysis device 2 prepares a learning image by himself / herself and generates a model that has been independently learned using the model generation device 1. Therefore, as described below, it is preferable to enable the generation of a learning image by a UI that can assign a label to an image by an intuitive operation.
[0057] First, in step S11, the user inputs an image, and the acquisition unit 11 acquires the image. Next, in step S12, the acquisition unit 11 receives a label specification for the pixels of the image from the user. FIG. 5 is an example of a UI related to learning image generation. The model generation device 1 displays a screen as shown in FIG. 5 on the display and receives various inputs from the user. The learning image display unit LI displays the image acquired by the acquisition unit 11. In step S12, the user selects a label to be assigned to the pixel by performing an input via the label button LB.
[0058] Next, in step S13, when the user selects a position (pixel) in the learning image display unit LI, the acquisition unit 11 performs a process of assigning the label selected in step S12 to the pixel selected in step S13 (step S14). The label assigned to the pixel is displayed superimposed on the image in the learning image display unit LI. Note that the order of step S12 and step S13 may be reversed. Also, the selection of pixels in step S13 may be such that the position on the image is specified within a range, and pixels within the range can be selected collectively.
[0059] After steps S11 to S13 are repeated for arbitrary pixels, when the "learning" button is selected by the user, the feature amount calculation unit 12, the estimation unit 13, and the learning unit 14 execute the following learning process based on the learning image including the labeled pixels corresponding to the input.
[0060] When the user selects the "learning result confirmation" button, the acquisition unit 11 analyzes the acquired image using the model for which the learning process has been performed based on the learning image. Then, in the learning image display unit LI, the analysis result by the model, that is, the estimation result of the label for each pixel, is further superimposed and displayed. The user can check the result, further assign labels, and perform additional learning. Also, if the labels are appropriately assigned to the learning image, the user can finish learning for that learning image, complete the creation of the model, or input other learning images.
[0061] <4. Learning model generation method> Next, with reference to FIG. 6, the learning process by the model generation device 1 will be described. FIG. 6 shows the procedure for learning one learning image. The process of FIG. 6 may be performed for each learning image. As described above, in this embodiment, it is assumed that the general user operates to assign labels to the pixels of the learning image, and only some points (pixels) in the learning image are labeled. The number of labeled pixels is not limited to this, but for example, it is 1 / 2 or less of the number of pixels in the learning image, 1 / 10 or less of the number of pixels in the learning image, 1 / 100 or less of the number of pixels in the learning image, 1 / 1000 or less of the number of pixels in the learning image, or 1 / 10000 or less of the number of pixels in the learning image. Alternatively, the number of labeled pixels is 100 or less, 75 or less, 50 or less, 30 or less, or 10 or less. In each learning image, it is sufficient that there is at least one or more labeled pixels.
[0062] Also, there is no particular limitation on the number of learning images used for learning. However, in the present embodiment, learning is performed using 200 or fewer, 100 or fewer, 75 or fewer, 50 or fewer, 30 or fewer, 20 or fewer, or 10 or fewer learning images. The number of learning images may be at least 1 or more.
[0063] Here, it is not necessary for all types of labeled pixels to exist in each learning image. However, in the present embodiment, the learning images and the labeled pixels are selected such that at least 1 or more labeled pixels exist for each type of label when considering all the learning images. That is, in the present embodiment, for example, when generating a model that assigns three types of labels, A, B, and C, using two learning images, at least one labeled pixel with each of the labels A, B, and C is included in either of the two learning images.
[0064] In the learning process, first, in step S21, the acquisition unit 11 acquires a learning image composed of labeled pixels and unlabeled pixels by the above-described generation process. Next, in the feature amount calculation step of step S22, the feature amount calculation unit 12 calculates the feature amount of each pixel by a multi-layer CNN, generates an integrated feature in which the features are integrated by the feature integration layer for each pixel, acquires the integrated feature corresponding to the labeled pixel, and passes it to the estimation unit 13. Note that when generating a model for estimating a label for each fixed range including a plurality of pixels instead of for each pixel, the integrated feature may also be generated for each such range.
[0065] Next, in the estimation step of step S23, the estimation unit 13 inputs the integrated feature received from the feature amount calculation unit 12 into a linear discriminator. The estimation unit 13 estimates the label of the pixel according to the output of the linear discriminator.
[0066] In the learning process of step S24, the learning unit 14 corrects the calculation parameters of the linear discriminator so that the label at the target pixel of the learning image matches the estimation result by the estimation unit 13. Then, in step S25, it is determined whether there is an unlearned labeled pixel, and steps S22 to S24 are repeated for one learning image until there is no unlearned labeled pixel. Thereby, learning is performed for all the labeled pixels in the target learning image.
[0067] <5. Image analysis method> Next, with reference to FIG. 7, the image analysis process by the analysis device 2 will be described. The analysis device 2 estimates the label of a pixel for an unknown image (an image other than the learning image) using the learned model created by the above-described procedure. The estimation of the label is performed by executing the same processing as steps S22 and S23 of the learning process described above for all the pixels of the input image using the learned model. This will be described in detail below.
[0068] First, in step S31, the acquisition unit 21 acquires an input image to be the target of label estimation. Next, in step S32, the feature amount calculation unit 22 calculates the feature amount of the pixel using the learned model stored in the storage device DB. The feature amount is calculated by the feature extraction layer constituted by CNN as described above and integrated by the feature integration layer. Note that it is not always necessary to integrate the feature amounts so as to correspond to each pixel. For example, the feature amounts may be integrated so as to correspond to a certain range in the image to calculate an integrated feature, and the integrated feature may be regarded as the feature amount corresponding to each of the plurality of pixels. That is, it is not necessary to make the resolution the same as that of the input image, and the integrated feature may be set to a resolution appropriately reduced, such as 1 / 4, 1 / 8, 1 / 16, 1 / 32, etc. of the input image as described above.
[0069] In step S33, the estimation unit 23 estimates the label of a pixel using the learned model. Specifically, when the integrated feature is input to the discrimination layer of the learned model, the scores of the respective labels corresponding to the integrated feature are output. The estimation unit 23 estimates the label with the highest score as the label of the pixel corresponding to the integrated feature.
[0070] Then, in step S34, it is confirmed whether there is a pixel for which the above-described estimation process has not been performed (unclassified pixel). If there is an unclassified pixel, the target pixel is changed and steps S23 to S33 are repeated. Thereby, the label is estimated for all pixels.
[0071] As described above, according to the configuration of the present invention, by performing learning using a training image in which labels are attached only to some pixels, a learned model for estimating the label of a pixel for an unknown image can be generated. Further, using the learned model generated in this way, analysis of an image for estimating the label of a pixel can be performed. The present invention can be applied to, for example, an inspection apparatus for discovering defects of a product based on a product image.
[0072] <6. Remarkable Effects of the Present Invention> Hereinafter, the remarkable effects of the present invention will be described. First, the differences between the prior art and the present invention will be described. As described in the description of the background art, in the technique described in Non-Patent Document 1, at the time of model learning, by performing labeling learning only for some pixels, the batch size, which was conventionally reduced due to memory constraints, was increased to improve the accuracy.
[0073] In the technique described in Non-Patent Document 1, a multi-layer perceptron (MLP) and a stochastic gradient descent method (SGD) are adopted in the head part. In contrast, the present invention is different from the technique described in Non-Patent Document 1 in that a linear discriminator is used in the head part. In addition, the technique described in Non-Patent Document 1 assumes learning using an image in which annotations are attached to all pixels. In contrast, the present invention is different from the technique described in Non-Patent Document 1 in that learning is performed using a training image composed of labeled pixels and unlabeled pixels.
[0074] <6.1 Estimation Accuracy Comparison Experiment> As described above, in machine learning, it is common to need to learn a large number of images, and even in the technique described in Non-Patent Document 1, thousands of images are learned. However, when trying to generate a model specialized for a specific application, it is often difficult to prepare such a large number of teacher images. Therefore, a comparative experiment on the estimation accuracy was conducted using a model generated under the condition that the number of training images was extremely small.
[0075] (1) Model Configuration A comparative experiment was conducted using Model A (with the same configuration as the model described in Non-Patent Document 1) that employed a multi-layer perceptron (MLP), a non-linear discriminator, in the discrimination layer, and Model B (with the same configuration as the model of the present invention) that employed a linear discriminator in the discrimination layer. For both models, a convolutional neural network is used in the feature extraction layer, and the configuration of the feature integration layer is the same. The two models are common except for the configuration of the discrimination layer.
[0076] (2) Training Images The image shown in Fig. 8(a) was used as the training image. Labels were attached to the pixels at the positions indicated by dots in Fig. 8(a). Specifically, a total of 8 labeled pixels were set, with 1 point on the circular object, 1 point on the rectangular object, 1 point on the linearly rising object, 1 point on the linearly falling object, and 4 points on the background, each with a label indicating the respective object or background. As the training image, only 1 image shown in Fig. 8(a) was used.
[0077] (3) Results Figure 8(b) is a diagram showing the estimation result by Model A when an image identical to the learning image is used as the input image. Figure 8(c) is a diagram showing the estimation result by Model B under the same conditions. In both cases, the pixels of the input image are labeled with five types: circular object, rectangular object, linear object with an upward right slope, linear object with a downward right slope, and background.
[0078] From the output of Model A, it can be seen that the discrimination between the background and each object is insufficient, and the shape of the object is not well represented. On the other hand, from the output of Model B, it can be seen that the range of each object can be correctly recognized, and for the part of the unlabeled pixels, the label of the input image is accurately estimated.
[0079] When calculating the accuracy rate for each label, the following results were obtained. Background Model A: 97.58%, Model B: 99.02% Rectangular object Model A: 33.90%, Model B: 98.75% Circular object Model A: 25.39%, Model B: 98.73% Linear object with an upward right slope Model A: 62.00%, Model B: 74.39% Linear object with a downward right slope Model A: 12.51%, Model B: 82.96%
[0080] Thus, under the conditions where the learning image and the labeled pixels are extremely few, with Model A using a non-linear discriminator, the label could not be estimated with a practical accuracy. On the other hand, it was shown that with Model B using a linear discriminator, even under the conditions where the learning image and the labeled pixels are extremely few, the label can be estimated with high accuracy. Thus, by using a linear discriminator in the discrimination layer, the significant improvement in accuracy under specific conditions is a remarkable effect that cannot be predicted from the conventional technology.
[0081] <6.2 Computational Complexity> (1) Comparison with PixelNet The present invention has an advantageous effect in terms of computational complexity as compared with the prior art. For example, in the technique described in Non-Patent Document 1, the structures of the backbone and the neck are generally the same as those of the present invention, but a multi-layer perceptron is used in the head part. In the head of the technique described in Non-Patent Document 1, taking the features received from the neck as input, the weights of neurons with 4096 neurons in each layer are calculated.
[0082] In the technique described in Non-Patent Document 1, since there are two hidden layers, assuming the number of classes to be classified is k, in the head part, there are parameters of the dimensionality of the features input to the head × 4096 + 4096 × 4096 + 4096 × k. Since calculations are performed considering a large number of parameters, the computational complexity increases and the calculation speed decreases.
[0083] On the other hand, if a linear discriminator is used as in the present invention, the number of parameters decreases. Specifically, since scores are calculated linearly for each class to be classified, the number of parameters is the dimensionality of the features input to the head × k. As a result, the computational complexity decreases, the processing load is reduced, and the calculation speed is improved.
[0084] (2) Speed Comparison Experiment with General Models Furthermore, a comparison experiment of the calculation speed between the present invention's model and a general image labeling (Semantic Segmentation) model was conducted. The DeepLabV3 model was used as the comparison target. The DeepLabV3 model is a model configured such that the resolution is not significantly reduced in the backbone by atlas convolution (asymmetric convolution) (Reference: “DeepLabv3”, [online], Search date: October 11, 2023, Internet <URL:https: / / paperswithcode.com / method / deeplabv3>). Also, as the model of the present invention, a model with a linear discriminator adopted in the head was used.
[0085] The conditions are as follows. · Use the same GPU (GeForce (registered trademark) RTX 3080) in both models · Input image size: 288×288 pixels · Backbone of both models: ResNet50 (a convolutional neural network with 50 layers of depth) · Number of trials: Label the images 10 times and calculate the average time to label one image.
[0086] As a result, the calculation time of the DeepLabV3 model was 73 ms on average, and the calculation time of the model of the present invention was 11 ms. Thus, it was confirmed by experiments that the present invention can perform calculations at high speed compared to general models.
Explanation of Signs
[0087] 0: Image analysis system 1: Model generation device 2: Analysis device 9: Information processing device 11: Acquisition unit 12: Feature calculation unit 13: Estimation unit 14: Learning unit 21: Acquisition unit 22: Feature calculation unit 23: Estimation unit 901: Control unit 902: Memory unit 903: Communication unit 904: Input unit 905: Output unit LB: Label button LI: Learning image display unit NW: Network
Claims
1. A model generation device that generates a model for estimating labels of pixels in an image, comprising: a feature amount calculation unit, an estimation unit, and a learning unit, wherein the feature amount calculation unit inputs a training image composed of labeled pixels and unlabeled pixels into a convolutional neural network model with a multi-layer structure, and obtains, as an output of the convolutional neural network model, a feature amount corresponding to the labeled pixels based on the labeled pixels and the unlabeled pixels, the estimation unit inputs an input based on the feature amount into a linear discriminator, and estimates the label of the labeled pixel according to the output of the linear discriminator, the learning unit generates the model in which the convolutional neural network model and the linear discriminator are combined by correcting the calculation parameters of the linear discriminator so that the estimation result by the estimation unit matches the label of the labeled pixel in the training image, wherein pixels other than the labeled pixels in the training image are unlabeled pixels, the model generation device.
2. There are multiple types of the labels, the feature amount calculation unit obtains, based on the labeled pixels and the unlabeled pixels, the feature amounts corresponding to at least one or more of the labeled pixels for each type of label using the convolutional neural network model, the estimation unit inputs, for each of the labeled pixels for which the feature amount calculation unit has obtained the feature amount, an input based on the feature amount into the linear discriminator, and estimates the label of the labeled pixel according to the output of the linear discriminator, the learning unit generates the model in which the convolutional neural network model and the linear discriminator are combined by repeatedly correcting the calculation parameters of the linear discriminator so that the estimation result by the estimation unit matches the label of the labeled pixel in the training image for each of the labeled pixels for which the estimation unit has estimated the label, the model generation device according to Claim 1.
3. The number of the labeled pixels is 1 / 2 or less of the number of pixels in one training image, the model generation device according to Claim 1.
4. The number of the labeled pixels is 1 / 100 or less of the number of pixels in one training image, the model generation device according to Claim 1.
5. The model generation device according to claim 1, wherein the number of learning images used for learning is 100 or less.
6. The model generation device according to claim 1, wherein the learning unit corrects the operation parameters of the linear discriminator so that the estimation result by the estimation unit matches the label of the labeled pixel in the learning image without changing the parameters of the convolutional neural network model.
7. The model generation device according to claim 1, wherein the label includes a label for classifying the presence or absence of an abnormality or a label for classifying the type of an abnormality.
8. A model generation program that causes a computer to function as a device for generating a model for estimating the label of a pixel in an image, causing the computer to function as a feature quantity calculation unit, an estimation unit, and a learning unit, The feature quantity calculation unit inputs a learning image composed of labeled pixels and unlabeled pixels into a convolutional neural network model having a multi-layer structure, and obtains, as an output of the convolutional neural network model, a feature quantity corresponding to the labeled pixel based on the labeled pixel and the unlabeled pixel. The estimation unit inputs an input based on the feature quantity to a linear discriminator, and estimates the label of the labeled pixel according to the output of the linear discriminator. The learning unit corrects the operation parameters of the linear discriminator so that the estimation result by the estimation unit matches the label of the labeled pixel in the learning image, and performs learning. A model generation program, wherein pixels other than the labeled pixels in the learning image are unlabeled pixels.
9. A model generation method for generating a model for estimating the label of a pixel in an image, A feature quantity calculation step of inputting a learning image composed of labeled pixels and unlabeled pixels into a convolutional neural network model having a multi-layer structure, and obtaining, as an output of the convolutional neural network model, a feature quantity corresponding to the labeled pixel based on the labeled pixel and the unlabeled pixel; An estimation step of inputting an input based on the feature quantity to a linear discriminator, and estimating the label of the labeled pixel according to the output of the linear discriminator; A learning step of generating the model in which the convolutional neural network model and the linear discriminator are combined by correcting the calculation parameters of the linear discriminator so that the estimation result in the estimation step matches the label of the labeled pixel in the learning image; A model generation method in which pixels other than the labeled pixels of the learning image are pixels without labels.
10. A learned model for causing a computer to function, which has learned a function of estimating the label of a pixel of an image by using the image as an input, Composed of a convolutional neural network with a multi-layer structure and a linear discriminator combined so that information based on the output of the convolutional neural network is input and an estimation result of the label of a pixel is output, The convolutional neural network outputs feature amounts corresponding to some of the plurality of pixels based on the plurality of pixels of the input image, The linear discriminator corrects calculation parameters so that the label of the labeled pixel matches the estimation result for the labeled pixel for a learning image composed of labeled pixels and pixels without labels, and thus learns the function of outputting an estimation result of the label of a pixel from an input based on the output of the convolutional neural network, Pixels other than the labeled pixels of the learning image are pixels without labels, A learned model for causing a computer to function so as to perform a calculation based on the calculation parameters of the convolutional neural network and the linear discriminator for each pixel of an input image input to the convolutional neural network and output an estimation result of the label for each pixel of the input image.
11. An image analysis apparatus for estimating the label of a pixel of an image, Comprising an acquisition unit and a classification unit, The acquisition unit acquires an input image for which the label of a pixel is to be estimated, The classification unit outputs an estimation result of the label for each pixel of the input image by inputting the input image to a learned model, The learned model is Composed of a convolutional neural network with a multi-layer structure and a linear discriminator combined so that information based on the output of the convolutional neural network is input and an estimation result of the label of a pixel is output. The convolutional neural network outputs feature amounts corresponding to some of the plurality of pixels based on the plurality of pixels of the input image, The linear discriminator is trained to output an estimation result of a label of a pixel from an input based on the output of the convolutional neural network by correcting arithmetic parameters so that the label of the labeled pixel matches the estimation result for the labeled pixel for a training image composed of labeled pixels and unlabeled pixels, An image analysis apparatus, wherein pixels other than the labeled pixels of the training image are unlabeled pixels.
Citation Information
Patent Citations
Category discriminator generation apparatus, category discrimination device, and computer program
JP2015164012A
Detector, detection program, detection method, vehicle, parameter calculation device, parameter calculation program, and parameter calculation method
JP2016006626A
Information processing apparatus, information processing method, and program
JP2019192009A
Construction limit determination device
JP2020006788A
Sampling device and sampling method
JP2022159720A