Learning model generation device, learning model generation program, learning model generation method, trained model, and image analysis device

The learning model generation device uses a linear discriminator to estimate pixel labels efficiently, addressing the resource-intensive and accuracy challenges of conventional image analysis by reducing annotation needs and enhancing performance with minimal training data.

JP2025078639APending Publication Date: 2025-05-20RESEARCH INSTITUTE OF SYSTEMS PLANNING INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025020194
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

Existing image analysis technologies require large amounts of training data and manual annotation, leading to high costs and resource consumption, and conventional methods struggle with accuracy when using small datasets or randomly selected points for learning.

Method used

A learning model generation device that utilizes a multi-layered neural network with a feature calculation unit, estimation unit, and learning unit, employing a linear discriminator to estimate pixel labels based on labeled and unlabeled pixels, reducing the need for extensive annotation and enabling accurate image analysis with minimal training data.

Benefits of technology

The solution allows for high-accuracy image analysis with significantly reduced effort and resources, achieving practical results even with a small number of labeled pixels and training images, and improves calculation speed by minimizing parameter adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025078639000001_ABST
    Figure 2025078639000001_ABST
Patent Text Reader

Abstract

To provide a new technique for easily achieving image analysis by a machine learning technique while achieving practical accuracy appropriate for the application.SOLUTION: A learning model generation device generates a model to be applied to a device that labels pixels of an image, the learning model generation device comprising a feature calculation unit, an estimation unit, and a training unit. The feature calculation unit calculates, for a learning image composed of labeled pixels and unlabeled pixels, features for a labeled pixel according to a neural network model having a multilayer structure. On the basis of the features, the estimation unit estimates the label of the labeled pixel according to a linear discriminator. The training unit performs training by correcting a calculation parameter of the linear discriminator so that a result estimated by the estimation unit matches the label of the learning image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a learning model generation device, a learning model generation program, a learning model generation method, a trained model, and an image analysis device for inferring labels of pixels in an image. [Background technology]

[0002] Machine learning technology has been known for a long time and is used in various fields such as text generation, image processing, etc. Image processing in particular is generally considered to be an area where so-called AI (Artificial Intelligence) excels, and machine learning technology is also used, for example, in image analysis.

[0003] A technique for recognizing an image using a machine learning technique is known, for example, from Patent Document 1. Patent Document 1 discloses a technique for inputting an image into a multi-layer neural network and estimating the classification of pixels.

[0004] In such technologies, generally, a Convolutional Neural Network (CNN) is used to extract high-level features while reducing the resolution (the part that performs this process is called the backbone), then the features are integrated to increase the resolution (the part that performs this process is called the neck), and finally the integrated features are used to estimate the classification of the pixels (the part that performs this process is called the head).

[0005] In models with this structure, it is necessary to retain high-resolution and high-dimensional features in neck processing, which tends to use a large amount of memory. Therefore, due to memory constraints, it is not possible to increase the batch size in training, and there is a problem that the accuracy of image recognition by the generated model decreases due to biased loss during training.

[0006] Non-Patent Document 1 is known as a document that describes such a problem. Non-Patent Document 1 discloses a technology in which, instead of performing classification learning for all pixels included in an image, image features are calculated using a backbone, and then neck and head calculations are performed only for points randomly selected from among the pixels included in the image, and learning is performed from the results.

[0007] This improves memory efficiency and allows for larger batch sizes, since neck and head operations only need to be performed on selected points, resulting in highly accurate classification (labeling) of image pixels. [Prior art documents] [Patent documents]

[0008] [Patent Document 1] JP 2019-192009 A [Non-patent literature]

[0009] [Non-Patent Document 1] Aayush Bansal, 4 others, “PixelNet: Representation of the pixels, by the pixels, and for the pixels.” [online], February 21, 2017, [Retrieved September 14, 2020], Internet<URL:https: / / arxiv.org / pdf / 1702.06506.pdf> Summary of the Invention [Problem to be solved by the invention]

[0010] However, in order to perform highly accurate image analysis using the above-mentioned technologies, a large amount of training data is generally required. In addition, all pixels in the training image must be annotated (labeled with the correct answer), and the effort required to prepare one training image is very large. For this reason, the training process had to be performed by experts, and it took a lot of time, money, and human resources to obtain a trained model suitable for the user's purpose.

[0011] In addition, the technology described in Non-Patent Document 1 learns only a portion of randomly selected points, but uses a multilayer perceptron as the head, which requires adjustment of many items such as the model structure and learning settings. Furthermore, sufficient accuracy cannot be obtained when learning only a small number of randomly selected points or when there are a small number of training images.

[0012] Regarding such problems, the inventor conducted intensive research and discovered that by adopting a model with a specific structure, it is possible to perform image analysis (classification of pixels) with high accuracy even under unfavorable conditions such as a small amount of training data.

[0013] Based on the above research results, the objective of the present invention is to provide a new technology that can achieve practical accuracy according to the application and easily realize image analysis using machine learning technology. [Means for solving the problem]

[0014] In order to solve the above problems, the present invention provides a learning model generation device that generates a model to be applied to a device that estimates labels of pixels in an image, the device comprising a feature calculation unit, an estimation unit, and a learning unit, wherein the feature calculation unit obtains features corresponding to labeled pixels for a training image composed of labeled pixels and unlabeled pixels by using a multi-layered neural network model, the estimation unit estimates labels of the labeled pixels based on the features using a linear discriminator, and the learning unit performs learning by modifying calculation parameters of the linear discriminator so that the estimation result by the estimation unit matches the label of the training image.

[0015] With this configuration, it is possible to perform learning using a training image in which only some of the pixels are labeled, and generate a trained model. As a result of research by the inventors, it was found that by adopting a linear discriminator for feature-based estimation, it is possible to perform image analysis (pixel classification) with sufficient accuracy even when using such training images. Therefore, the present invention makes it possible to generate a trained model capable of analyzing images with sufficient accuracy using training images that can be easily created.

[0016] In a preferred embodiment of the present invention, there are a plurality of types of labels, and the learning section performs learning on at least one labeled pixel for each type of label.

[0017] With this configuration, it becomes possible to perform learning according to the label to be classified, which is expected to improve the accuracy of analysis. As a result of research by the inventors, it has been confirmed that, according to the model of the present invention, which employs a linear discriminator for feature-based estimation, it is possible to analyze images with sufficient accuracy by using training images that contain at least one labeled pixel for each label, even without adding a large number of annotations to the training images.

[0018] In a preferred embodiment of the present invention, the number of labeled pixels for each training image is equal to or less than half the number of pixels in the training image.

[0019] With this configuration, the effort required to create learning images can be significantly reduced.

[0020] In a preferred embodiment of the present invention, the number of labeled pixels for each training image is equal to or less than 1 / 100 of the number of pixels in the training image.

[0021] With this configuration, the effort required for creating learning images can be significantly reduced. Furthermore, as a result of research by the inventors, it has been confirmed that the model of the present invention, which employs a linear discriminator for feature-based estimation, makes it possible to analyze images with sufficient accuracy even in conditions where the number of annotations in the training images is extremely small compared to conventional techniques.

[0022] In a preferred embodiment of the present invention, the number of training images used in training is 100 or less.

[0023] With this configuration, the effort required for creating learning images can be significantly reduced. Generally, to generate a model for image annotation, it is necessary to prepare tens of thousands of training images. However, as a result of research by the inventors, it has been confirmed that the model of the present invention, which employs a linear discriminator for feature-based estimation, can perform image analysis with sufficient accuracy even with a small number of training images.

[0024] In a preferred embodiment of the present invention, the learning section performs learning by modifying calculation parameters of the linear discriminator without changing parameters of the neural network model.

[0025] With this configuration, the burden on the learner during learning is small, and the time required for learning is significantly reduced.

[0026] In a preferred embodiment of the present invention, the label classifies the presence or absence of an abnormality and / or the type of abnormality.

[0027] In order to solve the above-mentioned problems, the present invention provides a learning model generation program that causes a computer to function as an apparatus for generating a model to be applied to an apparatus that estimates labels of pixels in an image, the learning model generation program causing the computer to function as a feature calculation unit, an estimation unit, and a learning unit, the feature calculation unit obtains features corresponding to labeled pixels for a training image composed of labeled pixels and unlabeled pixels using a multi-layered neural network model, the estimation unit estimates labels of the labeled pixels based on the features using a linear discriminator, and the learning unit performs learning by modifying calculation parameters of the linear discriminator so that the estimation result by the estimation unit matches the label of the training image.

[0028] In order to solve the above-mentioned problems, the present invention provides a learning model generation method for application to an apparatus that estimates labels of pixels in an image, the method including: a feature calculation step of acquiring feature values ​​corresponding to labeled pixels of a learning image constituted by labeled pixels and unlabeled pixels using a multi-layered neural network model; an estimation step of estimating labels of the labeled pixels using a linear discriminator based on the feature values; and a learning step of correcting calculation parameters of the linear discriminator and performing learning so that an estimation result in the estimation step matches the label of the learning image.

[0029] In order to solve the above problem, the present invention provides a trained model for causing a computer to function to estimate labels of pixels in an image, the trained model being composed of a multi-layered neural network and a linear discriminator that is coupled to receive information based on the output of the neural network and output an estimated result of the label of the pixel, the neural network outputs features corresponding to the pixels of an input image, and the linear discriminator has calculation parameters trained so that the output of the linear discriminator matches the labels of the labeled pixels for a training image composed of labeled pixels and unlabeled pixels, and the computer is caused to function to perform a calculation based on the calculation parameters of the neural network and the linear discriminator for each pixel of the input image input to the neural network, and output an estimated result of the label for each pixel of the input image.

[0030] In order to solve the above problem, the present invention provides an image analysis device that estimates labels of pixels in an image, comprising an acquisition unit and a classification unit, wherein the acquisition unit acquires an input image for estimating pixel labels, and the classification unit inputs the input image to a trained model to output an estimation result of the label for each pixel of the input image, the trained model being composed of a multi-layered neural network and a linear discriminator that receives input information based on the output of the neural network and is coupled to output an estimation result of the label of a pixel, the neural network outputs features corresponding to pixels, and the linear discriminator has calculation parameters trained so that the output matches the labels of the labeled pixels for a training image composed of labeled pixels and unlabeled pixels. Effect of the Invention

[0031] According to the present invention, it is possible to provide a novel technology that can achieve practical accuracy according to the application and easily realize image analysis using machine learning technology. [Brief description of the drawings]

[0032] [Figure 1]FIG. 1 is a block diagram showing a configuration of a system according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a hardware configuration diagram of the information processing apparatus according to the embodiment. [Diagram 3] FIG. 2 is a schematic diagram showing the structure of a learning model according to the present embodiment. [Figure 4] 4 is a flowchart showing a process relating to learning image generation according to the embodiment. [Diagram 5] 4 is an example of a screen related to learning image generation according to the present embodiment. [Figure 6] 4 is a flowchart showing a processing procedure related to a learning model generation method of the present embodiment. [Figure 7] 4 is a flowchart showing a processing procedure of image analysis by the image analyzing device of the present embodiment. [Figure 8] FIG. 11 is a diagram showing an example of an image analysis result in the embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0033] The present invention will now be described in more detail with reference to the accompanying drawings, in which preferred embodiments are shown, but which may be embodied in many different forms and are not limited to the embodiments set forth herein.

[0034] For example, in this embodiment, the configuration, operation, etc. of the image analysis system will be described, but similar effects can be achieved by a device having similar functions, a method executed by the device, a computer program that causes a computer device to execute the method, etc. The program may be provided as a non-transitory computer-readable recording medium, or may be provided so as to be downloadable from an external server.

[0035] In the present invention, an "annotation" is a remark corresponding to a position in an image. In this embodiment, it particularly refers to a label indicating a classification of a pixel. Note that a label may be assigned to each pixel, or a label may be specified by indicating a certain range. In the latter case, the pixels present in the specified range may be processed assuming that the specified label is assigned.

[0036] A label indicates what is represented in the part of the image that corresponds to the pixel. For example, in an image that is composed of people, animals, and background, the label is information that indicates whether each pixel is part of a person, part of an animal, or part of the background. The type of label can be determined arbitrarily. For example, a label indicating normality or abnormality may be attached to detect defective parts of a product. Furthermore, there may be multiple types of labels indicating abnormality based on the above types.

[0037] In other examples, the label setting method can be used to classify cell types in a microscope image, classify ingredients in a lunch box, detect or classify organisms in an image (e.g., organisms in an aquarium), detect screw holes or gripping points in industrial products or parts, detect or count the number of parts in a tray, etc. Also, by adding a label indicating the location of movement between multiple images or frames of a video, it may be used to detect the location of movement.

[0038] <1. Functional configuration> Fig. 1 is a block diagram showing the configuration of an image analysis system according to this embodiment. As shown in Fig. 1, the image analysis system 0 includes a model generation device 1, an analysis device 2, and a storage device DB. The model generation device 1 and the storage device DB are connected by wire or wirelessly and configured to be able to communicate with each other. The storage device DB and the analysis device 2 are connected by wire or wirelessly and configured to be able to communicate with each other. There are no limitations on the type of communication protocol, the type of network, etc. applied between these devices.

[0039] The model generation device 1 generates a trained model by performing training based on a training image and stores the trained model in the storage device DB. The analysis device 2 analyzes an input image using the trained model stored in the storage device DB. Here, analysis of the input image refers to estimating labels for pixels of the input image.

[0040] <1.1. Functional configuration of model generating device 1> The model generating device 1 includes an acquisition unit 11, a feature amount calculation unit 12, an estimation unit 13, and a learning unit 14. This is a specific implementation of information processing by software using hardware.

[0041] The acquisition unit 11 acquires a training image. The training image is composed of labeled pixels and unlabeled pixels. In this embodiment, the acquisition unit 11 receives an image input from a user, and further receives designation of labels for pixels of the image, thereby generating a training image. Pixels with designated labels become labeled pixels. The generation of training images will be described later.

[0042] The feature calculation unit 12 calculates feature amounts for each pixel by using a multi-layered neural network model for the training image acquired by the acquisition unit 11. In this embodiment, the feature calculation unit 12 calculates image feature amounts by a convolutional neural network (CNN), and acquires feature amounts that correspond in particular to labeled pixels from the calculated image feature amounts.

[0043] The estimation unit 13 estimates labels for labeled pixels of the training image using a linear discriminator based on the features corresponding to the labeled pixels acquired by the feature calculation unit 12. As the linear discriminator, the estimation unit 13 can perform the estimation using, for example, a simple perceptron, a logistic regression, a linear support vector machine, Fisher's linear discriminator, or the like. The estimation here is a process of calculating the score of each label based on the feature amount of the pixel. The highest score can be estimated as the label of the pixel.

[0044] Here, in order to estimate the labels corresponding to the pixels of the image, the features extracted by the CNN (corresponding to the backbone) are integrated and made high-resolution to have the same resolution as the input image (corresponding to the neck). Note that the resolution does not necessarily have to be made high-resolution to the same extent as the input image. For example, features with an appropriately reduced resolution, such as 1 / 4, 1 / 8, 1 / 16, or 1 / 32 of the input image, may be used.

[0045] 3 is a schematic diagram showing the structure of a model in this embodiment. The model includes a feature extraction layer, a feature integration layer, and a discrimination layer. The feature amount calculation unit 12 in this embodiment acquires the output of the CNN in the feature extraction layer, and further executes a calculation to integrate and increase the resolution of this output in the feature integration layer, and passes the result to the estimation unit 13. The estimation unit 13 estimates a label corresponding to a pixel of the input image by inputting the result to a linear discriminator in the discrimination layer.

[0046] Any method can be used for the calculation of the feature integration layer. For example, a network with a HyperColumn structure may be used as the feature integration layer, similar to the technology described in Non-Patent Document 1. The network with a HyperColumn structure is one of the methods for integrating features extracted from multiple layers corresponding to the same pixel.

[0047] More specifically, a network with a HyperColumn structure can obtain the features of each pixel by enlarging features of different resolutions extracted by the backbone to a resolution similar to that of the input image or approximately one-quarter of the input image, and then combining them by concatenation.

[0048] The learning unit 14 performs learning by modifying the calculation parameters of the linear discriminator so that the estimation result by the estimation unit 13 matches the label of the learning image (the label of the target labeled pixel). The linear discriminator can perform learning at high speed because the number of parameters to be adjusted is significantly smaller than that of a non-linear discriminator.

[0049] <1.2. Functional configuration of the analysis device 2> The analysis device 2 includes an acquisition unit 21, a feature calculation unit 22, and an estimation unit 23. The analysis device 2 uses a trained model to analyze an unknown image. The acquisition unit 21 acquires an input image to be analyzed for pixel labeling. The feature calculation unit 22 and the estimation unit 23 function as a classification unit of the present invention.

[0050] Here, the process of estimating the labels of the pixels is similar to the learning process performed by the model generating device 1. That is, the feature calculation unit 22 and the estimation unit 23 execute the same process as the feature calculation unit 12 and the estimation unit 13, respectively, on pixels of an unknown input image using a trained model stored in the storage device DB. Therefore, detailed description of the processes will be omitted here.

[0051] The model generating device 1, the analysis device 2, and the storage device DB may be partially or entirely realized by a plurality of computers operating in cooperation with each other. These devices may also be realized by a single computer. In this case, the acquiring unit 11 and the acquiring unit 21, the feature calculating unit 12 and the feature calculating unit 22, and the estimating unit 13 and the estimating unit 23 may each serve as a single computer.

[0052] <2. Hardware configuration> An information processing device 9 (computer device) such as a general-purpose server or a personal computer can be used as the model generating device 1, the analysis device 2, and the storage device DB. In this embodiment, the model generating device 1 is the information processing device 9 in which a computer program (learning model generating program) that executes the learning model generating method is installed.

[0053] Fig. 2 is a hardware configuration diagram of the information processing device 9. As shown in Fig. 2, the information processing device 9 has a control unit 901, a storage unit 902, a communication unit 903, an input unit 904, and an output unit 905, which are used to perform the functions of each unit and each process.

[0054] The control unit 901 has a processor such as a CPU capable of executing an instruction set, and executes an OS, programs, and the like. The storage unit 902 includes a volatile memory such as a RAM capable of storing an instruction set, and a non-volatile recording medium such as an HDD or SSD capable of recording an OS, a determination program, and the like. The communication unit 903 has an interface for physically connecting to a network, and controls communications with the network NW to input and output information. The input unit 904 includes an operation input device capable of input processing, such as a touch panel or a keyboard, and an audio input device capable of audio input, such as a microphone. The output unit 905 includes a display device capable of display processing, such as a display, and an audio output device, such as a speaker.

[0055] <3. Learning image generation> 4 is a flowchart showing the process of generating training images. As described above, in this embodiment, the acquisition unit 11 generates training images by receiving an input of an image from a user and a designation of labels for pixels of the image.

[0056] In this embodiment, it is assumed that a user who intends to use the analysis device 2 prepares training images by himself / herself and generates a model that has been independently trained using the model generation device 1. Therefore, as described below, it is preferable to be able to generate training images using a UI that allows a user to assign labels to images through intuitive operations.

[0057] First, in step S11, the user inputs an image and the image is acquired by the acquisition unit 11. Next, in step S12, the acquisition unit 11 accepts designation of labels for pixels of the image from the user. Fig. 5 is an example of a UI for training image generation. The model generation device 1 displays a screen such as that shown in Fig. 5 on the display and accepts various inputs from the user. The training image display unit LI displays images acquired by the acquisition unit 11. In step S12, the user selects a label to be assigned to the pixel by inputting via the label button LB.

[0058] Next, in step S13, the user selects a position (pixel) in the learning image display section LI, and the acquisition section 11 assigns the label selected in step S12 to the pixel selected in step S13 (step S14). The label assigned to the pixel is displayed in the learning image display section LI, superimposed on the image. The order of steps S12 and S13 may be reversed. Furthermore, the selection of pixels in step S13 may be such that a range of positions on the image is specified and all pixels within the range are selected at once.

[0059] After steps S11 to S13 are repeated for an arbitrary pixel, when the user selects the “Learning” button, the feature calculation unit 12, the estimation unit 13, and the learning unit 14 execute the learning process described below based on a learning image including labeled pixels according to the input.

[0060] When the user selects the "Check learning result" button, the model that has undergone learning processing based on the learning image analyzes the image acquired by the acquisition unit 11. Then, in the learning image display area LI, the analysis results by the model, i.e., the estimation results of the labels of each pixel, are further superimposed and displayed. The user can check the results and add labels to perform additional learning. Furthermore, if the training images are properly labeled, the user can finish training the training images and complete model creation or input other training images.

[0061] <4. Learning model generation method> Next, the learning process by the model generating device 1 will be described with reference to Fig. 6. Fig. 6 shows the procedure for learning one training image. The process in Fig. 6 can be performed for each training image. As described above, in this embodiment, it is assumed that the pixels of the training image are labeled by an operation of a general user, and only some points (pixels) in the training image are labeled. The number of labeled pixels is not limited to this, but may be, for example, half or less of the number of pixels of the training image, one tenth or less of the number of pixels of the training image, one hundredth or less of the number of pixels of the training image, one thousandth or less of the number of pixels of the training image, or one ten thousandth or less of the number of pixels of the training image. Alternatively, the number of labeled pixels is 100 or less, 75 or less, 50 or less, 30 or less, or 10 or less. It is sufficient that at least one labeled pixel is present in each training image.

[0062] There is also no particular limit to the number of training images used in training, but in this embodiment, training is performed using training images that are 200 or less, 100 or less, 75 or less, 50 or less, 30 or less, 20 or less, or 10 or less. The number of training images may be at least 1.

[0063] Here, it is not necessary for all types of labeled pixels to be present in each training image, but in this embodiment, the training images and labeled pixels are selected such that, when all training images are taken into consideration, there is at least one labeled pixel for each type of label. That is, in this embodiment, for example, when generating a model in which three types of labels, A, B, and C, are assigned using two training images, at least one labeled pixel assigned with each of the labels A, B, and C is included in either of the two training images.

[0064] In the learning process, first, in step S21, the acquisition unit 11 acquires a learning image composed of labeled pixels and unlabeled pixels by the generation process described above. Next, in the feature amount calculation process in step S22, the feature amount calculation unit 12 calculates the feature amount of each pixel by multi-layer CNN, generates an integrated feature for each pixel by integrating the features by the feature integration layer, acquires an integrated feature corresponding to the labeled pixel, and passes it to the estimation unit 13. Note that, when generating a model that estimates a label for each certain range including multiple pixels, rather than for each pixel, the integrated feature may also be generated for that range.

[0065] Next, in the estimation step of step S23, the estimation unit 13 inputs the integrated feature received from the feature amount calculation unit 12 to a linear discriminator. The estimation unit 13 estimates the label of the pixel in question according to the output of the linear discriminator.

[0066] In the learning process of step S24, the learning unit 14 corrects the calculation parameters of the linear discriminator so that the label of the target pixel in the learning image matches the estimation result by the estimation unit 13. Then, in step S25, it is determined whether there are any unlearned labeled pixels, and steps S22 to S24 are repeated for one learning image until there are no unlearned labeled pixels. In this way, learning is performed for all labeled pixels in the target learning image.

[0067] 5. Image analysis method Next, the image analysis process by the analysis device 2 will be described with reference to Fig. 7. The analysis device 2 estimates the labels of pixels for an unknown image (an image other than the learning image) using the trained model created by the above-mentioned procedure. The label estimation is performed by executing the same processes as steps S22 and S23 of the learning process described above for all pixels of the input image using the trained model. This will be described in detail below.

[0068] First, in step S31, the acquisition unit 21 acquires an input image to be subjected to label estimation. Next, in step S32, the feature amount calculation unit 22 calculates the feature amount of pixels using the trained model stored in the storage device DB. The feature amount is calculated by the feature extraction layer configured by CNN as described above, and is integrated by the feature integration layer. It is not necessary to integrate the feature amounts so as to correspond to each pixel, but for example, the feature amounts may be integrated so as to correspond to a certain range in the image to calculate an integrated feature, and the integrated feature may be regarded as the feature amount corresponding to each of the multiple pixels. In other words, the resolution does not need to be the same as that of the input image, and the integrated feature may have an appropriately reduced resolution, such as 1 / 4, 1 / 8, 1 / 16, or 1 / 32 of the input image, as described above.

[0069] In step S33, the estimation unit 23 estimates the label of the pixel using the trained model. Specifically, when the integrated feature is input to the discrimination layer of the trained model, the score of each label corresponding to the integrated feature is output. The estimation unit 23 estimates the label with the highest score as the label of the pixel corresponding to the integrated feature.

[0070] Then, in step S34, the presence or absence of pixels (unclassified pixels) for which the above-mentioned estimation process has not been performed is confirmed, and if an unclassified pixel exists, the target pixel is changed and steps S23 to S33 are repeated, thereby estimating the labels for all pixels.

[0071] In this way, according to the configuration of the present invention, by performing learning using a learning image in which only some pixels are labeled, a trained model for estimating pixel labels for an unknown image can be generated. Also, using the trained model thus generated, an image analysis can be performed to estimate pixel labels. The present invention can be applied to, for example, an inspection device for finding product defects based on product images.

[0072] <6. Remarkable Effects of the Present Invention> The remarkable effects of the present invention will now be described. First, the differences between the conventional technology and the present invention will be described. As described in the description of the background technology, in the technology described in Non-Patent Document 1, when learning a model, labeling learning is performed only for some pixels, thereby increasing the batch size, which was previously small due to memory constraints, and improving accuracy.

[0073] The technology described in Non-Patent Document 1 employs a multilayer perceptron (MLP) and a stochastic gradient descent (SGD) in the head part. In contrast, the present invention differs from the technology described in Non-Patent Document 1 in that a linear classifier is used in the head part. In addition, the technology described in Non-Patent Document 1 is premised on performing learning using images in which all pixels are annotated. In contrast, the present invention differs from the technology described in Non-Patent Document 1 in that learning is performed using training images composed of labeled pixels and unlabeled pixels.

[0074] <6.1 Estimation accuracy comparison experiment> As already mentioned, machine learning generally requires learning from a large number of images, and the technology described in Non-Patent Document 1 also learns from several thousand images. However, when attempting to generate a model specialized for a specific purpose, it is often difficult to prepare such a large number of training images. We therefore conducted a comparative experiment on the estimation accuracy using a model generated under conditions where there were an extremely small number of training images.

[0075] (1) Model configuration A comparative experiment was conducted using model A (with the same configuration as the model described in non-patent document 1) that employs a multilayer perceptron (MLP), a non-linear discriminator, in the discrimination layer, and model B (with the same configuration as the model of the present invention) that employs a linear discriminator in the discrimination layer. Both models use a convolutional neural network for the feature extraction layer, and the configuration of the feature integration layer is the same. Both models are common except for the configuration of the discrimination layer.

[0076] (2) Learning images The image shown in Fig. 8(a) was used as the learning image. The pixels at the positions indicated by dots in Fig. 8(a) were labeled. Specifically, a total of eight labeled pixels were set, one on a circular object, one on a square object, one on a linear object sloping upward to the right, one on a linear object sloping downward to the right, and four on the background. As a learning image, only one image shown in FIG. 8(a) was used.

[0077] (3) Results Fig. 8(b) is a diagram showing the estimation results by model A when the same image as the learning image is used as the input image. Fig. 8(c) is a diagram showing the estimation results by model B under similar conditions. In both cases, the pixels of the input image are labeled in five ways: circular objects, square objects, linear objects sloping upward to the right, linear objects sloping downward to the right, and the background.

[0078] In the output of model A, it can be seen that the background and each object are not sufficiently distinguished, and the shape of the object is not well represented. On the other hand, in the output of model B, the range of each object is correctly recognized, and the label of the input image is accurately estimated even for unlabeled pixels.

[0079] When calculating the accuracy rate for each label, the results were as follows: background Model A: 97.58%, Model B: 99.02% Rectangle Object Model A: 33.90%, Model B: 98.75% Circular objects Model A: 25.39%, Model B: 98.73% Upward-sloping linear object Model A: 62.00%, Model B: 74.39% Sloping linear object Model A: 12.51%, Model B: 82.96%

[0080] Thus, under conditions where there were very few training images or labeled pixels, model A, which used a nonlinear classifier, was unable to estimate labels with an accuracy that was practical. On the other hand, model B, which used a linear classifier, was shown to be able to estimate labels with high accuracy, even under conditions where there were very few training images or labeled pixels. In this way, the use of a linear discriminator in the discrimination layer significantly improves accuracy under specific conditions, which is a remarkable effect that could not be predicted from conventional techniques.

[0081] <6.2 Computation amount> (1) Comparison with PixelNet The present invention has an advantageous effect compared to the conventional technology in terms of the amount of calculation. For example, in the technology described in Non-Patent Document 1, the backbone and neck structures are roughly similar to those of the present invention, but a multilayer perceptron is used in the head part. The head of the technology described in Non-Patent Document 1 uses the features received from the neck as input and calculates the weights for the neurons in each layer (4096).

[0082] In the technology described in Non-Patent Document 1, since there are two hidden layers, if the number of classes to be classified is k, the head part has parameters of the number of dimensions of the features input to the head × 4096 + 4096 × 4096 + 4096 × k. Since calculations are performed taking into account a large number of parameters, the amount of calculation increases and the calculation speed decreases.

[0083] On the other hand, if a linear discriminator is used as in the present invention, the number of parameters is reduced. Specifically, since the score is calculated linearly for each class to be classified, the number of parameters is the number of dimensions of the features input to the head × k. This reduces the amount of calculation, reducing the processing load and improving the calculation speed.

[0084] (2) Speed ​​comparison experiment with a general model Furthermore, a comparative experiment was conducted on the calculation speed between a general image labeling (semantic segmentation) model and the model of the present invention. The DeepLabV3 model was used for comparison. The DeepLabV3 model is a model that uses atlas convolution (asymmetric convolution) to avoid a significant loss of resolution in the backbone (Reference: “DeepLabv3”, [online], search date October 11, 2023, Internet<URL:https: / / paperswithcode.com / method / deeplabv3> ). As the model of the present invention, a model employing a linear discriminator in the head was used.

[0085] The conditions are as follows: Both models use the same GPU (GeForce RTX 3080) Input image size: 288 x 288 pixels Backbone of both models: ResNet50 (a convolutional neural network with a depth of 50 layers) Number of trials: Labeling of images is performed 10 times and the average time required to label one image is calculated.

[0086] As a result of the above, the calculation time of the DeepLabV3 model was 73 ms on average, while the calculation time of the model of the present invention was 11 ms. In this way, it was confirmed by the experiment that the present invention can perform calculations faster than general models. [Explanation of symbols]

[0087] 0: Image analysis system 1: Model generation device 2: Analyzer 9: Information processing equipment 11: Acquisition part 12: Feature calculation unit 13: Estimation part 14: Learning Department 21: Acquisition section 22: Feature calculation unit 23:Estimation part 901: Control unit 902: Storage section 903: Communications Department 904: Input section 905: Output section LB: Label button LI: Learning image display section NW: Network

Claims

1. 1. A model generating apparatus for generating a model for estimating labels of pixels of an image, comprising: The apparatus includes a feature calculation unit, an estimation unit, and a learning unit, the feature calculation unit inputs a learning image composed of labeled pixels and unlabeled pixels into a multi-layered convolutional neural network model, and acquires, as an output of the convolutional neural network model, feature amounts corresponding to the labeled pixels based on the labeled pixels and the unlabeled pixels; the estimation unit performs an input based on the feature amount to a linear discriminator, and estimates a label of the labeled pixel according to an output of the linear discriminator; the learning unit generates the model in which the convolutional neural network model and the linear discriminator are combined by modifying an operation parameter of the linear discriminator so that an estimation result by the estimation unit matches a label of the labeled pixel in the learning image; A model generating device, wherein pixels of the training image other than the labeled pixels are unlabeled pixels.

2. There are multiple types of the label, the feature calculation unit acquires the feature corresponding to at least one of the labeled pixels for each type of label, based on the labeled pixels and unlabeled pixels, by using the convolutional neural network model; the estimation unit estimates a label of the labeled pixel according to an output of the linear discriminator by inputting, to the linear discriminator, an input based on the feature for each of the labeled pixels for which the feature calculation unit has acquired the feature; 2. The model generation device according to claim 1, wherein the learning unit generates the model in which the convolutional neural network model and the linear discriminator are combined by repeatedly correcting an operation parameter of the linear discriminator so that, for each of the labeled pixels whose labels have been estimated by the estimation unit, an estimation result by the estimation unit matches a label of the labeled pixel in the training image.

3. The model generating device according to claim 1 , wherein the number of the labeled pixels for each training image is equal to or less than half the number of pixels of the training image.

4. The model generating device according to claim 1 , wherein the number of the labeled pixels for each training image is equal to or less than 1 / 100 of the number of pixels of the training image.

5. The model generating device according to claim 1 , wherein the number of training images used in training is 100 or less.

6. 2. The model generating device according to claim 1, wherein the learning unit corrects an operation parameter of the linear discriminator so that an estimation result by the estimation unit matches a label of the labeled pixel in the training image, without changing a parameter of the convolutional neural network model.

7. The model generating device according to claim 1 , wherein the labels include a label for classifying the presence or absence of an abnormality or a label for classifying a type of an abnormality.

8. A model generation program that causes a computer to function as an apparatus for generating a model for estimating labels of pixels in an image, the program comprising: A computer is caused to function as a feature calculation unit, an estimation unit, and a learning unit; the feature calculation unit inputs a learning image composed of labeled pixels and unlabeled pixels into a multi-layered convolutional neural network model, and obtains, as an output of the convolutional neural network model, feature amounts corresponding to the labeled pixels based on the labeled pixels and the unlabeled pixels; the estimation unit performs an input based on the feature amount to a linear discriminator, and estimates a label of the labeled pixel according to an output of the linear discriminator; the learning unit performs learning by modifying an operation parameter of the linear discriminator so that an estimation result by the estimation unit matches a label of the labeled pixel in the learning image; Pixels in the training images other than the labeled pixels are unlabeled pixels.

9. 1. A method for generating a model for estimating labels of pixels of an image, comprising the steps of: a feature calculation step of inputting a learning image composed of labeled pixels and unlabeled pixels into a multi-layered convolutional neural network model, and acquiring, as an output of the convolutional neural network model, feature amounts corresponding to the labeled pixels based on the labeled pixels and the unlabeled pixels; an estimation step of inputting the feature amount to a linear discriminator and estimating the label of the labeled pixel according to an output of the linear discriminator; a learning process of generating the model in which the convolutional neural network model and the linear discriminator are combined by modifying an operation parameter of the linear discriminator so that an estimation result in the estimation process matches a label of the labeled pixel in the training image, A model generating method in which pixels in the training images other than the labeled pixels are unlabeled pixels.

10. A trained model for causing a computer to function, which has been trained to estimate labels of pixels in an image using an image as input, The method includes: a multi-layered convolutional neural network; and a linear discriminator that is coupled to receive information based on an output of the convolutional neural network and output an estimation result of a label of a pixel; The convolutional neural network outputs, based on a plurality of pixels of an input image, feature amounts corresponding to a portion of the plurality of pixels; the linear discriminator has been trained to have a function of outputting an estimation result of a pixel label from an input based on an output of the convolutional neural network by modifying an operation parameter so that the label of the labeled pixel coincides with an estimation result for the labeled pixel for a training image composed of labeled pixels and unlabeled pixels; Pixels of the training image other than the labeled pixels are unlabeled pixels; A trained model for causing a computer to function so as to perform calculations based on calculation parameters of the convolutional neural network and a linear discriminator for each pixel of an input image input to the convolutional neural network, and output an estimated label result for each pixel of the input image.

11. 1. An image analysis apparatus for estimating labels of pixels of an image, comprising: An acquisition unit and a classification unit, The acquisition unit acquires an input image for estimating pixel labels; The classification unit inputs the input image to a trained model, and outputs an estimation result of a label for each pixel of the input image; The trained model is The method includes: a multi-layered convolutional neural network; and a linear discriminator that is coupled to receive information based on an output of the convolutional neural network and output an estimation result of a label of a pixel; The convolutional neural network outputs, based on a plurality of pixels of an input image, feature amounts corresponding to a portion of the plurality of pixels; the linear discriminator has been trained to have a function of outputting an estimation result of a pixel label from an input based on an output of the convolutional neural network by modifying an operation parameter so that the label of the labeled pixel coincides with an estimation result for the labeled pixel for a training image composed of labeled pixels and unlabeled pixels; Pixels of the training image other than the labeled pixels are unlabeled pixels.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, and program

    JP2019192009A