Image processing apparatus and image processing method
The image processing apparatus addresses overdetection in object detection models by allowing users to visually correct model outputs, enhancing accuracy through user-specified annotations, thus improving model training for non-experts.
Patent Information
- Application Number
- JP2023219709
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-07-08
AI Technical Summary
Existing object detection models trained by machine learning suffer from 'overdetection', where non-target objects are incorrectly identified, reducing accuracy, and it is challenging for non-experts to address this issue through model design or data augmentation.
An image processing apparatus that allows users to visually emphasize potential overdetections and omissions in object detection, enabling them to specify first and second annotations to train the model to avoid overdetection and ensure accurate object identification.
The apparatus enables non-experts to effectively mitigate overdetection and detection omissions by allowing intuitive user interaction for model training, improving object detection accuracy without requiring extensive expertise or additional data.
Smart Images

Figure 2025102344000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an object inspection technique for detecting an object included in an inspection target image (or an inspection target object imaged in the inspection target image; hereinafter the same).
Background Art
[0002] Conventionally, it has been known to detect an object included in an inspection target image obtained by imaging an inspection target object. The object here is assumed to be a wide variety of things, such as an abnormality (defect, scratch, etc.) or an object included in the inspection target image. In recent years, for example, it has been studied to detect an object included in an inspection target image using a learned model obtained by training a model expressed by a deep neural network by machine learning.
[0003] Patent Document 1 discloses an appearance inspection device for detecting an abnormality included in an inspection target image. The prior art disclosed in Patent Document 1 inputs an inspection target image into a convolutional neural network (learned neural network), outputs a feature map from each layer of the convolutional neural network (learned neural network), and inputs each of the feature maps from each layer into a model by an autoencoder corresponding to each layer. The model by the autoencoder corresponding to each layer calculates, as a loss function, the square of the norm of the difference in feature amounts (the sum in the channel direction) at each spatial position (spatial coordinates) between the feature map that is the input to the model and the feature map that is the output from the autoencoder, and treats the value of the loss function as the degree of abnormality. An abnormality degree map is formed by a set of abnormality degrees corresponding to each spatial position (spatial coordinates). The prior art disclosed in Patent Document 1 determines whether or not the inspection target image includes an abnormality using the abnormality degree map. Patent Document 2 discloses an appearance inspection apparatus that determines the quality of a work image obtained by imaging a work to be inspected. The prior art disclosed in Patent Document 2 inputs the work image into a machine learning network to determine the quality of the work image. In order to adjust the parameters of the machine learning network, machine learning is executed by inputting good product images and defective product images into the machine learning network. A label indicating that it is a good product image is assigned to the good product image. For the defective product image, there are cases where a label indicating that it is a defective product image is assigned, and cases where an annotation specifying a defective part is also assigned after a label indicating that it is a defective product image is assigned.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] When detecting an object included in an inspection target image using a model trained by machine learning, "overdetection" may occur, where an object that is not originally detected as an object is detected as an object. Various reasons are assumed for the occurrence of overdetection. For example, when there is bias in the training data (training images), a part of an object that is not originally detected as an object may be in a situation not learned by the model, leading to overdetection. Also, for example, when small dust or the like is reflected in the inspection target image, the object may be "overdetected". Since the "overdetection" of an object may reduce the accuracy of object detection using a model, it may be considered to train the model so that the object is not "overdetected" as much as possible. Here, several methods can be considered as techniques for training the model by machine learning so that overdetection does not occur. One method of preventing overdetection is to devise the design of the model, such as the definition of the loss function. However, in order to make such a devise, specialized knowledge of machine learning and the inspection object is often required. Therefore, it is often difficult for the person in charge at the site who actually operates the system for detecting an object included in the inspection target image to execute such a devise. Furthermore, depending on the content of the devise, the accuracy of the estimation of the presence or absence of an object in a part (area) where the estimation could be appropriately performed by the model before the devise may be adversely affected. Another method of preventing overdetection is to add learning data used for machine learning. However, in order to prevent overdetection, it is often difficult for the person in charge at the site to determine what kind of learning data should be added and in what quantity. Also, when the object is abnormal, it may be difficult to increase the number of learning images including the abnormality. In addition, in the case of an event with a low occurrence frequency, which is an event in the background part of the inspection target image and for which it is desired not to be overdetection as an object, it is highly likely that it cannot be improved by adding learning data. Patent Document 1 shows that, as a countermeasure against overdetection, in each of the hierarchical feature maps or the composite feature maps, those in which overdetection occurs should be used as little as possible. However, this countermeasure does not eliminate the overdetection itself by the model. Also, Patent Document 1 mentions another countermeasure against overdetection regarding the addition of learning images. However, as already pointed out, there are often difficulties in adding learning data. Patent Document 2 shows that labels and annotations are given to good product images and defective product images that are learning data. However, Patent Document 2 only shows that when it is found that there is a "detection omission" where a defective product image cannot be detected as a defective product image, an annotation specifying a defective part is given to the defective product image that has become a detection omission. That is, the prior art disclosed in Patent Document 2 does not take measures to prevent over-detection when over-detection is found.
[0006] Based on the above, when detecting an object included in an inspection target image obtained by imaging an inspection target using a model trained by machine learning, one of the purposes of the present disclosure is to prevent (or mitigate the degree of) "over-detection" in which a target that is not originally detected as an object is detected as an object. One of the purposes of the present disclosure is to realize training of a machine learning model for a person who does not have sufficient expertise in machine learning to prevent (or mitigate the degree of) "over-detection" by the above-described model. When it is found that the above-described model causes "over-detection", one of the purposes of the present disclosure is to realize training of a machine learning model to prevent (or mitigate the degree of) "over-detection" by the model.
Means for Solving the Problem
[0007] In order to achieve at least one of the above objects, the features that the present disclosure may include are as follows, for example. One aspect of the present disclosure is an image processing apparatus including a processor that detects an object included in an inspection target image in which an inspection target is imaged, using a machine learning model. The processor trains the machine learning model using a training image in which the inspection target is imaged, outputs a portion estimated to be the object in the inspection target image using the trained machine learning model, causes a display unit to display a display image in which a portion estimated to be the object based on the inspection target image is emphasized, receives, via the display unit, a designation of a first annotation indicating that a portion emphasized as the object in the displayed display image is a non-detection target that is not a target to be detected as the object, and trains the machine learning model so as not to detect the portion designated by the first annotation in the inspection target image as the object, using the inspection target image and the first annotation.
Advantages of the Invention
[0008] According to the present disclosure, when detecting an object included in an inspection target image using a machine learning model, it is possible to prevent (or mitigate the degree of) "overdetection" in which a target that is not originally detected as an object is detected as an object. The present disclosure can receive a designation of a first annotation indicating that a portion emphasized and displayed as an object is not a target to be detected as an object, after displaying a display image in which the portion estimated to be the object is emphasized. Therefore, the present disclosure can realize training of a machine learning model for a person who does not have sufficient expertise in machine learning to prevent (or mitigate the degree of) "overdetection" in the above-described model.
[0009] An image processing method and an image processing program that achieve the same processing as that realized by the above-described image processing apparatus can also obtain the same operational effects as the above-described image processing apparatus. Furthermore, in the case of a program, costs are often reduced. In a program, design changes related to processing are also easily made.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Mode for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Note that the embodiments described below do not limit the disclosure according to the claims, and not all of the elements and combinations thereof described in the embodiments are essential for the solution means of the present disclosure. The following description and drawings are examples for explaining the present disclosure, and for the sake of clarity of explanation, omissions and simplifications are made as appropriate. Each of the system, apparatus, model or part of the present disclosure may be integrated into one hardware-wise, or may be divided into a plurality of parts and these parts may cooperate to play a role.
[0012] First, the functional configuration of the embodiments of the present disclosure will be described with reference to FIGS. 1, 2 and 3. FIG. 1 shows the functional configuration of the embodiments of the present disclosure. In FIG. 1, rectangles indicate functional parts, and parallelograms that are not rectangles indicate information, data, etc. handled in the functional parts. FIG. 2 shows an example of the internal configuration of a model in the embodiments of the present disclosure. Note that the model in the embodiments of the present disclosure may be any model that can obtain object estimation information regarding objects that may be included in the inspection target image. Therefore, the embodiments of the present disclosure are not limited to models having the internal configuration shown in FIG. 2. FIG. 3 shows the training of a model using first and second annotations. As will be described in detail later, the first annotation indicates that the portion designated by the first annotation is not an object to be detected in the inspection target image. The second annotation indicates that the portion designated by the second annotation is an object to be detected in the inspection target image.
[0013] As shown in FIG. 1, the image processing apparatus 100 includes, as functional units, a model realization unit 102, a display output control unit 105, an annotation designation reception unit 106, and a learning control unit 107. Here, the model realization unit 102 realizes a model 122. The learning control unit 107 controls the training of the model. The image processing apparatus 100 further includes a learned neural network 101 as a functional unit. The annotation designation reception unit 106 includes, as a functional unit, a first annotation designation reception unit 161. Also, the annotation designation reception unit 106 includes, as a functional unit, a second annotation designation reception unit 162. The functional units of the image processing apparatus 100 are realized by a processor (including, for example, a CPU or a GPU) as the information processing apparatus 501 described later. The image processing apparatus 100 may handle, as information, data, etc. handled by the functional units, an inspection target image, an input to the model, object estimation information, an image with object highlighting display output, etc. Further, the image processing apparatus 100 may handle, as information, data, etc. handled by the functional units, one or more of a feature map and a frequency map (for example, an abnormality degree map; the same applies hereinafter). What the first annotation designation reception unit 161 receives may be the first annotation. Also, each strength 173 of the first annotation may be included in what the first annotation designation reception unit 161 receives. What the second annotation designation reception unit 162 receives may be the second annotation.
[0014] The inspection target image is an image obtained by imaging an inspection target object. For example, as shown in FIG. 4 described later, an inspection target image obtained by imaging an inspection target object with an imaging device 401 may be provided to the image processing device 100. Alternatively, a set of one or more inspection target images may be collectively provided to the image processing device 100, and the set of inspection target images may be stored in the image processing device 100 in a manner shown in 524 of FIG. 5 described later.
[0015] The learned neural network 101 inputs an inspection target image and outputs a feature map. For example, the feature map is an array of feature amounts formed by arranging the feature amounts output from an intermediate layer or an output layer of the learned neural network 101 along the spatial dimension (and the channel dimension if it exists). The input to the model may be the image itself or a feature map obtained from the image, depending on the specification of the model 122. In the following, there are some places where, for the sake of clarity, examples of images are described, but in these description places, a feature map may be used instead of the image. Also, in this specification, the feature map includes not only the one output from the learned neural network 101 but also an image obtained by filtering an inspection target image or a verification image with an edge detection filter or the like.
[0016] The model implementation unit 102 implements the model 122. For example, the model implementation unit 102 implements the model using a model program and model parameters corresponding to the model 122 to be implemented. Alternatively, the model 122 may be implemented on hardware specialized for executing a machine learning model (e.g., ASIC or FPGA). The model 122 outputs object estimation information regarding an object included in the inspection target image corresponding to the input to the model. The object estimation information output by the model 122 may be information regarding the objects included in the inspection target image corresponding to the input to the model. For example, the object estimation information is numerical information indicating the degree to which the presence of an object in the inspection target image is estimated. Hereinafter, this numerical information is referred to as the "object estimation degree". Further, when the object is an abnormality included in the inspection target image, the above numerical information is referred to as the "abnormality degree". FIG. 2 shows an example of the internal configuration of the model 122 in an embodiment of the present disclosure. The model 122 includes, as functional units, an in-model neural network 221, a loss function calculation unit 222, and an object estimation information generation unit 223. The model handles a neural network output and a loss function value (Loss) as information, data, etc. handled by the functional units.
[0017] The in-model neural network 221 is trained by machine learning using a learning image or a feature map obtained from the learning image. Through machine learning using the learning image or the feature map, the in-model neural network 221 is trained so that for pixels that are not detected as objects (parts), both the pixel value or feature amount that is the input data to the neural network and the pixel value or feature amount that is the output from the neural network substantially match (the difference between the two is close to zero). That is, the in-model neural network 221 is trained to output an image or a feature map in which no object exists as the neural network output by learning an image or a feature map in which no object exists. An example of the in-model neural network 221 as described above is an autoencoder.
[0018] The loss function calculation unit 222 calculates a loss function for the pixels constituting the image or the feature map using the input value and the output value to the model. For example, if the input to the model is a feature map, and the feature map is a three-dimensional array of input feature quantities OFM_x,y,c with the coordinates in the two-dimensional spatial dimension (two-dimensional spatial coordinates) x, y and the coordinates in the channel dimension (channel coordinates) c as indices, and the neural network output is a three-dimensional array of output feature quantities AEOFM_x,y,c, assuming the number of elements in the channel direction is N_c, the loss function Loss_x,y at the two-dimensional spatial coordinates x, y on the feature map may be calculated using the following formula. Note that the sigma (Σ) in the following formula is the sum with respect to the channel coordinate c. Loss_x,y=(1 / N_c)*Σ(OFM_x,y,c - AEOFM_x,y,c) 2 That is, the loss function may be calculated based on the square of the difference between the neural network input and the neural network output. If the neural network 221 in the model is trained as already pointed out, for the two-dimensional spatial coordinates x, y where it is estimated that there is no object, it can be expected that the loss function Loss_x,y is close to zero. Note that the calculation method of the loss function Loss_x,y is not limited to the above example. Alternatively, the loss function calculation unit 222 may obtain the loss function Loss_X,Y at the two-dimensional spatial coordinates X, Y on the inspection target image.
[0019] The object estimation information generation unit 223 obtains the object estimation degree at the two-dimensional spatial coordinates X, Y based on the loss function Loss_X,Y at the two-dimensional spatial coordinates X, Y on the inspection target image or the loss function Loss_x,y at the two-dimensional spatial coordinates x, y on the feature map. When the input to the model is a feature map and it is desired to obtain the object estimation information for the inspection target image, the object estimation information generation unit 223 may perform an operation (upsampling-like operation) for converting the loss function Loss_x,y at the two-dimensional spatial coordinates x, y on the feature map into the object estimation degree at the two-dimensional spatial coordinates X, Y on the inspection target image. From each of the multiple layers of the learned neural network 101, in a manner in which each of the feature maps is output, the model realization unit 102 realizes a model 122 corresponding to each of the above-described layers. When a neural network output and a loss function value (Loss) corresponding to each of the above-described layers are obtained, the object estimation information generation unit 223 calculates an object estimation degree in two-dimensional space coordinates X and Y using the loss function value (Loss) corresponding to each of the above-described layers.
[0020] The display output control unit 105 controls to perform a display based on the object estimation information. The display output control unit 105 may control to display or output one or both of an image with object highlighted display output or the like or a frequency map. Note that the object to be controlled to display or output an image with object highlighted display output or the like or a frequency map may be the display or output device 507 in FIG. 5 described later. Alternatively, the object to be controlled to display or output an image with object highlighted display output or the like or a frequency map may be any display device or output device capable of communicating with the image processing apparatus 100.
[0021] An image with object highlighted display output or the like is an image in any of the inspection target image, the feature map, the processed inspection target image, and the processed feature map, and is an image in a mode of highlighting and displaying or outputting a portion estimated to be an object. Here, examples of the processed image or feature map include a binarized original image or feature map, a simplified original image or feature map, and a reduced-resolution original image or feature map. The display output control unit 105 determines a portion estimated to be an object based on the object estimation information. For example, when the object estimation information is the object estimation degree in two-dimensional space coordinates X, Y, the display output control unit 105 estimates that an object exists at the two-dimensional space coordinates X, Y when the object estimation degree in the two-dimensional space coordinates X, Y is equal to or greater than a predetermined value (or greater than a predetermined value). Alternatively, the display output control unit 105 calculates the total value (or average value) of the object estimation degrees for each fixed set of two-dimensional space coordinates X, Y, and estimates that an object exists at the position (portion) of the fixed set of the two-dimensional space coordinates X, Y when the calculated total value (or average value) is equal to or greater than a predetermined value (or greater than a predetermined value). As long as a person using the image processing apparatus 100 (hereinafter also referred to as a user) can visually recognize a portion estimated to be an object, the mode of highlighting and displaying or outputting the portion estimated to be an object may be arbitrary. For example, a mode of coloring a portion estimated to be an object in what is to be displayed or output (any one of an inspection target image, a feature map, an image obtained by processing the inspection target image, and an image obtained by processing the feature map) with a specific color may be used. Here, the mode of coloring with a specific color may be to fill it with the specific color. Further, the mode of coloring with a specific color may be to color it thinly to such an extent that the image or feature map before coloring can still be seen. An image with object-emphasized display output or the like can clearly show which location in the image or feature map the user should consider when making a designation of the following first annotation or second annotation.
[0022] The frequency map is a visible map that reflects the object estimation information. When the object indicates an abnormality, the frequency map may be called an abnormality degree map. When the object estimation information is a set of object estimation degrees in two-dimensional spatial coordinates X and Y, the display or output mode of the two-dimensional spatial coordinates X and Y in the heat map reflects the values of the object estimation degrees in the two-dimensional spatial coordinates X and Y. For example, after setting one or more thresholds for the object estimation degree, the display in the heat map may be determined according to which numerical range separated by the threshold the object estimation degree is included in. By displaying the heat map, the user can visually grasp the numerical level of the object estimation degree and the distribution of the parts where the absolute value of the object estimation degree is large.
[0023] The annotation specification reception unit 106 receives an instruction for annotation regarding the inspection target image while utilizing the image with object highlighting display output or the frequency map that is displayed or output under the control of the display output control unit 105. The annotation specification reception unit 106 includes a first annotation specification reception unit 161 and a second annotation specification reception unit 162 as functional units.
[0024] The first annotation specification reception unit 161 receives a specification of a first annotation regarding the inspection target image that is the target of display by the image with object highlighting display output or the frequency map. The first annotation indicates that the part specified by the first annotation is not a target to be detected as an object.
[0025] An example of the specification of the first annotation will be described with reference to FIG. 3. Assume that the image with object highlighting display output based on the object estimation information obtained by inputting the inspection target image shown in the upper left of FIG. 3 into the model 122 is the one shown in the upper right of FIG. 3. In the inspection target image etc. in the upper left of FIG. 3, a certain part is regarded as the target 312 to be detected as an object by the user, and the other parts are not regarded as the targets to be detected as objects. On the other hand, in the image etc. with object emphasis display output in the upper right of FIG. 3, a part of the part that is not originally regarded as the target to be detected as an object is regarded as the estimated part 321 as an object. The estimated part 321 as an object is either emphasized and displayed or displayed according to a high object estimation degree. In this way, the fact that the part that is not originally regarded as the target to be detected as an object is regarded as the estimated part 321 as an object means that the part 321 is a part where the object is "overdetected". The user who has viewed the image etc. with object emphasis display output in the upper right of FIG. 3 makes a first annotation designation for the inspection target image as shown on the left side of the central part of FIG. 3. Specifically, in the image etc. with object emphasis display output in the upper right of FIG. 3, for the part 321 that is estimated as an object and is "overdetected", the user makes a first annotation designation. The first annotation designation reception unit 161 receives the first annotation designation for the inspection target image on the left side of the central part of FIG. 3. As a result of this reception of the first annotation designation, the part where the first annotation is designated for the inspection target image, that is, the first annotation designated part 371 is set. In this way, by using the first annotation designation reception unit 161, the user who has viewed the image etc. with object emphasis display output can make a first annotation designation indicating that the part that is regarded as the estimated part 321 as an object but is a part that is not regarded as the target to be detected as an object is a part where the object is "overdetected". In addition, in the designation of the first annotation, the strength 173 may also be designated.
[0026] The second annotation specifying reception unit 162 receives the specification of the second annotation regarding the inspection target image that is the target of display or output by an image with object highlighting output or the like or by a frequency map. The second annotation indicates that the portion specified by the second annotation is a target to be detected as an object.
[0027] An example of specifying the second annotation will be described with reference to FIG. 3. In the image with object highlighting output or the like in the upper right of FIG. 3, the portion that is originally supposed to be a target to be detected as an object is regarded as a portion 322 that is not estimated as an object. The portion 322 that is not estimated as an object is either not highlighted or is displayed according to a low object estimation degree. Thus, when the portion 312 that is originally supposed to be a target to be detected as an object is regarded as the portion 322 that is not estimated as an object, it can be said that the portion 322 is a portion where the object has been "missed in detection". After viewing the image with object highlighting output or the like in the upper right of FIG. 3, the user may specify the second annotation regarding the inspection target image as shown on the left side of the central part of FIG. 3. Specifically, in the image with object highlighting output or the like in the upper right of FIG. 3, the user may specify the second annotation for the portion 322 that is not estimated as an object and is a "missed in detection" portion. The second annotation specifying reception unit 162 receives the specification of the second annotation for the inspection target image on the left side of the central part of FIG. 3. As a result, the second annotation specified portion 372 is set in the inspection target image. In this way, by using the second annotation specifying reception unit 162, after viewing an image with object highlighting output or the like, the user can specify the second annotation indicating that a portion that is regarded as the portion 322 not estimated as an object but where the object has been "missed in detection" is a portion to be detected as an object.
[0028] The learning control unit 107 operates the model realization unit 102 and performs control for executing machine learning to train the model 122. When the first annotation is specified, the learning control unit 107 controls to execute machine learning using the first annotation and the inspection target image associated with the first annotation. The learning control unit 107 performs control for training the model 122 so that the first annotation specified portion 371 becomes an object estimation information such that it is not estimated to be an object or approaches the non - estimated one. When the second annotation is specified, the learning control unit 107 controls to execute machine learning using the second annotation and the inspection target image associated with the second annotation. The learning control unit 107 performs control for training the model 122 so that the second annotation specified portion 372, which is the portion that has received the second annotation, becomes object estimation information such that it is estimated to be an object. When both the first annotation and the second annotation are specified, the learning control unit 107 performs control for training the model 122 so as to satisfy both the training of the model 122 when the first annotation is specified and the training of the model 122 when the second annotation is specified.
[0029] Using FIG. 3, the training of the model 122 using one or both of the first annotation and the second annotation and the effects of the training will be described. The inspection target image on the left side of the central part of FIG. 3 shows a case where both the first annotation and the second annotation are specified. The lower part of FIG. 3 shows the situation after the model 122 is trained by machine learning using the first annotation and the second annotation. The inspection target image at the lower left of FIG. 3 is the same as the one at the upper left of FIG. 3. When such an inspection target image or the like is input to the model 122 after training by machine learning using annotations, an image with object highlighting output based on object estimation information can be as shown in the lower right of FIG. 3. Specifically, after the part 321 where the object was "over-detected" in the upper right of FIG. 3 is designated as the first annotation designated part 371 and the model 122 is trained, in the lower right of FIG. 3, the first annotation designated part 371 can become the part 331 that is no longer estimated as an object. Also, after the part 322 where the object was "undetected" in the upper right of FIG. 3 is designated as the second annotation designated part 372 and the model 122 is trained, in the lower right of FIG. 3, the second annotation designated part 372 can become the part 332 that is correctly estimated as an object. In this way, since the training of the model 122 reflecting one or both of the first annotation and the second annotation specified by the user is realized, a model 122 desired by the user can be constructed. Also, when the accuracy of object detection by the model 122 is not sufficient, after an annotation is specified for the handheld inspection target image and the model 122 is trained by machine learning using the annotation, a model 122 desired by the user can be constructed without adding the image data itself used for learning. In other words, even if the number of image data used for learning is small, sufficient accuracy of object detection by the model 122 can be realized. When the training of the model 122 by machine learning using both the first annotation and the second annotation is executed, it is also possible to make the model 122 learn in a well-balanced manner the part where the object is not desired to be detected and the part where the object is desired to be actively detected in the inspection target image. When the training of the model 122 by machine learning using both the first annotation and the second annotation is executed, more detailed adjustments in the training of the model 122 may be achieved than when adding the image data itself used for learning.
[0030] Since the image processing apparatus 100 in the present disclosure has the functional configuration as described above, it can have the effects shown in the above [Advantages of the Invention].
[0031] FIG. 4 shows the system configuration of the entire system including the image processing apparatus 100 which is an embodiment of the present disclosure. An imaging device 401 that images an inspection object 441, 442, or 443 which is an object to be detected and the image processing apparatus 100 which is an embodiment of the present disclosure are capable of communicating with each other. In FIG. 4, although the image processing apparatus 100 and the imaging device 401 appear to be capable of communicating with each other by a wired connection, they may be capable of communicating with each other wirelessly. The imaging device 401 transmits the captured image data to the image processing apparatus 100. During object detection operation, the image processing apparatus 100 determines the presence or absence of an object included in the inspection target image received from the imaging device 401. The image data captured by the imaging device 401 may be given to the image processing apparatus 100 in batches. The imaging device 401 may be integrated with the image processing apparatus 100. What is integrated with the imaging device 401 and the image processing apparatus 100 may be, for example, a smartphone with a camera or what is called a so-called smart camera. Separate from the image processing apparatus 100, there may be some display device or output device for displaying or outputting an image with object emphasis display output, a frequency map, etc. Separate from the image processing apparatus 100, there may be a display device for displaying the annotation specification screen 1000 in FIG. 10 and some input device for receiving an input using the annotation specification screen 1000.
[0032] FIG. 5 shows a computer architecture for implementing an embodiment of the present disclosure. To implement the image processing apparatus 100, an information processing apparatus (e.g., CPU or processor) 501, a storage device (e.g., memory) 502, a non-volatile recording medium or recording apparatus (e.g., non-volatile memory) 503, an external recording medium drive (e.g., disk drive) 504, an input device (e.g., mouse, keyboard, imaging device, sensor, touch panel, pointing device) 506, a display or output device, or a display unit (e.g., display) 507, a communication device 508, an external input / output port 509, and a part or all of an imaging device connection port 510 are interconnected by an interconnecting unit (e.g., bus) 511. One or more of a functional unit program group 521 (e.g., a program for implementing the functional configuration according to the present disclosure), a model program group 522, a model parameter group 523, image data, etc. 524, various databases 525, and various information 526 may be recorded in the non-volatile recording medium or recording apparatus 503. The external recording medium drive 504 can connect an external recording medium 505. Part or all of the information shown in 521 to 526 above may be provided via the input device 506, the communication device 508, or the external input / output port 509 and recorded or stored in the non-volatile recording medium or recording apparatus 503 and the storage device 502.
[0033] Hereinafter, the processing of the embodiment of the present disclosure will be described along the flow of processing that can be performed in the embodiment of the present disclosure.
[0034] FIG. 6 shows an overall flow of an example of the flow of processing that can be performed in the embodiment of the present disclosure. In step 601 of FIG. 6, training of the model 122 is performed to construct the model 122 used in the image processing apparatus 100. The training of the model 122 is realized by machine learning using a group of learning images.
[0035] The execution entity of step 601 is the image processing apparatus 100 in FIG. 1. Not limited to this, model parameters and the like obtained in a machine learning system, which is a separate image processing apparatus different from the image processing apparatus 100 in FIG. 1, may be transferred or transplanted to the image processing apparatus 100 that operates the model 122.
[0036] In step 602 of FIG. 6, the image processing apparatus 100 uses the inspection target image as an input to the model 122 to obtain object estimation information, and controls to display an image with object highlighting output or a frequency map based on the object estimation information. Through the display performed in step 602, it can be confirmed whether the model 122 trained in step 601 can achieve object detection desired by the user. As a result of this confirmation, if it is found that there are "over-detections" or "detection omissions" of objects, steps 603 and 604 will be executed.
[0037] In step 603 of FIG. 6, the image processing apparatus 100 utilizes the image with object highlighting output or the frequency map obtained in step 602 to accept the specification of annotations for the inspection target image used in step 602. If a part that is not a target to be detected as an object is estimated as a part to be detected as an object, the user assigns a first annotation to the part to specify that the part is not a target to be detected as an object. Also, if a part that is a target 312 to be detected as an object is estimated as a part that is not detected as an object, the user assigns a second annotation to the part to specify that the part is a target to be detected as an object.
[0038] In step 604 of FIG. 6, the image processing apparatus 100 trains the model 122 by performing machine learning using the inspection target image used in step 602 and the annotation specified in step 603. Here, the model 122 is trained so that the first annotation specified portion 371, which is the portion of the inspection target image where the first annotation is specified, does not become or approaches something that is not estimated to be an object, and becomes object estimation information. Also, the model 122 is trained so that the second annotation specified portion 372, which is the portion of the inspection target image where the second annotation is specified, becomes or approaches something that is estimated to be an object, and becomes object estimation information.
[0039] In step 605 of FIG. 6, the image processing apparatus 100, similar to step 602, uses the inspection target image as an input to the model 122 to obtain object estimation information, and controls to display an image with object emphasis display output or a frequency map based on the object estimation information. The model 122 in step 605 is the one after being trained in step 604. By executing this step 605, the user can confirm the degree of improvement in the accuracy of object detection using the model 122 by training the model 122 using the annotation in step 604.
[0040] When the model 122 is first trained by machine learning, the learning images are used as the input to the model. What kind of images are suitable as learning images depends on the machine learning method. For example, if the model 122 is to be trained so that it can reproduce the features of an image without an object, etc., the learning image may be an image obtained by imaging an object without an object. If it is an image obtained by imaging an object without an object, the collection of learning images is easier than an image obtained by imaging an object with an object. Note that depending on the ingenuity of the machine learning method, it is also possible to use an image obtained by imaging an object including an object as a learning image. One of such ingenuities of the machine learning method is the ingenuity in the definition of the loss function. The learning control unit 107 controls to train the model 122 based on the output of the model 122. The learning control unit 107 controls, for example, to execute the training of the model based on one or more of object estimation information, loss function value (Loss), neural network output, or input to the model. The training of the model may include, for example, the update calculation of the parameters of the in-model neural network 221. If each of the feature maps from each of the multiple layers of the learned neural network 101 is used, the learning control unit 107 controls to execute the training of the model by machine learning for each of the models 122 corresponding to each of the above-mentioned layers.
[0041] FIG. 7 shows a flowchart of the process when training a model by machine learning. In step 701 of FIG. 7, the learning control unit 107 selects one of the learning images to be used for machine learning. In step 702 of FIG. 7, the image processing apparatus 100 determines whether the input to the model 122 is a feature map obtained from an image. If the determination result in step 702 is affirmative, the control proceeds to step 703. If the determination result in step 703 is negative, the control proceeds to step 706. If it is clear in the image processing apparatus 100 whether the input to the model 122 is the image itself or a feature map obtained from the image according to the specifications of the model 122, step 702 may be omitted and the flowchart of FIG. 7 may be started from step 703 or step 706. In step 703 of FIG. 7, the learned neural network 101 takes in the selected learning image as the input to the learned neural network 101. In step 704 of FIG. 7, the learned neural network 101 outputs a feature map obtained from the selected learning image. In step 705 of FIG. 7, the model realization unit 102 takes in the feature map output by the learned neural network 101 as the input to the model. After step 705, the control proceeds to step 723. In step 706 of FIG. 7, the model realization unit 102 takes in the learning image itself as the input to the model. After step 706, the control proceeds to step 723. In step 723 of FIG. 7, the learning control unit 107 obtains a neural network output or a loss function value (Loss) from the model 122. In step 724 of FIG. 7, the learning control unit 107 determines whether all of the learning images planned to be used for machine learning have been selected in step 701. If the determination result in step 724 is affirmative, the control proceeds to step 725. If the determination result in step 724 is negative, the control proceeds to step 701 and the processing continues for the next selected learning image. In step 725 of FIG. 7, the learning control unit 107 controls to train the model 122. For example, the learning control unit 107 controls to perform update calculation of the parameters (model parameters) of the in-model neural network 221 so that the value (absolute value) of the loss function approaches the optimal value. The optimal value is, for example, zero. However, depending on the model design, noise in the learning data, the purpose of preventing overfitting, etc., the optimal value may be other than zero.
[0042] FIG. 8 shows a flowchart of a process of obtaining object estimation information for an image to be inspected and a process of performing display based on the object estimation information. In step 802 of FIG. 8, the image processing apparatus 100 determines whether the input to the model 122 is a feature map obtained from the image to be inspected. If the determination result in step 802 is affirmative, the control transitions to step 803. If the determination result in step 803 is negative, the control transitions to step 806. Although it has been described above as if a conditional branch determination is made in step 802, if it is clear in the image processing apparatus 100 whether the input to the model 122 is the image to be inspected itself or a feature map obtained from the image to be inspected due to the specification of the model 122, step 802 may be omitted and the flowchart of FIG. 8 may be started from step 803 or step 806. In step 803 of FIG. 8, the learned neural network 101 takes in the image to be inspected as an input to the learned neural network 101. In step 804 of FIG. 8, the learned neural network 101 outputs a feature map obtained from the image to be inspected. In step 805 of FIG. 8, the model realization unit 102 takes in the feature map output by the learned neural network 101 as an input to the model. After step 805, the control transitions to step 807. In step 806 of FIG. 8, the model realization unit 102 takes in the image to be inspected itself as an input to the model. After step 806, the control transitions to step 807.
[0043] In step 807 of FIG. 8, the model 122 realized by the model realization unit 102 outputs object estimation information.
[0044] After obtaining the object estimation information in step 807, one or both of step 808 for an image with object highlighted display output or the like and step 809 for the frequency map are executed. In step 808 of FIG. 8, the display output control unit 105 controls to display an image with object highlighted display output or the like based on the object estimation information. In step 809 of FIG. 8, the display output control unit 105 controls to display the frequency map based on the object estimation information.
[0045] After obtaining the image with object highlighted display output or the like displayed in step 808 and the frequency map displayed in step 809, the user can check whether there are inconsistencies such as "overdetection" or "detection omission" regarding the positions of the part to be detected as an object 312 and the part not to be detected as an object in the inspection target image, as shown at the upper part of FIG. 3, and the positions of the part estimated as an object 321 and the part not estimated as an object 322 in the image with object highlighted display output or the like.
[0046] If there seems to be no inconsistency such as "overdetection" or "detection omission", since the accuracy of object detection by the model 122 is likely to be acceptable for practical use by the user, steps after step 603 in FIG. 6 do not need to be executed. In that case, the image processing apparatus 100 will perform the actual operation of object detection by the model 122.
[0047] FIG. 9 shows a flowchart of a process for specifying annotation. The processing of the flowchart in FIG. 9 starts, for example, when the user who has viewed an image with object highlighting output such as that displayed in step 808 or step 809 of FIG. 8, or a frequency map, determines that it is necessary to train the model 122 by machine learning using annotations.
[0048] In step 901 of FIG. 9, any one of the annotation specifying reception unit 106, the display output control unit 105, or the image processing apparatus 100 controls to display the annotation specifying screen 1000. The annotation specifying screen 1000 is used to specify one or both of the first annotation and the second annotation for the inspection target image while utilizing an image with object highlighting output or a frequency map. FIG. 10 shows an example of the annotation specifying screen 1000. The upper half of FIG. 10 shows an example of the annotation specifying screen 1000 in a state before starting the work of specifying an annotation. The lower half of FIG. 10 shows an example of the annotation specifying screen 1000 after completing the work of specifying an annotation and before clicking the specification confirmation icon 1035.
[0049] The annotation specifying screen 1000 displays, for example, an annotation specifying work window 1001, an object estimation partial highlighting window 1002, an annotation type specifying column 1011, an intensity specifying column 1012, a partial identification method specifying column 1013, a specification start icon 1031, a specification save icon 1032, an undo icon 1033, a specification deletion icon 1034, and a specification confirmation icon 1035. The annotation specification work window 1001 may display any one of the inspection target image for which annotation is to be specified, the feature map obtained from the inspection target image, or the processed inspection target image or the processed feature map. Examples of the processed image or feature map include the binarized original image or feature map, the simplified original image or feature map, and the image or feature map with reduced resolution. The user can use an input device 506 such as a mouse or a pointing device to perform an operation of overlaying an indication that the first annotation or the second annotation has been specified on the inspection target image or the like displayed in the annotation specification work window 1001.
[0050] The object estimation part highlighting window 1002 displays an image with object highlighting output or the like, or a frequency map, for the inspection target image for which annotation is to be specified.
[0051] The annotation type specification column 1011 is for the user to select what type of annotation to specify in the annotation specification work window 1001. The annotation type specification column 1011 includes, for example, the label "Negative", the label "Positive", and radio button icons 1021 corresponding to each of the labels. The user clicks on one of the two radio button icons 1021 present in the annotation type specification column 1011 to make a selection.
[0052] The strength specification column 1012 is for the user to specify the strength of the annotation for the annotation to be the target of the operation specified in the annotation specification work window 1001. The strength of the annotation is the strength of the training of the model 122 in the part (area) specified by the annotation. The strength specification column 1012 includes the label "Strong (S)", the label "Weak (W)", and radio button icons 1021 corresponding to each of the labels.
[0053] The partial specific method designation column 1013 is for the user to select how to specify the part for which the annotation is specified in the operation of specifying the annotation in the annotation specification work window 1001. As the specific method for specifying the part for which the first annotation is specified, for example, there are label specification 1101 (overall specification), frame specification 1121, painted picture specification 1122, or precise specification 1103 (pixel specification, element specification). Among these, the frame specification 1121 and the painted picture specification 1122 are also collectively called region specification 1102. As the specific method for specifying the part for which the second annotation is specified, for example, there are label specification 1201 (overall specification), frame specification 1221, painted picture specification 1222, or precise specification 1203 (pixel specification, element specification). Among these, the frame specification 1221 and the painted picture specification 1222 are called region specification 1202. Note that the frame specification 1121 can specify a frame of any shape, for example, it can be a polygon or a circle.
[0054] FIG. 11 shows an example of a method for specifying a portion for designating a first annotation. As shown in the upper part of FIG. 11, when a portion that is the target 411 not detected as an object in the inspection target image is also a portion estimated as an object in an image with object highlighting output obtained by object inspection using the model 122, it can be said that the object is "over-detected". In order to eliminate such "over-detection" of the object, the user designates a first annotation for the inspection target image. Here, since the user may designate the first annotation so as to include the portion 321 where the object is "over-detected", it is possible to designate it by any of the label designation 1101, frame designation 1121, painting designation 1122, or precise designation 1103 as shown in the lower part of FIG. 11. The label designation 1101 is for designating a first annotation for the entire inspection target image. The frame designation 1121 is for designating a first annotation for the area surrounded by the frame. The painting designation 1122 is for designating a first annotation for the area designated in a painting mode on the annotation designation work window 1001 by an input device 506 such as a mouse or a pointing device (designated by performing an operation as if painting). The precise designation 1103 is for designating a first annotation for the area designated at the pixel or element granularity in the inspection target image.
[0055] FIG. 12 shows an example of a method for specifying a portion for which a second annotation is to be specified. As shown at the top of FIG. 12, when a portion that is the target 312 to be detected as an object in the inspection target image is also a portion 322 that is not estimated as an object in an image with object emphasis display output obtained by object inspection using the model 122, it can be said that the object has been "undetected". In order to eliminate such "undetected" objects, the user specifies a second annotation for the inspection target image. Here, the user can specify the second annotation in any of the label specification 1201, frame specification 1221, painting specification 1222, or precise specification 1203 as shown at the bottom of FIG. 12 in order to specify the second annotation so as to include the portion 322 where the object has been "undetected". (However, if a second annotation is specified for a portion other than the portion 322 where the object has been "undetected" and including a portion that is not the target to be detected as an object, there is a risk of causing new "false detection" in object detection by the model 122 after training the model 122 by machine learning using the second annotation. Therefore, in the task of specifying the portion for which the second annotation is to be specified, the user needs to consider specifying a portion that does not induce new "false detection".) With the above configuration, the user can choose whether to emphasize the ease of specifying the annotation or the fineness of specifying the direction of training the model 122 by the annotation. Depending on the cause of the insufficient accuracy of object detection by the model 122, it is also assumed that the shapes of the portions where "false detection" or "undetected" of the object occur are different. As described above, since there are multiple methods for specifying the portion for which the annotation is to be specified, it is possible to specify the annotation according to the shape of the portion where "false detection" or "undetected" of the object occurs. Also, as a method for specifying the part for annotation, when precise specification 1103 or 1203 (pixel specification, element specification) is selected, it is possible to eliminate or reduce the parts that do not necessarily require annotation within the parts specified by the annotation. Therefore, the content of training the model 122 by machine learning using annotation can be made more appropriate.
[0056] The specification start icon 1031 is for starting a mode in which, when the specification start icon 1031 is clicked, the annotation reception unit 106 performs an operation of specifying an annotation in the annotation specification work window 1001. The specification save icon 1032 is for causing the annotation reception unit 106 to save the content of the operation performed in the mode of specifying an annotation in the annotation specification work window 1001 that has been ongoing until that point and for ending the mode when the specification save icon 1032 is clicked. The Undo icon 1033 is for causing the annotation reception unit 106 to cancel the most recent operation performed using the annotation specification screen 1000 when the Undo icon 1033 is clicked. The specification deletion icon 1034 is for causing the annotation reception unit 106 to delete information regarding the annotation selected on the annotation specification work window 1001 at the time of the click and for erasing the display of the selected annotation from the annotation specification work window 1001 when the specification deletion icon 1034 is clicked. The specified confirmation icon 1035 is used to complete the entire operation on the annotation specification screen 1000 in the annotation specification reception unit 106 when the specified confirmation icon 1035 is clicked, to confirm the information regarding the first annotation saved up to that point in the first annotation specification reception unit 161, and to confirm the information regarding the second annotation saved up to that point in the second annotation specification reception unit 162.
[0057] Hereinafter, the operation of specifying an annotation using the annotation specification screen 1000 as displayed in step 901 of FIG. 9 will be described. The operations performed by the user before and after steps 902 to 905 in the flowchart of FIG. 9 and the processes performed by the annotation specification reception unit 106 will be described.
[0058] First, before clicking the specified start icon 1031, the user selects the settings for starting the mode of performing the operation of specifying an annotation using the annotation specification work window 1001. Specifically, the user clicks one radio button icon 1021 for each of the annotation type specification column 1011, the strength specification column 1012, and the partial specification method specification column 1013 to select the settings for starting the mode of performing the operation of specifying an annotation using the annotation specification work window 1001.
[0059] After selecting the settings in the various specification columns, the user clicks the specified start icon 1031 using the input device 506. Then, the determination result of whether the specified start icon 1031 has been clicked by the annotation specification reception unit 106 in step 902 of FIG. 9 is affirmative, and the control transitions to step 903. In step 903 of FIG. 9, the annotation specification reception unit 106 starts a mode for performing an operation of specifying an annotation using the annotation specification work window 1001. The annotation specification reception unit 106 may determine the setting of the above mode based on the selection status of each radio button icon 1021 in the annotation type specification column 1011, the strength specification column 1012, and the partial specification method specification column 1013 at the time when the specification start icon 1031 is clicked. In the mode for performing an operation of specifying an annotation using the annotation specification work window 1001, the user performs an operation of arranging an annotation on the annotation specification work window 1001 using the input device 506 in accordance with the setting of the mode. As a result of the user performing the operation of arranging the first annotation, for example, the first annotation is arranged as in the first annotation specification portion 371 in the lower half of the annotation specification work window 1001 in FIG. 10.
[0060] When the specification of the annotation using the mode setting determined by the selection status of each radio button icon 1021 in the annotation type specification column 1011, the strength specification column 1012, and the partial specification method specification column 1013 at the time when the specification start icon 1031 is clicked is completed, the user clicks the specification save icon 1032. Then, in step 904 of FIG. 9, the determination result of whether or not the specification save icon 1032 has been clicked by the annotation specification reception unit 106 becomes affirmative, and the control transitions to step 905. In step 905 of FIG. 9, the annotation specification reception unit 106 ends the mode for performing an operation of specifying an annotation using the annotation specification work window 1001. The annotation specification reception unit 106 saves the information for specifying the annotation that has been input so far in the above mode. The information for specifying the annotation here may include the type of the annotation (type such as the first annotation or the second annotation), the strength of the annotation, and the position information of the portion (area) where the annotation is specified.
[0061] After step 905, the user may change the selection status of each radio button icon 1021 in the annotation type designation field 1011, the strength designation field 1012, and the partial specification method designation field 1013, and then click the designation start icon 1031 again. The user can also arrange both the first annotation designation part 371 and the second annotation designation part 372 like the annotation designation work window 1001 in the lower half of FIG. 10.
[0062] Hereinafter, the usage of the designation deletion icon 1034 on the annotation designation screen 1000 will be described. Prior to clicking the designation deletion icon 1034, the user selects the annotation already arranged in the annotation designation work window 1001 using the input device 506. After selecting any one of the annotations already arranged in the annotation designation work window 1001, the user clicks the designation deletion icon 1034 using the input device 506. Then, in step 906 of FIG. 9, the determination result by the annotation designation reception unit 106 as to whether the designation deletion icon 1034 has been clicked becomes affirmative, and the control transitions to step 907. In step 907 of FIG. 9, the annotation designation reception unit 106 discards the information designating the saved annotation regarding the annotation selected at the time when the designation deletion icon 1034 was clicked. Further, the annotation designation reception unit 106 erases the display of the annotation from the annotation designation work window 1001.
[0063] Prior to clicking the designation confirmation icon 1035, the user completes the work of designating all the annotations that the user wants to designate for the inspection target image handled on the annotation designation screen 1000. When the user clicks the specified confirmation icon 1035, in step 908 of FIG. 9, the annotation specifying reception unit 106 determines that the determination result as to whether the specified confirmation icon 1035 has been clicked is affirmative. As a result, the control proceeds to step 909. In step 909 of FIG. 9, the annotation specifying reception unit 106 determines the information specifying the annotation saved in step 905 until then as the confirmation of the annotation for the inspection target image being handled on the annotation specification screen 1000. Since the annotation specification screen 1000 as described above is presented to the user, it is possible to easily specify an annotation for an inspection target image in which "over-detection" or "detection omission" of an object has occurred. Further, according to the annotation specification screen 1000, since it is possible to execute both the specification of the first annotation used to eliminate the "over-detection" of the object and the specification of the second annotation used to eliminate the "detection omission" of the object, the user can handle various causes leading to insufficient accuracy of object detection of the model 122 with an operation on one annotation specification screen 1000.
[0064] Hereinafter, the training of the model 122 by machine learning using an annotation, which is step 604 in the overall flow of FIG. 6, will be described in the order of processing.
[0065] FIG. 13 shows a flowchart of processing when training a model by machine learning using an annotation. In step 1301 of FIG. 13, the learning control unit 107 selects one of the inspection target images to be handled in the machine learning using an annotation. The inspection target image selected here is an inspection target image for which one or both of the first annotation and the second annotation have been specified by the processing in step 603 in the overall flow of FIG. 6. From step 1302 to step 1306 in FIG. 13, the same processing as that from step 702 to step 706 in FIG. 7 may be performed. However, the object input to model 122 from step 1302 to step 1306 in FIG. 13 is the inspection target image selected in step 1301. In step 1323 of FIG. 13, the learning control unit 107 obtains a neural network output or a loss function value (Loss) from model 122. In step 1324 of FIG. 13, the learning control unit 107 determines whether all of the inspection target images to be handled in the machine learning using the annotation have been selected in step 1301. If the determination result in step 1324 is affirmative, the control transitions to step 1325. If the determination result in step 1324 is negative, the control transitions to step 1301, and the processing continues for the next selected inspection target image.
[0066] In step 1325 of FIG. 13, the learning control unit 107 controls to perform training of model 122 in consideration of the annotation specification. For example, the learning control unit 107 performs adjustments such as operating the loss function value reflecting the annotation specification or defining the definition of the loss function reflecting the annotation specification. Then, the learning control unit 107 intends for the absolute value of the loss function value to approach the optimal value for some or all of the two-dimensional spatial coordinates X, Y on the inspection target image (or the two-dimensional spatial coordinates x, y on the feature map), and controls to perform update calculation of the parameters of the neural network included in the model. As described above, by controlling the learning control unit 107 to execute the training of the model by machine learning using annotations, among the inspection target images, the first annotation designated portion 371, which is the portion where the first annotation is designated, is not estimated to be an object or approaches the non-estimated one. The model 122 is trained to be object estimation information. Also, among the inspection target images, the second annotation designated portion 372, which is the portion where the second annotation is designated, is estimated to be an object or approaches the estimated one. The model 122 is trained to be object estimation information. That is, detection of an object using the model 122 is realized in accordance with the intention indicated by the user by specifying an annotation.
[0067] FIG. 14 shows the training of the model 122 by machine learning using annotations in the case where the first annotation is designated only for a part of the portions estimated to be objects as a result of using the inspection target image as the input of the model 122. The annotation specification and the training of model 122 by machine learning using the annotation will be described when there are the following assumptions. That is, in the inspection target image shown in the upper left part of FIG. 14, it is assumed that there are a part 312 that is the target to be detected as an object and a part that is not the target to be detected as an object. And, assuming that the inspection target image is used as the input to model 122, the result of object detection using model 122 is an image with object emphasis display output shown in the upper right part of FIG. 14. Here, assuming that there are two parts 321 and 1422 estimated as objects in the image with object emphasis display output and the like. Part 321 corresponds to a part that is not the target to be detected as an object in the inspection target image and the like, and is a part 321 where the object is "over-detected". On the other hand, part 1422 corresponds to a part 312 that is the target to be detected as an object in the inspection target image, and is a part 1422 where the presence of the object is correctly estimated.
[0068] When there are assumptions as above, the user makes an annotation specification for the inspection target image shown on the left side of the center of FIG. 14. Specifically, the user makes a first annotation specification for part 321 where the object is "over-detected". The first annotation specification part 371, which is the part where the first annotation specification is made so as to include part 321 where the object is "over-detected" and not to include part 1422 where the presence of the object is correctly estimated, may be set. The part 1422 where the presence of the object is correctly estimated may be called the object estimation maintenance part 1472. In the inspection target image, when there are a first annotation designated part 371 which is a part that has received the designation of the first annotation and an object estimation maintenance part 1472 which is a part that has not received the designation of the first annotation among the parts estimated to be objects, the learning control unit 107 trains the model 122 using the first annotation so that the object estimation maintenance part 1472 is still estimated to be an object. Further, the image processing apparatus 100 causes the display unit to display a display image in which a first part estimated to be a first object and a second part estimated to be a second object in the inspection target image are emphasized, and by receiving the designation of the first annotation only for the first part among the first part and the second part via the display unit, the learning control unit 107 uses the inspection target image and the first annotation so that the part where the first annotation of the inspection target image is designated is not detected as the first object, and the machine learning model may be trained so that the second object is still detected.
[0069] As described above, for the two parts 321 and 1422 estimated as objects in the inspection target image, after distinguishing between the first annotation designated part 371 and the object estimation maintenance part 1472, the model 122 is trained by machine learning using the annotation. At the same time, for the object estimation maintenance part 1472, the model 122 is trained so as to be object estimation information such that it is still estimated to be an object even after the model 122 is trained by machine learning using the annotation. For example, if the object estimation information is a set of object estimation degrees at two-dimensional spatial coordinates X and Y on the inspection target image and is based on the loss function value Loss_X,Y of the model 122 at the two-dimensional spatial coordinates X and Y, the model 122 may be trained in the following manner. That is, for the loss function value Loss_X,Y of the model 122 at the two-dimensional spatial coordinates X and Y included in the first annotation specifying portion 371, the parameter update calculation of the in-model neural network 221 is performed so that Loss_X,Y approaches the optimum value after the training of the model 122.
[0070] There can be at least two methods for handling the object estimation maintenance portion 1472. The first method is to prevent the parameter update calculation of the in-model neural network 221 from being performed to correct the value of Loss_X,Y or Loss_x,y, which is the loss function value of the model 122 at the two-dimensional spatial coordinates X and Y or the two-dimensional spatial coordinates x and y included in the object estimation maintenance portion 1472, during the training of the model 122 using machine learning with annotations. The second method is to change the definition of the loss function at the two-dimensional spatial coordinates X and Y or the two-dimensional spatial coordinates x and y included in the object estimation maintenance portion 1472, and then perform the parameter update calculation of the in-model neural network 221 so that the value of the changed loss function approaches the optimum value. Here, the definition of the loss function at the two-dimensional spatial coordinates X and Y or the two-dimensional spatial coordinates x and y included in the object estimation maintenance portion 1472 may be such that the state where the input and output to the in-model neural network 221 at the two-dimensional spatial coordinates X and Y or the two-dimensional spatial coordinates x and y are somewhat different in value results in the loss function value becoming the optimum value. For example, the input to the model is a feature map, the feature map is a three-dimensional array of input feature amounts OFM_x,y,c with two-dimensional spatial coordinates x, y and channel coordinate c as indices, the neural network output is a three-dimensional array of output feature amounts AEOFM_x,y,c, and assuming the number of elements in the channel direction is N_c, the loss function Loss_Positive_x,y at the two-dimensional spatial coordinates x, y in the object estimation maintenance part 1472 may be redefined using the following formula. Note that the sigma (Σ) in the following formula is the sum with respect to the channel coordinate c. Also, max{A,B} means the larger of A and B. α is a positive constant. Loss_Positive_x,y = max{α - (1 / N_c)*Σ(OFM_x,y,c - AEOFM_x,y,c) 2 ,0} The above loss function is similar to the loss function Loss_x,y shown in FIG. 2, but it is different from the loss function Loss_x,y shown in FIG. 2 in that the loss function Loss_Positive_x,y becomes the optimum value when the one based on the square of the difference between the neural network input and the neural network output is α or more. For the object maintenance part 147, by using the loss function Loss_Positive_x,y using the above max function, the object maintenance part 147 will be maintained even after the training of the model 122 by machine learning using annotations.
[0071] Below FIG. 14 shows the state where object detection is performed with the inspection target image as the input of the model 122 after the training of the model 122 by machine learning using annotations as described above. As shown in the image with object emphasis display output shown in the lower right part of FIG. 14, the part 321 where the object was previously "over-detected" is no longer estimated as an object. That part is indicated by 1421. On the other hand, the part 1422 that has correctly estimated the object from before is still maintained. As described above, even when both the part where the object is "over-detected" and the part where the object is correctly estimated (object estimation maintenance part 1472) exist in the object detection result by the model 122, through the training of the model 122 by machine learning using annotations, it is possible to prevent the existence of the object from being estimated in the part where the object is "over-detected" while correctly maintaining the object estimation maintenance part 1472.
[0072] As shown in FIG. 15, when the model 122 is trained by machine learning using annotations, among the parts in the inspection target image that are not the targets to be detected as objects, the handling of the first annotation designated part 371, which is the part where the first annotation is designated, and the non-annotation designated part 1575, which is the part where the first annotation is not designated (since the object is not "over-detected"), is distinguished. The learning control unit 107 trains the model 122 based on the first annotation so that in the inspection target image, the part estimated to be an object is preferentially estimated to be a non-detection target over the non-detection target part existing outside that part. Also, the learning control unit 107 trains the model 122 so that in the inspection target image, the loss function value of the part corresponding to the first annotation converges earlier than the loss function value of the non-detection target part existing outside the part corresponding to the first annotation. Further, the learning control unit 107 trains the model 122 so that the convergence speed of the loss function value of the part corresponding to the first annotation toward the optimal value is greater than that of the loss function value of the non-detection target part existing outside the part corresponding to the first annotation. The learning control unit 107 trains the model 122 to increase the loss function value of the part corresponding to the first annotation or, after weighting, use the loss function value to accelerate the convergence of the loss function value of the part corresponding to the first annotation to the optimal value. For example, when training the model 122 by machine learning using annotations, the model 122 is trained so as to be object estimation information such that the first annotation designated portion 371 is changed more greatly toward a state where it is less likely to be estimated as an object than the non-annotation designated portion 1575. For example, when training the model 122 by machine learning using annotations, the convergence rate of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) of the two-dimensional spatial coordinates X, Y (or x, y) included in the first annotation designated portion 371 is greater toward the optimum value than the value of the loss function Loss_NonAnnotation_X,Y (or Loss_NonAnnotation_x,y) of the two-dimensional spatial coordinates X, Y (or x, y) included in the non-annotation designated portion 1575, and the model 122 is trained. To realize such training of the model 122, when performing the update calculation of the parameters (model parameters) of the in-model neural network 221, the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) of the two-dimensional spatial coordinates X, Y (or x, y) included in the first annotation designated portion 371 is handled as if it were large. Alternatively, when performing the update calculation, it may be handled as if a measure of weighting the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) (for example, multiplying by a number greater than 1) has been taken. Thus, if the update calculation of the parameters of the in-model neural network 221 is performed using the value of the handled loss function, it can be expected that the amount of change in the model parameters that can be involved in the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) will increase. By doing so, it can be expected that the time required for training the model 122 until the detection of the object by the model 122 reaches the desired accuracy for the user and the like will be shortened. This is because more powerful correction will be performed on the first annotation designated portion 371 where the object is "over-detected". In addition, when there is a situation where it is desired to reduce the sensitivity of object detection (make it difficult to detect an object) for a predetermined position (for example, the background portion) in the inspection target image, there may also be a utilization method of assigning a first annotation (setting the first annotation designated portion 371) to the predetermined position. The handling of the loss function Loss_NonAnnotation_X,Y (or Loss_NonAnnotation_x,y) in the non-annotation designated portion 1575 may remain the same as the loss function Loss_x,y described with reference to FIG. 2. Therefore, for the portion (area) where the determination of the presence or absence of an object by the model 122 has already been correctly made before training the model 122 by machine learning using annotations, the accuracy of the object of the model 122 can be maintained even after training the model 122 by machine learning using annotations.
[0073] FIG. 16 shows another method of training the model 122 by machine learning using annotations for the same first annotation designated portion 371 and non-annotation designated portion 1575 as in FIG. 15. In the method shown in FIG. 15, in the training of the model 122 by machine learning using annotations, the handling of the first annotation designated portion 371 and the handling of the non-annotation designated portion 1575 were always distinguished. In the method shown in FIG. 16, such a distinction is made only when such a distinction is appropriate.
[0074] First, the learning control unit 107 determines whether to distinguish the handling of the first annotation designated portion 371 and the handling of the non-annotation designated portion 1575 in the training of the model 122 by machine learning using annotations. The determination material for making the determination may be, for example, the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) of the two-dimensional spatial coordinates X,Y (or x,y) included in the first annotation specification part 371, which is a representative value of the loss function value (Loss_Negative) of the first annotation specification part, and the value of the loss function Loss_NonAnnotation_X,Y (or Loss_NonAnnotation_x,y) of the two-dimensional spatial coordinates X,Y (or x,y) included in the non-annotation specification part 1575, which is a representative value of the loss function value (Loss_NonAnnotation) of the non-annotation specification part. Here, the definition of the "representative value" for specifying Loss_Negative and Loss_NonAnnotation may be determined statistically, and may be, for example, any of the average value, median value, or mode value. For example, the input to the model is a feature map, the feature map is a three-dimensional array of input feature amounts OFM_x,y,c with two-dimensional spatial coordinates x,y and channel coordinates c as indices, the neural network output is a three-dimensional array of output feature amounts AEOFM_x,y,c, and assuming the number of elements in the channel direction is N_c, both Loss_Negative_x,y and Loss_NonAnnotation_x,y may be calculated using the following formula for the loss function Loss_x,y at the two-dimensional spatial coordinates x,y. Note that the sigma (Σ) in the following formula is the sum with respect to the channel coordinate c. Loss_x,y=(1 / N_c)*Σ(OFM_x,y,c - AEOFM_x,y,c) 2 That is, for the two-dimensional spatial coordinates x,y where it is estimated that no object exists, it can be expected that the loss function Loss_x,y is close to zero. When the difference between the loss function value of the portion corresponding to the first annotation and the loss function value of the portion corresponding to the non-detection target portion existing outside the portion corresponding to the first annotation in the inspection target image is less than or equal to the threshold value, the learning control unit 107 trains the model 122 without distinguishing between the portion corresponding to the first annotation and the non-detection target portion. In order to determine whether to distinguish between the handling of the first annotation designated portion 371 and the handling of the non-annotation designated portion 1575, the learning control unit 107 may, for example, determine whether Loss_Negative and Loss_NonAnnotation satisfy a predetermined relationship. The predetermined relationship is, for example, whether the difference (Loss_Negative - Loss_NonAnnotation) obtained by subtracting Loss_NonAnnotation from Loss_Negative is greater than a predetermined non-negative threshold value ε.
[0075] When the determination regarding the above difference is negative (when Loss_Negative - Loss_NonAnnotation > ε is not satisfied), as shown in FIG. 16(A), the learning control unit 107 may not distinguish between the handling of the first annotation designated portion 371 and the handling of the non-annotation designated portion 1575 in the training of the model 122 by machine learning using the annotation. The learning control unit 107 may control to execute the training of the model 122 by machine learning using the annotation without performing any processing on the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) of the two-dimensional spatial coordinates X, Y (or x, y) included in the first annotation designated portion 371 (that is, without accelerating the speed at which the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) converges to zero). By doing so, when the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) of the first annotation specified part 371 is not significantly larger than the value of the loss function Loss_NonAnnotation_X,Y (or Loss_NonAnnotation_x,y) of the non-annotation specified part 1575, the model 122 can be trained while balancing both the first annotation specified part 371 and the non-annotation specified part 1575.
[0076] When the determination for the above difference is affirmative (for example, when Loss_Negative - Loss_NonAnnotation > ε), as shown in FIG. 16(B), the learning control unit 107 may distinguish the handling of the first annotation specified part 371 and the handling of the non-annotation specified part 1575 in the training of the model 122 by machine learning using annotations. The learning control unit 107 may, for example, perform some processing on the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) of the two-dimensional spatial coordinates X,Y (or x,y) included in the first annotation specified part 371, and then (that is, increase the speed at which the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) converges to zero,) control to execute the training of the model 122 by machine learning using annotations. The method of processing performed by the learning control unit 107 on the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) of the two-dimensional spatial coordinates X,Y (or x,y) included in the first annotation specification part 371 is such that when performing the update calculation of the model parameters, the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) of the two-dimensional spatial coordinates X,Y (or x,y) included in the first annotation specification part 371 is treated as if it were large, or when performing the update calculation, a measure of weighting the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) (for example, multiplying by a number greater than 1) is treated as having been performed. That is, when calculating the corrected loss function Loss_2_X,Y (or Loss_2_x,y) for model parameter update calculation, the learning control unit 107 may set a value greater than the value of the loss function Loss_Negative_X,Y (or Loss_Negative_x,y) of the two-dimensional spatial coordinates X,Y (or x,y) included in the first annotation specification part 371 as the value of Loss_2_X,Y (or Loss_2_x,y), or may set as the value of Loss_2_X,Y (or Loss_2_x,y) a value obtained by performing a measure of weighting the value of Loss_Negative_X,Y (or Loss_Negative_x,y) (for example, multiplying by a number greater than 1). Alternatively, the learning control unit 107 may calculate the loss function Loss_2_X,Y (or Loss_2_x,y) for model parameter update calculation of the two-dimensional spatial coordinates X,Y (or x,y) included in the first annotation specification part 371 by weighting the difference between Loss_NonAnnotation and Loss_Negative_X,Y (or Loss_Negative_x,y).
[0077] In the case of (A) in FIG. 16, the learning control unit 107 applies Loss_NonAnnotation to the loss function Loss_2_X,Y (or Loss_2_x,y) for updating the model parameters of the two-dimensional space coordinates X, Y (or x, y) included in the first annotation designation part 371. On the other hand, in the case of (B) in FIG. 16, a method of weighting the above-described difference may be used for the loss function Loss_2_X,Y (or Loss_2_x,y) for updating the model parameters of the two-dimensional space coordinates X, Y (or x, y) included in the first annotation designation part 371. As described above, when training the model 122 in machine learning using annotations, the handling of the first annotation designation part can be determined flexibly.
[0078] When an annotation is specified for the image to be inspected, the strength of the annotation may be specified. It can be said that the strength of the annotation indicates the speed of convergence to the optimum value of the loss function value corresponding to the annotation. FIG. 17 illustrates the training of the model 122 by machine learning using annotations when the strength is specified for the first annotation. In the example of FIG. 17, for the image to be inspected, there are two specifications of the first annotation. One is specified as "Strong (S)" with a strength of 173, and the other is specified as "Weak (W)" with a strength of 173. That is, in the example of FIG. 17, for the image to be inspected, etc., the part that is not the object to be detected as an object is divided into three parts: the first annotation specified part 371 (NS) with the specification of "Strong (S)", the first annotation specified part 371 (NW) with the specification of "Weak (W)", and the non-annotation specified part 1575. Here, the loss function of the two-dimensional spatial coordinates X, Y (or x, y) included in the first annotation specified part 371 (NS) with the specification of "Strong (S)" is set as Loss_Negative_S_X,Y (or Loss_Negative_S_x,y). The loss function of the two-dimensional spatial coordinates X, Y (or x, y) included in the first annotation specified part 371 (NW) with the specification of "Weak (W)" is set as Loss_Negative_W_X,Y (or Loss_Negative_W_x,y). The loss function of the two-dimensional spatial coordinates X, Y (or x, y) included in the non-annotation specified part 1575 is set as Loss_NonAnnotation_X,Y (or Loss_NonAnnotation_x,y). In the training of the model 122 by machine learning using annotations, the learning control unit 107 performs an update calculation of the parameters of the neural network 221 in the model so that the loss function value approaches the optimal value for any of the first annotation specified part 371 (NS) with the specification of "Strong (S)", the first annotation specified part 371 (NW) with the specification of "Weak (W)", and the non-annotation specified part 1575.
[0079] When training this model 122, as shown in FIG. 17, the learning control unit 107 sets the degree of training strength in the order of the first annotation specified part 371 (NS) with the specification of "Strong (S)", the first annotation specified part 371 (NW) with the specification of "Weak (W)", and the non-annotation specified part 1575 to be strong. For example, when obtaining the corrected loss function Loss_2_X,Y (or Loss_2_x,y) for model parameter update calculation from the loss function Loss_Negative_S_X,Y (or Loss_Negative_S_x,y) of the two-dimensional spatial coordinates X,Y (or x,y) included in the first annotation designation part 371(NS) with the designation of "strong (S)", the learning control unit 107 relatively increases the degree of amplification of the loss function value. When obtaining the corrected loss function Loss_2_X,Y (or Loss_2_x,y) for model parameter update calculation from the loss function Loss_Negative_W_X,Y (or Loss_Negative_W_x,y) of the two-dimensional spatial coordinates X,Y (or x,y) included in the first annotation designation part 371(NW) with the designation of "weak (W)", the learning control unit 107 relatively decreases the degree of amplification of the loss function value. When obtaining the corrected loss function Loss_2_X,Y (or Loss_2_x,y) for model parameter update calculation from the loss function Loss_NonAnnotation_X,Y (or Loss_NonAnnotation_x,y) of the two-dimensional spatial coordinates X,Y (or x,y) included in the annotation non-designation part 1575, the learning control unit 107 does not perform amplification of the loss function value. As described above, after obtaining the corrected loss function Loss_2_X,Y (or Loss_2_x,y) for model parameter update calculation, the learning control unit 107 may perform update calculation of the model parameters using Loss_2_X,Y (or Loss_2_x,y). It can be expected that the greater the degree of amplification of the loss function value, the stronger the degree of correction of the value of the loss function (the faster the convergence speed of the value of the loss function to the optimum value) in object detection by the model 122 using the model parameters after the update calculation. In this way, if the update calculation of the parameters of the neural network 221 in the model is performed using the value of the corrected loss function Loss_2_X,Y (or Loss_2_x,y) for model parameter update calculation that has been handled, it can be expected that the amount of change in the model parameters that can affect the value of the loss function will be adjusted in accordance with the user's intention. If the annotation is specified with strength as described above, the user can adjust the training strength of the model 122 for each position on the image to be inspected. For example, when an operation such as making a difference in the sensitivity of object detection (the degree to which an object is easily detected) depending on the position on the image to be inspected is desired, the setting of the annotation strength can be utilized. For example, when there is a position where the sensitivity of object detection is set low (when it is desired to make it difficult to detect an object), the strength 173 of the first annotation given to that position may be set to "strong (S)". For example, when an object is "over-detected" in the background portion of the image to be inspected, the strength 173 of the first annotation given to that background portion is set to "strong (S)".
[0080] The portion 322 where the object shown in the upper right part of FIG. 3 is "undetected" is designated as the second annotation designation portion 372, and training of the model by machine learning using the second annotation is performed. The result of object detection by the trained model 122 is as shown in the lower right part of FIG. 3. The portion 322 where the object was previously "undetected" becomes the portion 332 that is correctly estimated as an object in the lower right part of FIG. 3 after training of the model 122.
[0081] In order for the learning control unit 107 to realize the training of the model 122 by machine learning using the second annotation, for example, after changing the definition of the loss function in the two-dimensional space coordinates X, Y or the two-dimensional space coordinates x, y included in the second annotation designation portion 372, the parameter update calculation of the in-model neural network 221 is performed so that the value of the changed loss function approaches the optimal value. Here, as the definition of the loss function in the two-dimensional space coordinates X, Y or the two-dimensional space coordinates x, y included in the second annotation designation portion 372, it is defined such that the state where the input and output to the in-model neural network 221 in the two-dimensional space coordinates X, Y or the two-dimensional space coordinates x, y are somewhat different in value results in the value of the loss function being the optimal value. For example, the input to the model is a feature map, the feature map is a three-dimensional array of input feature quantities OFM_x,y,c with two-dimensional spatial coordinates x, y and channel coordinate c as indices, the neural network output is a three-dimensional array of output feature quantities AEOFM_x,y,c, and assuming the number of elements in the channel direction is N_c, the loss function Loss_Positive_x,y at the two-dimensional spatial coordinates x, y in the second annotation specification part 372 may be redefined using the following formula. Note that the sigma (Σ) in the following formula is the sum with respect to the channel coordinate c. Also, max{A,B} means the larger of A and B. α is a positive constant. Loss_Positive_x,y = max{α - (1 / N_c) * Σ(OFM_x,y,c - AEOFM_x,y,c) 2 , 0} The above loss function is similar to the loss function Loss_x,y shown in FIG. 2, but differs from the loss function Loss_x,y shown in FIG. 2 in that the loss function Loss_Positive_x,y becomes the optimal value when the square of the difference between the neural network input and the neural network output is α or more. Regarding the second annotation specification part 372, by performing update calculation of the model parameters using the loss function Loss_Positive_x,y using the above max function, after training the model 122 by machine learning using the annotation, the input and output to the in-model neural network 221 at the two-dimensional spatial coordinates X, Y or the two-dimensional spatial coordinates x, y included in the second annotation specification part 372 are expected to be in a state where their values are somewhat separated. That is, in object detection by the model 122, the presence of an object in the second annotation specification part 372 is estimated to occur. That is, the image processing apparatus 100 can accept the designation of a second annotation indicating that any of the portions not estimated to be objects in the inspection target image is a target to be detected as an object. The learning control unit 107 uses one or both of the first annotation and the second annotation, and the inspection target image. If the first annotation is specified, the portion corresponding to the first annotation is not estimated to be an object, and if the second annotation is specified, the portion corresponding to the second annotation is estimated to be an object, and trains the model 122 accordingly. As described above, the image processing apparatus 100 according to the embodiment of the present disclosure can appropriately eliminate not only the "over-detection" of objects but also the "detection omission" of objects.
[0082] In the overall flow of FIG. 6, first, an example of executing "machine learning without annotation" and then executing "machine learning using annotation" was shown. Further, "machine learning using annotation" includes "machine learning using the first annotation", "machine learning using the second annotation", and "machine learning using the first annotation and the second annotation". FIG. 18 shows modes of multiple executions of machine learning in the embodiment of the present disclosure. As shown in FIG. 18, the machine learning 1801 without annotation, the machine learning 1802 using the first annotation, the machine learning 1803 using the second annotation, and the machine learning 1804 using the first annotation and the second annotation may be executed in any temporal order. That is, the machine learning 1801 without annotation, the machine learning 1802 using the first annotation, the machine learning 1803 using the second annotation, and the machine learning 1804 using the first annotation and the second annotation may be executed in any order and any number of times. As described above, the training of the model 122 by machine learning can be executed step by step or by trial and error. Moreover, any of machine learning without annotation 1801, machine learning using the first annotation 1802, machine learning using the second annotation 1803, and machine learning using the first and second annotations 1804 may be executed using the same inspection target image. That is, the same inspection target image may be repeatedly used as learning data for machine learning. As described above, the same inspection target image or the like can be reused as learning data for machine learning. That is, the cost of collecting learning data can be reduced.
[0083] The present disclosure is not limited to the above-described embodiments and includes various modifications. Part of the configuration and processing of the embodiment may be replaced with the configuration and processing of other conceivable embodiments. The configuration and processing of other conceivable embodiments may be added to the configuration and processing of the embodiment. For example, in the present disclosure, there may be the following modifications of the embodiments.
[0084] (α) Limitation of Annotation Designated Portion In the above-described embodiment, as shown in FIG. 11, the first annotation is designated. Here, the annotation designation reception unit 106 may perform a restriction such that the designation of the first annotation is received only within the range of the portion estimated as an object (or the portion having an object estimation degree equal to or higher than a predetermined value) in the image with the object emphasized display output. By providing such a restriction, even if the user roughly designates the target area to which the first annotation is applied in the designation of the first annotation, among the roughly designated areas, the portion estimated as an object (or the portion having an object estimation degree equal to or higher than a predetermined value) The first annotation designated portion 371 can be set within the range. That is, the work load of the user can be reduced.
[0085] (β) Applications of Object Detection and Image Classification In the description of the above embodiment, as an example of an object, the presence of abnormalities (defects and scratches) included in the inspection target image has been mainly dealt with. However, as already mentioned, as an object, for example, a wide range of things such as objects and characters included in the inspection target image are assumed. And the object detection according to the embodiment of the present disclosure can also be applied to these objects. For example, when the image processing apparatus 100 which is an embodiment of the present disclosure is used to detect an object as an object, it becomes possible to realize automatic recognition of the presence or absence of an object included in the inspection target image, the type and number of the object. For example, when the image processing apparatus 100 which is an embodiment of the present disclosure is used to detect an image segment (region on an image) as an object, it becomes possible to realize image segmentation of extracting an image segment from the inspection target image (or dividing the inspection target image into several image segments). As described above, the present disclosure can be applied to a wide range of targets as an object.
[0086] (γ) Priority of the first annotation and the second annotation In the description of the above embodiment, there is a mention of being able to specify a strength 173 for each of the first annotations and being able to specify a strength 173 for each of the second annotations. Here, the relative strength relationship between the training strength of the model 122 in the part where the first annotation is specified and the training strength of the model 122 in the part where the second annotation is specified may be set in any way. Note that the training strength of the model 122 in the annotation non - specified part 1575 may be set weaker than the training strength of the model 122 in the part where the first annotation is specified or the part where the second annotation is specified. As described above, the content of the training of the model 122 by machine learning using annotations can be determined flexibly.
[0087] The technical matters shown in each of the embodiments of the present disclosure and the modified examples of the embodiments described above can be combined as appropriate as long as no technical contradiction occurs.
Claims
1. An image processing apparatus comprising a processor that detects an object included in an inspection target image in which an inspection target object is imaged, using a machine learning model, wherein the processor, trains the machine learning model using a learning image in which the inspection target object is imaged, outputs a portion estimated to be the object in the inspection target image using the trained machine learning model, causes a display unit to display a display image in which a portion estimated to be the object based on the inspection target image is emphasized, receives, via the display unit, a designation of a first annotation indicating that a portion emphasized as the object in the display image displayed on the display unit is a non-detection target that is not a target to be detected as the object, trains the machine learning model so as not to detect a portion of the inspection target image designated by the first annotation as the object, using the inspection target image and the first annotation, an image processing apparatus.
2. wherein the processor, trains the machine learning model based on the first annotation so that a portion estimated to be the object in the inspection target image is estimated to be the non-detection target preferentially over a portion of the non-detection target existing outside the portion, the image processing apparatus according to claim 1.
3. wherein the processor, trains the machine learning model so that a loss function value of a portion corresponding to the first annotation converges earlier than a loss function value of a portion of the non-detection target existing outside the portion corresponding to the first annotation in the inspection target image, the image processing apparatus according to claim 2.
4. wherein the processor, trains the machine learning model so that a convergence speed of the loss function value of the portion corresponding to the first annotation toward an optimum value is greater than a loss function value of a portion of the non-detection target existing outside the portion corresponding to the first annotation, the image processing apparatus according to claim 3.
5. wherein the processor, trains the machine learning model to accelerate convergence of the loss function value of the portion corresponding to the first annotation to the optimum value by making the loss function value of the portion corresponding to the first annotation larger or multiplying it by a weight and then using the loss function value, the image processing apparatus according to claim 4.
6. The processor In the inspection target image, when the difference between the loss function value of the part corresponding to the first annotation and the loss function value of the part corresponding to the non-detection target part existing outside the part corresponding to the first annotation is equal to or less than a threshold value, the machine learning model is trained without distinguishing between the part corresponding to the first annotation and the part corresponding to the non-detection target part The image processing apparatus according to claim 1.
7. The processor In the inspection target image, among the parts estimated to be the object, when there is a first annotation designated part which is the part that has received the designation of the first annotation and an object estimation maintenance part which is the part that has not received the designation of the first annotation, even if the machine learning model is trained using the first annotation, the machine learning model is trained so that the object estimation maintenance part is still estimated to be the object The image processing apparatus according to claim 1.
8. The processor Using the machine learning model, outputs parts estimated to be a first object and a second object in the inspection target image respectively Based on the inspection target image, causes the display unit to display a display image in which a first part estimated to be the first object and a second part estimated to be the second object are emphasized Among the first part and the second part of the display image, accepts the designation of the first annotation only for the first part via the display unit By using the inspection target image and the first annotation, trains the machine learning model so that the part designated with the first annotation in the inspection target image is not detected as the first object and the second object is still detected The image processing apparatus according to claim 1.
9. The object indicates an abnormality The processor By causing the display unit to display an abnormality degree map that makes the abnormality that may be included in the inspection target image visible, it is possible to accept the designation of the first annotation based on the abnormality degree map The image processing apparatus according to claim 1.
10. The processor For any part in the inspection target image that is not estimated to be the object, it is possible to receive a designation of a second annotation indicating that the part is a target to be detected as the object. Using one or both of the first annotation and the second annotation, and the inspection target image, the part corresponding to the first annotation is not estimated to be the object, and the part corresponding to the second annotation is estimated to be the object, and training the machine learning model. The image processing apparatus according to claim 1.
11. The processor The image used for training the machine learning model together with the first annotation can be configured to be used again for training the machine learning model together with the second annotation. The image processing apparatus according to claim 10.
12. The processor The image used for training the machine learning model can be configured to be used again for training the machine learning model together with the first annotation. The image processing apparatus according to claim 1.
13. The first annotation can be designated with a strength. The processor is configured to be able to adjust the degree of non-detection as the object according to the strength by training the machine learning model using a plurality of the first annotations with different strengths designated. The image processing apparatus according to claim 1.
14. The strength indicates the speed of convergence to the optimal value of the loss function value corresponding to the first annotation. The image processing apparatus according to claim 13.
15. An image processing method for detecting an object included in an inspection target image obtained by imaging an inspection target object using a machine learning model, comprising: Training the machine learning model using a learning image obtained by imaging the inspection target object; Outputting, using the trained machine learning model, a part estimated to be the object in the inspection target image; Based on the inspection target image, displaying, on a display unit, a display image in which the part estimated to be the object is emphasized. Receiving, via the display unit, a designation of a first annotation indicating that a portion emphasized as the object in the display image displayed on the display unit is a non-detection target that is not a target to be detected as the object; Training the machine learning model so as not to detect a portion designated by the first annotation of the inspection target image as the object using the inspection target image and the first annotation; and An image processing method.
Citation Information
Patent Citations
Setting device and setting method
JP2023077051A
Appearance inspection device and appearance inspection method
JP2023077059A