Image inspection device

The image inspection device enhances human evaluation of machine learning model outputs by providing an interactive confusion matrix display for clearer insights.

JP2025150299APending Publication Date: 2025-10-09KEYENCE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024051110
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing confusion matrices in machine learning models are inadequate for human evaluation, as they only represent numerical data and do not provide clear insights into the output results.

Method used

An image inspection device that includes a report output unit displaying a confusion matrix, allowing users to specify components and select target image data for enhanced evaluation.

Benefits of technology

Enables accurate and easy evaluation of machine learning model output results by visually presenting confusion matrix data and allowing user interaction for targeted analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025150299000001_ABST
    Figure 2025150299000001_ABST
Patent Text Reader

Abstract

To provide an image inspection device which enables a person to accurately and easily evaluate an output result of a machine learning model.SOLUTION: An image inspection device (1) inspects inspection image data using a model whose parameter is updated by machine learning based on learning image data presented by a user. The image inspection device (1) includes: a report output part (100) for outputting report display including a first confusion matrix (104), on the basis of target image data which is at least one of the learning image data and the inspection image data, and the output result of the model to the target image data; and a reception part (100) for receiving user operation of designating a component (104a) included in the first confusion matrix (104). The report output part (100) outputs report display in which target image data (107a and 108a) corresponding to the component (104a) specified by the user operation are selected.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image inspection device. [Background technology]

[0002] In recent years, devices that use machine learning models to identify and classify image data have become known (see, for example, Patent Documents 1 and 2). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2023-077051 [Patent Document 2] Japanese Patent Publication No. 2023-077054 Summary of the Invention [Problem to be solved by the invention]

[0004] When humans evaluate the output results of machine learning models, they often use a confusion matrix. However, a typical confusion matrix can only represent the number of image data that correspond to its components (true positives, true negatives, false positives, false negatives, etc.). Therefore, it has not always been easy for humans to evaluate the output results of machine learning models.

[0005] In view of the above-mentioned problems, the present invention aims to provide an image inspection device that allows humans to accurately and easily evaluate the output results of a machine learning model. [Means for solving the problem]

[0006] For example, an image inspection device according to the present invention is an image inspection device that inspects inspection image data using a model whose parameters are updated by machine learning based on training image data presented by a user, and is equipped with target image data which is at least one of the training image data and the inspection image data, a report output unit that outputs a report display including a first confusion matrix based on the output result of the model for the target image data, and a reception unit that receives a user operation that specifies components to be included in the first confusion matrix, and the report output unit outputs the report display in which the target image data corresponding to the component specified by the user operation is selected.

[0007] Still other features, elements, steps, advantages, and characteristics will become more apparent from the detailed description that follows and the accompanying drawings related thereto. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide an image inspection device that allows humans to accurately and easily evaluate the output results of a machine learning model. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a schematic diagram showing a configuration of a visual inspection apparatus according to an embodiment of the present invention; [Figure 2] FIG. 2 is a block diagram showing a hardware configuration of the appearance inspection device. [Figure 3] 10A and 10B are diagrams illustrating input / output processing in the learning stage and the operation stage. [Figure 4] FIG. 10 is a diagram showing a schematic flow of a report output function and related functions. [Figure 5] FIG. 10 is a diagram showing a first display example of an AI detection report. [Figure 6] FIG. 10 is a diagram showing a second display example of an AI detection report. [Figure 7] FIG. 10 is a diagram showing a third example of display of an AI detection report. [Figure 8]FIG. 10 is a diagram showing an example of a display of an AI classification report. [Figure 9] FIG. 10 shows a first confusion matrix and a second confusion matrix for multiple classes. [Figure 10] FIG. 10 is a diagram showing an example of a separation degree graph. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. Note that the following description of the preferred embodiments is merely exemplary in nature and is not intended to limit the present invention, its applications, or its uses.

[0011] (Configuration of visual inspection device 1) FIG. 1 is a schematic diagram showing the configuration of an appearance inspection apparatus 1 according to an embodiment of the present invention. The appearance inspection apparatus 1 is an apparatus for determining whether an image of a workpiece to be inspected, such as various parts or products, is acquired and the workpiece image is passed or failed, and can be used in production sites such as factories. Specifically, a machine learning network is built inside the appearance inspection apparatus 1. Multiple machine learning networks are built in the appearance inspection apparatus 1, including a machine learning network generated by learning images of good products corresponding to good products and images of defective products corresponding to defective products, and a machine learning network generated only from images of good products corresponding to good products. A workpiece image captured of the workpiece to be inspected is input to the generated machine learning network, and the machine learning network can determine whether the workpiece image is passed or failed. The appearance inspection apparatus 1 can be understood as one aspect of an image inspection apparatus.

[0012] The entire workpiece may be the object of inspection, or only a portion of the workpiece may be the object of inspection. Also, one workpiece may contain multiple inspection objects. Also, a workpiece image may contain multiple workpieces.

[0013] The appearance inspection device 1 comprises a control unit 2, which is the device main body, an imaging unit 3, a display device (display section) 4, and a personal computer 5. The personal computer 5 is not essential and can be omitted. The personal computer 5 can be used instead of the display device 4 to display various information and images, and the functions of the personal computer 5 can be incorporated into the control unit 2 or the display device 4.

[0014] 1 illustrates a control unit 2, an imaging unit 3, a display device 4, and a personal computer 5 as an example of the configuration of the visual inspection device 1, but any two or more of these can be combined and integrated. For example, the control unit 2 and the imaging unit 3 can be integrated, or the control unit 2 and the display device 4 can be integrated. Furthermore, the control unit 2 can be divided into multiple units and some of them can be incorporated into the imaging unit 3 or the display device 4, or the imaging unit 3 can be divided into multiple units and some of them can be incorporated into other units.

[0015] (Configuration of imaging unit 3) As shown in FIG. 2, the imaging unit 3 includes a camera module (imaging section) 14 and an illumination module (illumination section) 15, and is a unit that acquires workpiece images. The camera module 14 includes an AF motor 141 that drives the imaging optical system, and an imaging board 142. The AF motor 141 is a part that automatically adjusts focus by driving the lens of the imaging optical system, and can perform focus adjustment using a conventionally well-known method such as contrast autofocus. The imaging board 142 includes a CMOS sensor 143 as a light receiving element that receives light incident from the imaging optical system. The CMOS sensor 143 is an imaging sensor configured to acquire color images. Instead of the CMOS sensor 143, a light receiving element such as a CCD sensor can be used.

[0016] The lighting module 15 includes an LED (light emitting diode) 151 as a light emitting element that illuminates an imaging area including a workpiece, and an LED driver 152 that controls the LED 151. The timing, duration, and amount of light emitted by the LED 151 can be arbitrarily controlled by the LED driver 152. The LED 151 may be provided integrally with the imaging unit 3, or may be provided separately from the imaging unit 3 as an external lighting unit.

[0017] (Configuration of display device 4) The display device 4 has a display panel made of, for example, a liquid crystal panel or an organic EL panel. A work image, a user interface image, etc. output from the control unit 2 are displayed on the display device 4. If the personal computer 5 has a display panel, the display panel of the personal computer 5 can be used in place of the display device 4.

[0018] (operation equipment) Examples of operation devices for a user to operate the visual inspection apparatus 1 include, but are not limited to, the keyboard 51 and mouse 52 of the personal computer 5, and any device configured to be able to accept various operations by the user may be used. For example, a pointing device such as the touch panel 41 of the display device 4 is also included in the operation devices.

[0019] User operations on the keyboard 51 and mouse 52 can be detected by the control unit 2. The touch panel 41 is a conventionally known touch-type operation panel equipped with, for example, a pressure-sensitive sensor, and user touch operations can be detected by the control unit 2. The same applies when other pointing devices are used.

[0020] (Configuration of control unit 2) The control unit 2 includes a main board 13, a connector board 16, a communication board 17, and a power supply board 18. The main board 13 is provided with a processor 13a. The processor 13a controls the operation of each connected board and module. For example, the processor 13a outputs a lighting control signal to an LED driver 152 of the lighting module 15 to control the turning on / off of the LED 151. In response to the lighting control signal from the processor 13a, the LED driver 152 switches the turning on / off of the LED 151 and adjusts the lighting time, and also adjusts the light intensity of the LED 151.

[0021] In addition, the processor 13a outputs an imaging control signal to the imaging board 142 of the camera module 14 to control the CMOS sensor 143. The CMOS sensor 143 starts imaging in response to the imaging control signal from the processor 13a and adjusts the exposure time to any desired time to perform imaging. That is, the imaging unit 3 captures an image within the field of view of the CMOS sensor 143 in response to the imaging control signal output from the processor 13a. If a workpiece is present within the field of view, the image of the workpiece is captured. However, if an object other than the workpiece is present within the field of view, the image of the object can also be captured. For example, the visual inspection device 1 can use the imaging unit 3 to capture non-defective product images corresponding to non-defective products and defective product images corresponding to defective products as learning images for the machine learning network. The learning images do not have to be images captured by the imaging unit 3, but may be images captured by another camera, etc.

[0022] On the other hand, when the visual inspection device 1 is in operation, the workpiece can be imaged by the imaging unit 3. The CMOS sensor 143 is configured to be able to output a live image, i.e., a currently captured image, at a short frame rate at any time.

[0023] When the CMOS sensor 143 has finished capturing an image, the image signal output from the imaging unit 3 is input to the processor 13a of the main board 13 for processing, and is also stored in the memory 13b of the main board 13. Specific details of the processing performed by the processor 13a of the main board 13 will be described later. The main board 13 may be provided with a processing device such as an FPGA or a DSP. The processor 13a may also be an integrated processor such as an FPGA or a DSP.

[0024] The connector board 16 is a part that receives power from an external source via a power connector (not shown) provided on the power interface 161. The power supply board 18 is a part that distributes the power received by the connector board 16 to each board and module, and specifically distributes power to the illumination module 15, the camera module 14, the main board 13, and the communication board 17. The power supply board 18 is equipped with an AF motor driver 181. The AF motor driver 181 supplies drive power to the AF motor 141 of the camera module 14 to achieve autofocus. The AF motor driver 181 adjusts the power supplied to the AF motor 141 in response to an AF control signal from the processor 13a on the main board 13.

[0025] The communication board 17 is a part that executes communication between the main board 13 and the display device 4 and the personal computer 5, and communication between the main board 13 and an external control device (not shown). An example of the external control device is a programmable logic controller. The communication may be wired or wireless, and either form of communication can be realized by a conventionally known communication module.

[0026] The control unit 2 is provided with a storage device (storage unit) 19, which may be, for example, a solid state drive, a hard disk drive, or the like. The storage device 19 stores program files 80, setting files, and the like (software) that enable the hardware to execute the various controls and processes described below. The program files 80 and setting files can be stored in a storage medium 90, such as an optical disk, and the program files 80 and setting files stored in the storage medium 90 can be installed in the control unit 2. The program files 80 may be downloaded from an external server via a communication line. The storage device 19 can also store, for example, the image data and parameters for constructing a machine learning network for the appearance inspection apparatus 1.

[0027] That is, the processor 13a of the visual inspection device 1 is configured to read parameters and the like stored in the storage device 19 to construct a machine learning network, input a workpiece image obtained by photographing the workpiece to be inspected into the constructed machine learning network, and determine whether the workpiece is good or bad based on the input workpiece image. By using this visual inspection device 1, a visual inspection method can be performed in which the quality of the workpiece is determined based on the workpiece image. The machine learning network may also be understood as a machine learning model. Note that in this embodiment, for the sake of convenience, the visual inspection device 1 performs a quality determination, but it may also perform a determination to classify the workpiece image into any class. That is, the configuration may be such that "good" and "defective" described for the workpiece image are treated as any class.

[0028] (Input / output processing) 3 is a diagram showing input / output processing in the learning stage and the operation stage of the visual inspection device 1. As shown in this diagram, in the learning stage of the visual inspection device 1, learning of a machine learning model is performed based on learning data presented by a user. In this embodiment, such learning of a machine learning model based on learning data presented by a user is called user learning.

[0029] The training data includes training image data and instruction content. The training image data includes at least one of image data of a good product and image data of a defective product. The instruction content includes classes (labels) such as "This image data is a good product," "This image data is a defective product," and "This part is defective."

[0030] In the above learning, the parameters of the machine learning model are updated (adjusted) so that the output of the machine learning model approaches an expected value according to the teaching content. Multiple machine learning models may be prepared (model 1 to model 3). With this configuration, it becomes possible to arbitrarily select a learning target or an operation target depending on the application of the visual inspection device 1.

[0031] It is not necessary for the user to perform all of the above learning processes. For example, learning with a relatively large amount of calculation (shipment learning) may be completed by the vendor before shipping the visual inspection apparatus 1, and then it may be sufficient for the user to perform only learning with a relatively small amount of calculation, i.e., user learning, before operating the visual inspection apparatus 1.

[0032] That is, the machine learning model of the visual inspection device 1 may include a parameter-fixed portion. The parameter-fixed portion is a layer in which parameters are fixed by learning at the time of shipment by the vendor, in other words, a layer that does not require user learning. Furthermore, the machine learning model of the visual inspection device 1 may include a segmentation model that is easy for the user to learn.

[0033] With this configuration, there is no need for the user to prepare equipment with the high processing power required for deep learning (such as a graphics processing unit (GPU)), or for a vendor to provide an advanced learning environment using a GPU or the like as a cloud service (such as SaaS), thereby lowering the barrier to introducing the appearance inspection device 1.

[0034] As such, the above-mentioned learning should be interpreted broadly to include not only deep learning, which requires a large amount of computation, but also learning with a small amount of computation, i.e., user learning.

[0035] Meanwhile, in the operation stage of the visual inspection apparatus 1, inspection image data is inspected using the trained machine learning model. The inspection may include image classification, anomaly detection, and segmentation.

[0036] In image classification, objects (workpieces) shown in an image are classified as good or bad. In image anomaly detection, abnormal parts are extracted from the image. For example, anomaly detection using an autoencoder is well known. An autoencoder can be understood as a machine learning model that has been trained (parameter adjustment) so that when normal and abnormal images are input, abnormal parts in the abnormal image are more likely to stand out. In image region segmentation, each pixel that forms the image is classified as defective or non-defective.

[0037] The visual inspection device 1 also includes a report output unit (model evaluation result generation unit) that outputs a report display including a first confusion matrix (described later) based on the output result of the machine learning model in the learning stage or the operation stage. In other words, the target image data input to the machine learning model when displaying the report may be at least one of the learning image data and the inspection image data.

[0038] The report output unit can be understood as, for example, one function of editor software executed on the personal computer 5. In other words, the personal computer 5 functions as the report output unit by executing the editor software.

[0039] As mentioned above, the visual inspection device 1 can be easily learned at the customer's site. Because learning is easy, a report display that allows the effect of learning to be easily understood can be useful. The report output function will be described in detail below.

[0040] (Report output function) 4 is a diagram showing an outline of the flow of the report output function and related functions. The flow of this diagram can basically be executed by the personal computer 5. When the flow of this diagram starts, image data is acquired in step S1. The image data may be image data captured in step S1, or may be internal data stored in advance in a storage area of ​​the personal computer 5.

[0041] In steps S2 and S3, the settings of the learning image data and the verification image data are accepted, respectively. A class (e.g., good product / defective product) is set in advance for the learning image data. On the other hand, a class is preferably set in advance for the verification image data, but a class is not necessarily set. Verification image data without a set class may be understood as equivalent to inspection image data in the operational stage.

[0042] In step S4, a learning process of the machine learning model is performed using the training image data and the verification image data.

[0043] In step S5, it is determined whether the learning process has been successful. If the determination in step S5 is YES, the flow proceeds to step S6. For example, if all of the verification image data have been classified into the correct class, it is determined that the learning process has been successful.

[0044] On the other hand, if a negative determination is made in step S5, the flow returns to step S2. If the learning process fails, the user may be prompted to appropriately reset the learning image data and verification image data so that the next learning process will be successful.

[0045] In addition, possible cases in which the learning process fails include: (1) the number of labels in the learning image data is insufficient; (2) more classification labels than specified in the specifications are registered in the AI ​​classification; (3) a larger amount of learning image data than specified in the specifications is registered; (4) the learning process is canceled midway; (5) there is no learning image data; or (6) the prerequisites for the learning process are not met.

[0046] If the determination in step S5 is YES, a report output instruction is accepted in step S6. For example, the user may issue a report output instruction by clicking a function call button on the editor software with the mouse 52. Note that the report output may be any of multiple types of reports (for example, an AI detection report and an AI classification report, which will be described later).

[0047] In step S7, a database of the learning results of the machine learning model is created. For example, a learning tool on the editor software judges the training image data and the verification image data. During this process, the judgment results for each of the training image data and the verification image data are linked to the image file name and recorded. Note that the judgment results may include, for example, the "class (label) determined by the AI" and the "anomaly score."

[0048] In step S8, image data for displaying a report is output. For example, an image for displaying a report may be created from a thumbnail image.

[0049] In step S9, the report data is output. The report data may be output in HTML (hyper text markup language) format, for example. However, if image data is directly embedded in the HTML file when outputting a report in HTML format, the image size (number of pixels) becomes a fixed value. This makes it difficult to change the size of the image displayed in the report later. Also, the resolution of the image displayed in the browser is likely to deteriorate.

[0050] Therefore, when outputting a report in HTML format, it is desirable to save the original-sized image files (.gif / .jpg / .png, etc.) and the HTML file (.html) for displaying the report in separate folders. This type of report output makes it possible to dynamically change the size of images displayed in the browser. It also helps prevent degradation of the resolution of the displayed images. Furthermore, if you create a PDF (portable document format) file using the browser's print function, you can obtain results that are nearly equivalent to report output with image data directly embedded.

[0051] The above-mentioned report data can be used, for example, to evaluate an inspection tool (evaluate the learning results) when a user creates one using a learning function. Examples of such tools include "AI detection" and "AI classification." In this embodiment, "AI detection" is a tool that detects defective areas from image data, and if a defective area of ​​a certain size or larger is detected, the image data is treated as a "defective" image, and if no defective area of ​​a certain size or larger is detected, the image data is treated as a "good" image.

[0052] (AI detection report) 5 is a diagram showing a first display example of an AI detection report. When an instruction to output an AI detection report is accepted in step S6 of FIG. 4, the personal computer 5 displays a user interface screen 100 as shown in this diagram.

[0053] The user interface screen 100 includes, for example, a creation date and time display area 101, an image number display area 102, a result display area 103, a first confusion matrix 104, an image size setting area 105, a display mode switching area 106, a training image display area 107, a verification image display area 108, and an unclassified image display area (not shown).

[0054] The creation date and time display area 101 displays the creation date and time (year, month, date, and time) of the AI ​​detection report.

[0055] The number of images display area 102 displays the number of images acquired in step 1 of FIG. 4 (the total number of training images and verification images, the number of training images only, and the number of verification images only).

[0056] The result display area 103 displays the output results of the machine learning model (correct rate, overlooked rate, and overdetection rate). The correct rate is calculated as the ratio of true positive images (where both the setting and the result are good) and true negative images (where both the setting and the result are bad) to all images. The overdetection rate is calculated as the ratio of false negative images (where the setting is good and the result is bad) to all images. The overdetection rate is calculated as the ratio of false positive images (where the setting is bad and the result is good) to all images.

[0057] The first confusion matrix 104 displays the classes (good / defective / unclassified) set by the user and the classes (good / defective) obtained as the output result of the machine learning model in a matrix format. Verification image data for which the user did not set a class in step 1 of Figure 4 is treated as the unclassified class.

[0058] In this figure, the first row, first column displays the number of true positive images where both the settings and the results are good. The second row, first column displays the number of false negative images where the settings are good but the results are bad. The third row, first column displays the number of images where the settings are good (the total number of true positive and false negative images).

[0059] Meanwhile, the second row, first column displays the number of false positive images where the settings are faulty and the results are good. The second row, second column displays the number of true negative images where both the settings and the results are faulty. The second row, third column displays the number of images where the settings are faulty (the total number of false positive and true negative images).

[0060] Furthermore, row 3, column 1 displays the number of images for which no setting was made, i.e., the number of images for which the setting was unclassified and the result was good. row 3, column 2 displays the number of images for which the setting was unclassified and the result was bad. row 3, column 3 displays the total number of images for which the setting was unclassified. In this way, by showing the number of images for which the setting was unclassified in the confusion matrix, it is possible to notify the user that they have forgotten to make a setting.

[0061] Additionally, the fourth row, first column displays the number of images that are considered to be good products (the total number of true positive images, false positive images, and images that are unclassified in the settings but are considered to be good products). The fourth row, second column displays the number of images that are considered to be bad products (the total number of false negative images, true negative images, and images that are unclassified in the settings but are considered to be bad products). The fourth row, third column displays the total number of images.

[0062] As such, the first confusion matrix 104 includes multiple components 104a. In particular, the first confusion matrix 104 displays the quantity (number of images) of target image data that corresponds to each of the nine components 104a, among the target image data input to the machine learning model when displaying the report. This display makes it easier to visually grasp the inference performance of the machine learning model that is the target of the AI ​​detection report. For example, the larger the numbers on the diagonal components of the first confusion matrix 104 (the total number of true positive images and true negative images), the higher the inference performance of the machine learning model.

[0063] The first confusion matrix 104 also functions as one of the reception units that receives a user operation to specify (select) one of the multiple components 104a contained therein as a display target. For example, any of the multiple components 104a contained in the first confusion matrix 104 can be clicked with the mouse 52. In this figure, the component 104a in the third row and third column of the first confusion matrix 104 (all eight pieces of target image data) has been clicked. The training image display area 107 and the verification image display area 108 each selectively display the target image data corresponding to the clicked component 104a.

[0064] It should be noted that displaying the quantity of target image data (number of images) in the first confusion matrix 104 is not essential for accepting the above-mentioned designation (selection) of target image data.

[0065] The image size setting area 105 displays a slider bar for setting the size (e.g., the number of pixels in the width direction) of the images to be displayed in the learning image display area 107 and the verification image display area 108. The user can dynamically set the image size by dragging the handle of the slider bar left or right with the mouse 52.

[0066] The display mode switching area 106 displays a pull-down menu for selecting the display mode of the images to be displayed in the learning image display area 107 and the verification image display area 108. The pull-down menu may provide options such as "normal image" as exemplified in this figure, as well as "normal image + heat map horizontal combination (described later)."

[0067] In the learning image display area 107, for example, a normal image 107a, a check box 107b, a file name 107c, an extract display button 107d, and a display reset button 107e are displayed.

[0068] As normal image 107a, the learning image data corresponding to component 104a clicked with mouse 52 is selectively displayed. In accordance with this figure, out of component 104a in the third row and third column clicked in first confusion matrix 104 (all eight target image data), four learning image data ("0001.bmp", "0002.bmp", "0003.bmp", and "0004.bmp") are displayed in a list.

[0069] The check box 107b can be understood as a display showing the selection state of the normal image 107a. The selection state of the normal image 107a (= whether or not the check box 107b is checked) may be linked to a click operation on the component 104a included in the first confusion matrix 104. That is, in the check box 107b, the selection state of the normal image 107a may be changed according to the component 104a specified by the user's click operation.

[0070] Furthermore, the check box 107b functions as one of the reception units that receives a user operation to specify (select) one of the normal images 107a displayed in the learning image display area 107 as a display target. For example, each time the check box 107b is clicked with the mouse 52, it alternates between "checked (display on)" and "unchecked (display off)."

[0071] The file names of the normal images 107a are displayed as the file names 107c. For example, when any of the file names 107c is clicked, the corresponding normal image 107a may be enlarged and displayed.

[0072] The extract display button 107d functions as one of the reception units that receives a user operation to extract and display only the normal images 107a designated (selected) as the display target by the check boxes 107b. For example, if the extract display button 107d is clicked with the check box 107b for "0001.bmp" among the four normal images 107a checked, only "0001.bmp" is displayed, and "0002.bmp," "0003.bmp," and "0004.bmp" are all hidden.

[0073] The display reset button 107e functions as one of the reception units that receives a user operation to reset the extraction display by the extraction display button 107d. For example, when the display reset button 107e is clicked while only "0001.bmp" is extracted and displayed, the designation (selection) of the check box 107b is canceled, and the state returns to a list display of all four learning image data ("0001.bmp," "0002.bmp," "0003.bmp," and "0004.bmp").

[0074] In this way, by implementing the check box 107b, the extract display button 107d, and the display reset button 107e, the user can make a final decision on the learning image data to be displayed in the learning image display area 107.

[0075] In the verification image display area 108, for example, a normal image 108a, a check box 108b, a file name 108c, an extract display button 108d, and a display reset button 108e are displayed.

[0076] As the normal image 108a, the verification image data corresponding to the component 104a clicked with the mouse 52 is selectively displayed. In accordance with this figure, of the component 104a in the third row and third column clicked in the first confusion matrix 104 (all eight target image data), four verification image data ("0005.bmp," "0006.bmp," "0007.bmp," and "0008.bmp") are displayed in a list.

[0077] The check box 108b can be understood as an indication of the selection state of the normal image 108a. Note that the selection state of the normal image 108a (= whether or not the check box 108b is checked) may be linked to a click operation on a component 104a included in the first confusion matrix 104. That is, in the check box 108b, the selection state of the normal image 108a may be changed according to the component 104a designated by the user's click operation.

[0078] Furthermore, the check box 108b functions as one of the reception units that receives a user operation to specify (select) as a display target one of the normal images 108a displayed in the verification image display area 108. For example, the check box 108b alternates between "checked (display on)" and "unchecked (display off)" each time it is clicked with the mouse 52.

[0079] The file names of the normal images 108a are displayed as the file names 108c. For example, when any of the file names 108c is clicked, the corresponding normal image 108a may be enlarged and displayed.

[0080] The extract display button 108d functions as one of the reception units that receives a user operation to extract and display only the normal images 108a designated (selected) as the display target by the check boxes 108b. For example, if the extract display button 108d is clicked with the check box 108b for "0005.bmp" among the four normal images 108a checked, only "0005.bmp" is displayed, and "0006.bmp," "0007.bmp," and "0008.bmp" are all hidden.

[0081] The display reset button 108e functions as one of the reception units that receives a user operation for resetting the extraction display by the extraction display button 108d. For example, when the display reset button 108e is clicked in a state where only "0005.bmp" is extracted and displayed, the designation (selection) of the check box 108b is cancelled, and the state returns to a state where all four verification image data ("0005.bmp," "0006.bmp," "0007.bmp," and "0008.bmp") are displayed in a list.

[0082] In this way, by implementing the check box 108b, the extract display button 108d, and the display reset button 108e, the user can make a final decision on the verification image data to be displayed in the verification image display area 108.

[0083] The unclassified image display area (not shown) is included in the display mode switching area 106 and displays unclassified images. As described above, there is no setting for comparing unclassified images with the judgment result, so the unclassified image display area may be configured to display images without dividing the image area according to whether the judgment is correct or incorrect. Furthermore, the unclassified image display area may be configured to display images that can accept operations similar to those of the learning image display area 107 and the verification image display area 108 described above.

[0084] In this way, in the AI ​​detection report, information is formatted and displayed so that the output results of the machine learning model can be easily confirmed by humans. Specifically, the multiple components 104a included in the first confusion matrix 104 are linked to the corresponding target image data and saved as a database. Then, the first confusion matrix 104 accepts a user operation to designate (select) one of the multiple components 104a as the display target, and only the corresponding target image data is selectively displayed. As a result, a report output is obtained that allows humans to accurately and easily evaluate the output results of the machine learning model.

[0085] Figure 6 shows a second display example of an AI detection report. In this second display example, unlike the first display example (Figure 5), the component 104a in the second row and second column (four true negative images) of the multiple components 104a included in the first confusion matrix 104 has been clicked. Therefore, the training image display area 107 and the verification image display area 108 each selectively display the target image data ("0001.bmp," "0003.bmp," "0006.bmp," [0007.bmp]) corresponding to the clicked component 104a.

[0086] 7 shows a third example of the display of an AI detection report. In this third example, the handle of the slider bar in the image size setting area 105 has been dragged to the right (i.e., in the direction in which the setting value increases) compared to the first example of the display (FIG. 5). Therefore, the sizes of the images displayed in the training image display area 107 and the verification image display area 108 (not shown in this figure) have increased.

[0087] AI detection is a tool that detects and judges abnormalities based on the locations in the learning data that the user has set as abnormal. Therefore, AI detection can display the locations that the AI ​​is focusing on when making judgments, i.e., the locations with a high degree of abnormality, as a heat map (= anomaly map).

[0088] Therefore, in the third display example, "Normal image + heat map horizontal combination" is selected in the pull-down menu of the display form switching area 106. Referring to this figure, in the learning image display area 107, a heat map 107f is combined adjacent to the right side of the normal image 107a.

[0089] In this way, by displaying the normal image 107a and the heat map 107f side by side, it is easier to see which part of the target image data the machine learning model is focusing on, compared to when the normal image 107a and the heat map 107f are displayed overlapping each other. Of course, the normal image 107a and the heat map 107f may be partially overlapped as long as the abnormal part of the target image data is not hidden. Furthermore, the display position of the heat map 107f is not necessarily limited to the right side of the normal image 107a, and it may be adjacent to the normal image 107a, above, below, left, or right.

[0090] Furthermore, displaying the normal image 107a and the heat map 107f side by side requires a larger display area than displaying them overlapping each other. Therefore, it is desirable to narrow down the images to be displayed so that the output results of the machine learning model (defect detection model) can be checked while ensuring a clear overview.

[0091] In this way, when the machine learning model is a model that includes a process of calculating the degree of abnormality of the pixels that make up the image, the target image data and the heat map 107f (= abnormality map) generated based on the degree of abnormality may be displayed in the report display without being superimposed.

[0092] The heat map 107f may also be accompanied by a setting / result label 107g and an image type label 107h. The setting / result label 107g displays the class set by the user (good product OK / defective product NG), the class (good product OK / defective product NG) and anomaly score obtained as the output result of the machine learning model. The anomaly score is lower for good images and higher for defective images. The image type label 107h displays either the image type "learning" or "verification."

[0093] For convenience of illustration, the verification image display area 108 is not shown in this figure, but it goes without saying that, like the learning image display area 107, a heat map can be displayed therein.

[0094] (AI classification report) 8 is a diagram showing an example of the display of an AI classification report. When an instruction to output an AI classification report is accepted in step S6 of FIG. 4, the personal computer 5 displays a user interface screen 100 as shown in this diagram.

[0095] The user interface screen 100 includes basically the same display items as the AI ​​detection report ( FIG. 5 ), except for the aforementioned display mode switching area 106, namely, a creation date and time display area 101, an image number display area 102, a result display area 103, a first confusion matrix 104, an image size setting area 105, a training image display area 107, a verification image display area 108, and an unclassified image display area. Note that the verification image display area 108 and the unclassified image display area are not shown in this figure.

[0096] In particular, a second confusion matrix 107i is displayed in the training image display area 107. In the second confusion matrix 107i, classes (good / defective / unclassified) set by the user and output results (good / defective) of the machine learning model are displayed in a matrix format.

[0097] In accordance with this diagram, the first row, first column displays a correctly determined image where the setting is good (i.e., a true positive image where both the setting and the result are good). The first row, second column displays an incorrectly determined image where the setting is good (i.e., a false negative image where the setting is good and the result is bad). The second row, first column displays a correctly determined image where the setting is bad (i.e., a true negative image where both the setting and the result are bad). The second row, second column displays an incorrectly determined image where the setting is bad (i.e., a false positive image where the setting is bad and the result is good).

[0098] In this way, when the machine learning model is a model that includes a process of classifying images, the report display may display the second confusion matrix 107i in which the target image data corresponding to the components that make up the matrix are arranged in the area that displays the components.

[0099] For convenience of illustration, the verification image display area 108 and the unclassified image display area are not explicitly shown in this figure, but it goes without saying that, like the training image display area 107, it is possible to display the second confusion matrix there.

[0100] Incidentally, in an AI classification report, it may be necessary to classify images into multiple classes (for example, up to 32 classes). In this way, in a multi-class AI classification report, displaying the second confusion matrix 107i described above can be effective in improving the visibility of the report output. The reason for this will be explained in detail below.

[0101] 9 is a diagram showing the multi-class first confusion matrix 104 and the second confusion matrix 107i. Note that, for convenience of illustration, the components and images included in each matrix are omitted in this diagram.

[0102] As shown in this figure, the first confusion matrix 104 for the number of classes n (where n is an integer equal to or greater than 2) contains n×n items (for example, 1024 items when n=32). Therefore, as the number of classes n increases, the visibility of the first confusion matrix 104 decreases. Therefore, it is difficult to evaluate the output results of the machine learning model at a glance of the multi-class first confusion matrix 104.

[0103] On the other hand, the second confusion matrix 107i, which displays the correct / incorrect judgment results for each class, has two columns, one for correct judgment and one for incorrect judgment, for n rows of classes CL(1) to CL(n). For example, the first row, first column displays a correctly judged image whose setting and result are both class CL(1). On the other hand, the first row, second column displays an incorrectly judged image whose setting is class CL(1) and whose result is one of classes CL(2) to CL(n). The same applies to the second row and below.

[0104] Therefore, in the case of the second confusion matrix 107i, the number of components included therein is significantly reduced to n × 2 items (for example, 64 items), and the visibility is dramatically improved. In addition, whether the output result of the machine learning model is valid or not is summarized in the number of components (number of images) included in the column of the positive judgment.

[0105] In this way, by displaying the second confusion matrix 107i in the aforementioned AI classification report (FIG. 8), humans can accurately and easily evaluate the output results of the machine learning model even in the case of multiple classes.

[0106] (Separation graph) 10 is a diagram showing an example of a separation degree graph. The horizontal axis of the separation degree graph 109 shows the anomaly score. As mentioned above, the anomaly score is lower for good images and higher for defective images. The vertical axis of the separation degree graph 109 shows the cumulative frequency. Therefore, the separation degree graph 109 displays the frequency distribution of both good images and defective images in a graphical format.

[0107] In the example shown in this figure, the region of the good product image (the hatched region marked OK in the figure) and the region of the defective product image (the hatched region marked NG in the figure) are clearly separated. In such cases, it can be determined that the inference performance of the machine learning model is sufficient and that a desirable output result has been obtained. On the other hand, although not shown separately, if there is overlap between the region of the good product image and the region of the defective product image, it is considered that the inference performance of the machine learning model is insufficient, and further machine learning (parameter adjustment) is necessary.

[0108] The AI ​​detection report (FIGS. 5 to 7) and the AI ​​classification report (FIG. 8) described above may each include a separation degree graph 109.

[0109] <Other> In addition to the above-described embodiments, the various technical features disclosed in this specification can be modified in various ways without departing from the spirit of the technical creation.

[0110] For example, in the above embodiment, an example is given of a configuration in which an image to be displayed in each of the training image display area 107 and the verification image display area 108 can be selected by clicking on an element 104a included in the first confusion matrix 104 specific to the visual inspection device 1. However, the function itself for selecting an image to be displayed in each of the AI ​​detection report (FIGS. 5 to 7) and the AI ​​classification report (FIG. 8) is novel, and the image selection method is in no way limited to the above embodiment.

[0111] In other words, the above-described embodiments should be considered to be illustrative in all respects and not restrictive. The technical scope of the present invention is defined by the claims, and it should be understood that all modifications within the meaning and scope of the claims are included. [Explanation of symbols]

[0112] 1. Visual inspection equipment (image inspection equipment) 2. Control Unit 3 Imaging unit 4 Display device (display section) 41 Touch Panel 5. Personal Computers 51 keyboard 52 Mouse 13 Main board 13a processor 13b Memory 14 Camera module (imaging unit) 141 AF motor 142 Imaging board 143 CMOS sensor 15 Lighting module (lighting section) 151 LED (light emitting diode) 152 LED drivers 16 Connector board 161 Power Interface 17 Communication board 18 Power supply board 181 AF motor driver 19 Storage device (storage unit) 80 Program Files 90 Storage medium 100 User Interface Screens 101 Creation date display area 102 Image number display area 103 Results display area 104 First confusion matrix 104a Component 105 Image size setting area 106 Display mode switching area 107 Learning image display area 107a Normal image 107b Checkbox 107c File name 107d Extraction display button 107e Display reset button 107f Heatmap 107g Settings / Results Label 107h Image type label 107i Second confusion matrix 108 Verification image display area 108a Normal image 108b Checkbox 108c File name 108d Extraction display button 108e Display reset button 109 Separation graph

Claims

1. An image inspection device that inspects inspection image data using a model in which parameters are updated by machine learning based on training image data presented by a user, a report output unit that outputs a report display including a first confusion matrix based on target image data, which is at least one of the training image data and the test image data, and an output result of the model for the target image data; a receiving unit that receives a user operation specifying components to be included in the first confusion matrix; Equipped with The report output unit outputs the report display in which the target image data corresponding to the component specified by the user operation is selected.

2. The image inspection device according to claim 1 , wherein the report display includes a display indicating a selection state of the target image data, and the report output unit changes the selection state depending on the component specified by the user operation.

3. The image inspection device according to claim 1 , wherein the first confusion matrix displays the number of pieces of target image data corresponding to each element of the matrix among the target image data.

4. 2. The image inspection device according to claim 1, wherein the model includes a step of calculating the degree of abnormality of pixels constituting an image, and the report output unit displays the target image data and an abnormality map generated based on the degree of abnormality without superimposing them on each other in the report display.

5. 2. The image inspection device according to claim 1, wherein the model includes a step of classifying images, and the report output unit displays, in the report display, a second confusion matrix in which target image data corresponding to components constituting the matrix are arranged in an area displaying the components.

6. The image inspection device of claim 1 , wherein the model includes a parameter-fixed portion.

Citation Information

Patent Citations

  • Setting device and setting method

    JP2023077051A

  • Appearance inspection device and appearance inspection method

    JP2023077054A