Image processing device, image processing method, and program

The image processing device addresses the limitation of users not being able to use their self-created AI models by incorporating a flow creation unit that allows users to select and order processing items, including AI processing items, enabling the utilization of user-created AI models for image processing.

WO2025094986A1PCT designated stage expired Publication Date: 2025-05-08OMRON CORP +1

Patent Information

Application Number
PCT/JP2024/038667
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-02
Filing Date
2024-10-30
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Users of image processing devices are unable to utilize AI models created by themselves due to lack of knowledge about the internal processing of the devices.

Method used

An image processing device with a flow creation unit that allows users to select and order processing items, including AI processing items that can acquire and utilize externally created AI models, enabling users to perform image processing using their own AI models.

Benefits of technology

Enables users to execute image processing flows that include AI processing items, allowing for the use of user-created AI models, thereby enhancing user flexibility and control over the image processing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024038667_08052025_PF_FP_ABST
    Figure JP2024038667_08052025_PF_FP_ABST
Patent Text Reader

Abstract

This image processing device comprises: a flow creation unit that creates an image processing flow by selecting one or more processing items to be executed from among a plurality of processing items and setting the order of execution of the one or more processing items; and an image processing unit that executes the image processing flow. The plurality of processing items include one or more AI processing items. Each of the one or more AI processing items defines: acquiring an AI model created externally; and outputting at least one of an inference result obtained by inputting a target image to the AI model and output information generated from the inference result.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device, image processing method and program

[0001] The present disclosure relates to an image processing device, an image processing method, and a program.

[0002] In product manufacturing, technologies have been developed that capture images of objects such as parts, intermediate products, or finished products and automatically determine attributes related to the object's appearance based on the captured images. Specifically, the attributes of an object are determined using a model obtained by performing machine learning using multiple images of the object whose attributes are known.

[0003] Japanese Patent Laid-Open Publication No. 2022-136563 (Patent Document 1) discloses an image processing device including a learning unit and a classification unit. The learning unit generates a model by performing machine learning using feature amounts of training images and label data indicating the classification of the training images. The classification unit inputs feature amounts of an image to be measured into the model and outputs a classification of the image to be measured from the model.

[0004] Japanese Patent Application Laid-Open No. 2022-136563

[0005] Mike Van Ness, Madeleine Udell, "CDF Normalization for Controlling the Distribution of Hidden Nodes", I (Still) Can't Believe It's Not Better Workshop at NeurIPS 2021 (October 19, 2021)

[0006] In the technology described in Patent Document 1, a model is generated by a learning unit included in the image processing device. The learning unit generates the model according to a learning algorithm adopted by the company that provides the image processing device.

[0007] In recent years, the so-called "democratization of AI (Artificial Intelligence)" has progressed. As a result, users of image processing devices can create their own AI models (trained models). However, because users of image processing devices do not know the details of the internal processing of the image processing devices, they cannot use their own AI models instead of models generated by the learning unit of the image processing devices.

[0008] The present disclosure has been made in consideration of the above-described situation, and its purpose is to provide an image processing device, an image processing method, and a program that are capable of performing image processing using an AI model that a user has created independently.

[0009] An image processing device according to one aspect of the present disclosure includes a flow creation unit and an image processing unit. The flow creation unit creates an image processing flow by selecting one or more processing items to be executed from a plurality of processing items and setting an execution order for the one or more processing items. The image processing unit executes the image processing flow. The plurality of processing items include at least one AI processing item. Each of the at least one AI processing item defines acquisition of an externally created AI model and output of at least one of an inference result obtained by inputting a target image into the AI ​​model and output information generated from the inference result.

[0010] According to this disclosure, the image processing device can perform image processing using an AI model created by the user.

[0011] In the above disclosure, each of the at least one AI processing item defines: obtaining a first parameter indicating an input image format of the AI ​​model; converting a target image to satisfy the input image format using the first parameter; and inputting the converted target image into the AI ​​model. According to this disclosure, the target image in an appropriate format is input into the AI ​​model.

[0012] In the above disclosure, the at least one AI processing item includes a first AI processing item using a first AI model that outputs an output image in which each pixel has a floating-point value as an inference result. The first AI processing item defines performing one or more image processes on the output image and outputting the output image after the one or more image processes as output information. The one or more image processes include at least one of a first process that converts the value of each pixel to an integer and a second process that normalizes the value of each pixel.

[0013] According to this disclosure, the output information output by executing the first AI processing item can be easily used in subsequent processing.

[0014] In the above disclosure, the value of each pixel in the output image represents the probability of belonging to a specific class, and the first AI processing item defines obtaining a second parameter representing the difference between a threshold optimized for determining classification into the specific class and a predetermined value, and performing a second processing using the second parameter.

[0015] According to this disclosure, the predetermined value can be used as an optimal threshold for determining classification into a particular class in the normalized output image.

[0016] In the above disclosure, at least one AI processing item includes a second AI processing item using a second AI model that performs an autoencoder on a target image to output a restored image as an inference result. The second AI processing item defines generating a difference image between the target image and the restored image, and outputting the difference image or a processed image obtained by performing one or more image processes on the difference image as output information.

[0017] According to this disclosure, the difference image or the processed image can be used to detect abnormalities in the target image.

[0018] In the above disclosure, the one or more image processes include at least one of a first process that converts the value of each pixel into an integer and a second process that normalizes the value of each pixel.

[0019] According to this disclosure, the output information output by executing the second AI processing item becomes easier to use in subsequent processing.

[0020] In the above disclosure, the second AI processing item defines obtaining a second parameter representing the difference between a threshold optimized for detecting abnormal portions from the difference image and a predetermined value, and performing a second processing using the second parameter.

[0021] According to this disclosure, the predetermined value can be used as an optimal threshold for detecting abnormalities in the normalized difference image.

[0022] In the above disclosure, the flow creation unit accepts designation of an AI model to be used in the target AI processing item in response to a target AI processing item being selected as one or more processing items among at least one AI processing item, and determines whether the designated AI model is compatible with the target AI processing item.

[0023] According to this disclosure, by checking the judgment results by the flow creation unit, the user can recognize that they have mistakenly specified an AI model that is not suitable for the AI ​​processing item.

[0024] In the above disclosure, at least one AI processing item includes a third AI processing item using a third AI model that outputs an output image as an inference result, the third AI processing item defining: generating a first processed image obtained by performing a first image processing on the output image and a second processed image obtained by performing a second image processing on the output image; displaying the second processed image on a display; and outputting the first processed image to a subsequent processing item among the one or more processing items.

[0025] According to this disclosure, the image processing device can display a second processed image on a display that is easy for the user to see, and can output a first processed image that is suitable for a subsequent processing item to that subsequent processing item.

[0026] In the above disclosure, the image processing device further includes an update unit that updates the AI ​​model to a retrained AI model. The retrained AI model is obtained by retraining the AI ​​model using a target image. According to this disclosure, the image processing device can perform image processing using the latest AI model.

[0027] An image processing method according to one aspect of the present disclosure includes: a processor selecting one or more processing items to be executed from a plurality of processing items and setting an execution order of the one or more processing items to create an image processing flow; and the processor executing the image processing flow. The plurality of processing items include at least one AI processing item. Each of the at least one AI processing item defines acquiring an externally created AI model and outputting at least one of an inference result obtained by inputting a target image into the AI ​​model and output information generated from the inference result.

[0028] A program according to one aspect of the present disclosure causes a computer to execute the image processing method. Based on these disclosures, the image processing method and program can also perform image processing using an AI model created by a user.

[0029] According to the present disclosure, an image processing device, an image processing method, or a program can perform image processing using an AI model created independently by a user.

[0030] 9 is a diagram illustrating an example of an image processing device according to an embodiment. FIG. 10 is a diagram illustrating an example of a procedure for creating an AI model. FIG. 11 is a diagram schematically illustrating the structure of an AI model. FIG. 12 is a diagram illustrating an example of an AI model. FIG. 13 is a diagram illustrating another example of an AI model. FIG. 14 is a diagram illustrating yet another example of an AI model. FIG. 15 is a schematic diagram illustrating an example of a hardware configuration of the image processing device shown in FIG. 1. FIG. 16 is a block diagram illustrating an example of a functional configuration of the image processing device shown in FIG. 1. FIG. 17 is a diagram illustrating an example of a screen provided by a flow creation unit. FIG. 18 is a diagram illustrating an example of an image processing flow created using the screen shown in FIG. 19 is a diagram illustrating an example of a screen for setting processing conditions for AI processing items. FIG. 19 is a diagram illustrating another example of a screen for setting processing conditions for AI processing items. FIG. 10 is a diagram illustrating a first example of the internal configuration of an item execution unit. FIG. 11 is a diagram illustrating a second example of the internal configuration of the item execution unit. FIG. 12 is a diagram illustrating a third example of the internal configuration of the item execution unit. FIG. 13 is a diagram illustrating a fourth example of the internal configuration of the item execution unit. FIG. 14 is a diagram illustrating a fifth example of the internal configuration of the item execution unit. FIG. 15 is a diagram illustrating an example of an operation procedure of an image processing device. FIG. 16 is a diagram illustrating a modified example of an item execution unit 14C. FIG. 17 is a block diagram illustrating a functional configuration of an image processing device according to a modified example.

[0031] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described in detail with reference to the accompanying drawings, in which the same or corresponding parts in the drawings are designated by the same reference numerals and the description thereof will not be repeated.

[0032] §1 Application Example First, an example of a situation in which the present invention is applied will be described using Fig. 1. Fig. 1 is a diagram showing an example of an image processing device according to an embodiment. As shown in Fig. 1, a camera 200 is externally attached to an image processing device 100 according to an embodiment.

[0033] The image processing device 100 performs an image processing flow that combines one or more processing items selected from a plurality of predetermined processing items.

[0034] A user M of the image processing device 100 creates his or her own AI model 15 using a learning device 300. The learning device 300 is, for example, a server that provides a known cloud service. Alternatively, the learning device 300 may be a device owned by the user M of the image processing device 100. The learning device 300 includes a learning tool 302 and an evaluation tool 304.

[0035] The learning tool 302 generates an AI model by performing machine learning using a training dataset including learning images and correct answer data.

[0036] The correct answer data indicates correct answer information corresponding to the training image. The type of information indicated by the correct answer data depends on the type of AI model to be generated. For example, when generating an AI model for region detection (segmentation), the correct answer data indicates a region in the training image (e.g., a region where a defect exists) and an attribute of the region (e.g., "defect").

[0037] The evaluation tool 304 evaluates the performance of the AI ​​model generated by the learning tool 302. For example, the evaluation tool 304 evaluates the performance of the AI ​​model by comparing an inference result obtained by inputting a plurality of evaluation images into the AI ​​model with correct answer data corresponding to the plurality of evaluation images.

[0038] A user M of the image processing device 100 acquires an AI model that he or she has created independently using the learning device 300.

[0039] The image processing device 100 includes a flow creation unit 11 and an image processing unit 12. The flow creation unit 11 creates an image processing flow by selecting one or more processing items to be executed from a plurality of processing items 44 and setting the execution order of the selected one or more processing items. The flow creation unit 11 creates an image processing flow in accordance with input from a user M. The image processing unit 12 executes the created image processing flow.

[0040] The plurality of processing items 44 includes at least one processing item (hereinafter referred to as a "general-purpose processing item 42") that does not use an external AI model, and at least one AI processing item 40 that uses an external AI model.

[0041] Each of the at least one AI processing item 40 defines the acquisition of an externally created AI model and the output of at least one of an inference result obtained by inputting an image of a measurement target (hereinafter referred to as a "target image") into the AI ​​model and output information generated from the inference result. The externally created AI model includes an AI model created by the learning device 300. The target image is typically acquired from a camera externally attached to the image processing device 100. Alternatively, the target image is output from a processing item in a previous stage.

[0042] This allows the user M of the image processing device 100 to cause the image processing device 100 to execute an image processing flow that executes one or more selected processing items 44 in a desired order, including an AI processing item that uses an AI model that the user M has created. In this way, the image processing device 100 can perform image processing using an AI model that the user M has created.

[0043] §2 Specific Examples <Examples of AI Models> In product manufacturing, AI models can be used for visual inspection of objects such as parts, intermediate products, or finished products. Therefore, the following describes AI models used for visual inspection on inspection lines. However, the use of AI models is not limited to visual inspection.

[0044] Fig. 2 is a diagram showing an example of a procedure for creating an AI model. Creating an AI model may require machine learning skills. Therefore, in the example shown in Fig. 2, an inspector and an AI model creator cooperate with each other to create the AI ​​model.

[0045] First, an inspector accumulates images of objects on the inspection line (Step S1). The accumulated images are used as learning images for generating an AI model or as evaluation images for evaluating the AI ​​model.

[0046] Next, the inspector performs annotation for each of the stored images (step S2). Annotation is a process of adding correct answer information to an image. That is, the inspector creates correct answer data indicating the correct answer information for each image. This generates a data set including the image and the correct answer data.

[0047] Next, the inspector provides the plurality of data sets to the AI ​​model creator (step S3). The AI ​​model creator uses some of the plurality of data sets as training data sets and the rest as evaluation data sets.

[0048] Next, the AI ​​model creator creates an AI model by applying multiple training datasets to the learning device 300 (step S4). The learning tool 302 included in the learning device 300 uses a known machine learning library such as TensorFlow (registered trademark) or Pytorch (registered trademark).

[0049] Generally, an AI model receives input images in which the brightness values ​​of each pixel are expressed in floating-point format (e.g., float or double). On the other hand, the brightness values ​​of each pixel in images accumulated on an inspection line are generally expressed in integer format (e.g., uchar or byte). Therefore, the learning tool 302 converts the learning images included in the training dataset from integer format to floating-point format.

[0050] Additionally, the training tool 302 typically performs a standardization process on the training images contained in the multiple training data sets by dividing, for each pixel, the deviation of the luminance value from the mean value μ by the standard deviation σ.

[0051] FIG. 3 is a diagram schematically illustrating the structure of an AI model. FIG. 3 illustrates an AI model configured by a neural network. The neural network includes an input layer, an intermediate layer, and an output layer. However, the structure of the neural network is not limited to this example and may be determined appropriately depending on the embodiment. For example, the number of intermediate layers is not limited to one, and may be two or more.

[0052] The number of neurons (nodes) in the input layer corresponds to the number of pixels in the input image. The number of neurons in the intermediate layer corresponds to the number of dimensions of the feature. The number of neurons in the output layer corresponds to the inference result.

[0053] In a neural network, neurons in adjacent layers are appropriately connected to each other, and a weight (connection weight) is set for each connection. A threshold is set for each neuron, and the output of each neuron is basically determined by whether the sum of the products of each input and each weight exceeds the threshold. The connection weights between neurons and the threshold of each neuron are examples of calculation parameters that define an AI model.

[0054] The learning tool 302 prepares a neural network that constitutes an AI model. The learning tool 302 executes a learning process for the neural network using multiple training data sets.

[0055] Because the training dataset includes correct data, the learning tool 302 can perform a neural network learning process using supervised learning. For example, in a first step, the learning tool 302 inputs each training image into the input layer and performs neural network calculations. In a second step, the learning tool 302 calculates the error between the output value obtained from the output layer and the correct data corresponding to each training image based on a loss function. In a third step, the learning tool 302 calculates the neural network's calculation parameters, i.e., the errors in the connection weights between neurons and the thresholds of each neuron, using the calculated output value errors, for example, by backpropagation. In a fourth step, the learning tool 302 updates the values ​​of the connection weights between neurons and the thresholds of each neuron based on the calculated errors. By repeating the first to fourth steps, the learning tool 302 adjusts the values ​​of the calculation parameters so as to reduce the sum of the errors between the output values ​​output from the output layer and the correct data.

[0056] The learning tool 302 may perform neural network learning using unsupervised learning. For example, the learning tool 302 may create an AI model that performs autoencoding using multiple training images of defect-free objects. The AI ​​model encodes an input image into a low-dimensional latent representation and outputs a reconstructed image obtained by decoding the latent representation as an inference result.

[0057] Returning to FIG. 2, next, the evaluation tool 304 of the learning device 300 evaluates the AI ​​model using multiple evaluation datasets (step S5).

[0058] The AI ​​model applied to classification outputs the probability of belonging to a specific class (also referred to as a "score"). A specific class is, for example, a class in which "defects" exist. The AI ​​model may output a score for an image, or may output a score for each pixel in the image. The score generally ranges from 0 to 1. Depending on the result of comparing the score with a threshold, the classification of the image or pixel into a specific class is determined. In step S5, the evaluation tool 304 may calculate a threshold (hereinafter referred to as the "optimal threshold Th1") optimized for classification into a specific class. For example, the evaluation tool 304 calculates the threshold that maximizes the F-measure when the threshold is varied between 0 and 1 as the optimal threshold Th1. The F-measure is the harmonic mean of precision and recall. Furthermore, the AI ​​model may output the processing results of the network within the AI ​​model. For example, the AI ​​model may use a Gradient-weighted Class Activation Mapping (Grad-CAM) technique to output information indicating where in the image it is focusing its attention in terms of classifying it into a particular class.

[0059] AI models that perform autoencoder are generally used to detect abnormalities in objects. Specifically, a difference image between a restored image, which is the inference result of the AI ​​model, and an input image represents an abnormal portion of the object. Therefore, the presence or absence of a defect in the object is determined based on the difference image between the restored image and the input image. The evaluation tool 304 inputs each evaluation image included in multiple evaluation datasets into the AI ​​model and generates a difference image between the restored image and the evaluation image. In the difference image, pixels having values ​​exceeding a threshold are determined to be pixels that depict a defect. The evaluation tool 304 compares the difference image with the correct data to calculate a threshold (hereinafter referred to as the "optimal threshold Th2") optimized for detecting abnormal portions.

[0060] Next, the AI ​​model creator provides the AI ​​model created by the learning device 300 to the inspector (step S6).

[0061] Furthermore, the AI ​​model creator acquires a parameter set for the created AI model from the learning device 300 and provides the parameter set to the inspector (step S7).

[0062] The parameter set includes a first parameter indicating the input image format of the AI ​​model. For example, the first parameter includes parameters indicating the size (image height and image width) of the input image of the AI ​​model. The first parameter includes a parameter indicating the type of luminance value (e.g., float type or double type) of each pixel of the input image of the AI ​​model. Furthermore, the first parameter may include a mean value μ and a standard deviation σ used in the standardization process by the learning tool 302.

[0063] The parameter set may include a second parameter for normalizing an image that is the inference result of the AI ​​model. The second parameter is also referred to as a "normalization parameter." One example of the second parameter represents the difference between an optimal threshold Th1 optimized for classification into a specific class and a predetermined value. The predetermined value is typically the median (0.5) of the range of values ​​that the probability (score) of belonging to a specific class can take. Another example of the second parameter represents the difference between an optimal threshold Th2 optimized for detecting an abnormal portion from a difference image and a predetermined value. The predetermined value is the median (0.5) of the range of values ​​that each pixel in the difference image can take.

[0064] 4 is a diagram showing an example of an AI model. Fig. 4 shows an AI model 15A that performs region detection (segmentation). Each training data set 90A includes a learning image 92A containing an object 91 and correct answer data 94A. The correct answer data 94A indicates a defect region 93 in the learning image 92A.

[0065] The AI ​​model 15A, which has been created by machine learning using multiple training data sets 90A, receives an input image 96A and outputs an output image 98A as an inference result. The value of each pixel in the output image 98A indicates the probability (score) of belonging to the "defect" class (in other words, the probability that a defect exists).

[0066] FIG. 5 is a diagram illustrating another example of an AI model. FIG. 5 illustrates an AI model 15B that reads character strings in a user-specific font. Each training data set 90B includes a training image 92B depicting a character string in a user-specific font, and correct answer data 94B. The correct answer data 94B represents the character string depicted in the training image 92B.

[0067] When an AI model 15B created by machine learning using multiple training data sets 90B receives an input image 96B, it outputs a character string 98B that appears in the input image 96B as an inference result.

[0068] FIG. 6 is a diagram showing yet another example of an AI model. FIG. 6 shows an AI model 15F that performs object detection. Each training data set 90F includes a learning image 92F depicting an object to be detected, and correct answer data 94F. In FIG. 6, codes (barcodes and two-dimensional codes) are shown as examples of objects to be detected. Note that the objects to be detected are not limited to codes. For example, the object to be detected may be a specific part of a product for determining an inspection area for the product. The correct answer data 94F indicates the center coordinates, width, and height of a partial region (bounding box) surrounding the object in the learning image 92F.

[0069] When an AI model 15F created by machine learning using multiple training datasets 90F receives an input image 96F, it outputs as an inference result the center coordinates, width, and height of a bounding box surrounding an object appearing in the input image 96F, the type label of the object, and the probability (score) of belonging to that type label.

[0070] <Example of Hardware Configuration of Image Processing Device> The image processing device 100 is typically a computer having a general-purpose architecture, and executes a pre-installed program (instruction code) to perform image processing according to this embodiment. Such a program is typically distributed in a state stored on various recording media, or is installed in the image processing device 100 via a network, etc.

[0071] When using such a general-purpose computer, an OS (Operating System) for executing basic computer processing may be installed in addition to the application for executing the image processing according to the present embodiment. In this case, the program according to the present embodiment may execute processing by calling necessary modules from among program modules provided as part of the OS in a predetermined sequence at a predetermined timing. In other words, the program according to the present embodiment itself may not include the above-mentioned modules, and may execute processing in cooperation with the OS. The program according to the present embodiment may also be in a form that does not include some of these modules.

[0072] Furthermore, the program according to the present embodiment may be provided by being incorporated into a part of another program. In this case, the program itself does not include the modules included in the other program to be combined as described above, and executes processing in cooperation with the other program. In other words, the program according to the present embodiment may be in a form incorporated into such other program. Note that some or all of the functions provided by the execution of the program may be implemented as dedicated hardware circuits.

[0073] Fig. 7 is a schematic diagram showing an example of the hardware configuration of the image processing device shown in Fig. 1. As shown in Fig. 7, the image processing device 100 includes a CPU (Central Processing Unit) 110, which is an example of a processor, a main memory 112, a hard disk 114, a camera interface 116, an input interface 118, a display controller 120, a communication interface 124, and a data reader / writer 126. These components are connected to each other via a bus 128 so as to be able to communicate data with each other.

[0074] The CPU 110 loads the programs 115 installed on the hard disk 114 into the main memory 112 and executes them in a predetermined order to perform various calculations. The main memory 112 typically includes a volatile storage device such as a dynamic random access memory (DRAM), and stores images acquired from the camera 200 in addition to the programs 115 read from the hard disk 114. The hard disk 114 also stores various data, as will be described later. Note that a semiconductor storage device such as a flash memory may be used in addition to or instead of the hard disk 114.

[0075] The camera interface 116 mediates data transmission between the CPU 110 and the camera 200. That is, the camera interface 116 is connected to the camera 200. The camera interface 116 issues an image capture command to the camera 200 in accordance with an internal command generated by the CPU 110. The image capture command may be output to the camera 200 in response to a detection signal from a photoelectric sensor. Alternatively, the image capture command may be output to the camera 200 in response to an external command from a programmable logic controller (PLC).

[0076] The camera interface 116 includes an image buffer 116a for temporarily storing images received from the camera 200. The camera 200 includes multiple image sensors. Each image sensor outputs a grayscale value for a corresponding pixel. The grayscale value is expressed as integer digital data, such as 8-bit, 10-bit, or 12-bit, depending on the image sensor. The camera 200 includes either a monochrome image sensor or a color image sensor. The color image sensor may include a single-chip image sensor using a Bayer filter. The image sensor may have a built-in A / D converter. In this case, digital data representing the grayscale value is output directly from the image sensor. The A / D converter may be external to the image sensor. The image stored in the image buffer 116a may be in a format in which pixel grayscale values ​​are arranged in raster scan order, for example. When receiving a color image from the camera 200, the image buffer 116a may store the image after demosaicing. Alternatively, image buffer 116a may store an image after performing some kind of compression processing on the image received from camera 200. For example, if the imaging element outputs a 10-bit gradation value, image buffer 116a may store the upper 8 bits of data for each pixel, excluding the lower 2 bits.

[0077] 7, the camera 200 is externally attached to the image processing device 100. However, the camera 200 may be built into the image processing device 100.

[0078] The input interface 118 mediates data transmission between the CPU 110 and the input device 160. That is, the input interface 118 accepts input information input to the input device 160 by the user.

[0079] The display controller 120 is connected to the display 150 and controls the screen of the display 150 so as to notify the user of the processing results of the CPU 110 and the like.

[0080] The communication interface 124 mediates data transmission between the CPU 110 and an external device (for example, a PC), and is typically implemented by an Ethernet (registered trademark) or a Universal Serial Bus (USB).

[0081] Data reader / writer 126 mediates data transmission between CPU 110 and memory card 106, which is a recording medium. That is, memory card 106 stores and distributes programs to be executed by image processing device 100, and data reader / writer 126 reads the programs from memory card 106. In addition, data reader / writer 126 writes images received from camera 200 and / or processing results in image processing device 100 to memory card 106 in response to internal commands from CPU 110. Note that memory card 106 may be a general-purpose semiconductor storage device such as an SD (Secure Digital), a magnetic storage medium such as a flexible disk, or an optical storage medium such as a CD-ROM (Compact Disk Read Only Memory).

[0082] <Example of functional configuration of image processing device> Fig. 8 is a block diagram showing an example of the functional configuration of the image processing device shown in Fig. 1. As shown in Fig. 8, the image processing device 100 includes a flow creation unit 11, an image processing unit 12, and a storage unit 13. The flow creation unit 11 and the image processing unit 12 are realized by the CPU 110 shown in Fig. 7 executing a program 115. The storage unit 13 is realized by the main memory 112 or the hard disk 114.

[0083] The storage unit 13 stores the AI ​​model 15 created using the external learning device 300 and the parameter set 16 related to the AI ​​model 15. For example, the inspector stores the AI ​​model 15 and the parameter set 16 received from the AI ​​model creator in the storage unit 13 according to the procedure shown in FIG.

[0084] The flow creation unit 11 creates an image processing flow by setting one or more process items to be executed from among a plurality of predetermined process items 44 and the execution order of the one or more process items.

[0085] As described above, one or more AI processing items 40 that perform image processing using an AI model are predefined in the image processing device 100. This allows the user of the image processing device 100 to appropriately select an AI processing item that is suitable for the desired appearance inspection.

[0086] Furthermore, the flow creation unit 11 sets various conditions for the process items to be executed in response to inputs to the input device 160 .

[0087] The image processing unit 12 executes the image processing flow created by the flow creation unit 11. The image processing unit 12 includes an item execution unit 14 corresponding to each processing item to be executed. The image processing unit 12 includes, for example, the OpenVINO (registered trademark) tool to perform inference using an AI model. The OpenVINO tool is provided by Intel and is an inference engine that uses an AI model.

[0088] <Example of Processing by Flow Creation Unit> Fig. 9 is a diagram showing an example of a screen provided by the flow creation unit 11. A screen 30 shown in Fig. 9 is provided by the flow creation unit 11 and displayed on the display 150.

[0089] The screen 30 includes a set item display area 32, a process item selection area 34, a camera image display area 36, ​​an insert / add process item button 38, and an execution order change button 39. The set item display area 32 graphically displays the contents of the currently set process flow.

[0090] Icons representing a plurality of pre-defined processing items 44 are displayed together with their names in the processing item selection area 34. As shown in Fig. 9, the list displayed in the processing item selection area 34 includes a plurality of AI processing items 40 and a plurality of general-purpose processing items 42 that do not use an external AI model.

[0091] The AI ​​processing items 40 include, for example, AI processing items 40a, 40c, 40d, 40e, and 40f. The AI ​​processing item 40c uses an AI model 15C that performs classification. That is, the AI ​​model 15C outputs a probability (score) of belonging to each class based on the characteristics of the input image.

[0092] The AI ​​processing item 40d uses the AI ​​model 15D that performs autoencoding. That is, the AI ​​model 15D encodes an input image into a low-dimensional latent representation, and outputs a reconstructed image obtained by decoding the latent representation as an inference result.

[0093] The AI ​​processing item 40e uses the AI ​​model 15E that performs object recognition. That is, the AI ​​model 15E outputs, as inference results, area information indicating a rectangular area in which an object appears in an input image and the probability (score) that the object moving to the rectangular area belongs to each class.

[0094] The general-purpose processing items 42 include known image processing items, for example, a general-purpose processing item 42a that performs labeling.

[0095] A user of the image processing device 100 selects a processing item required for the desired image processing in the processing item selection area 34 of the screen 30 ((1) Select Processing Item). Furthermore, the user selects the position (order) at which the selected processing item should be added in the set item display area 32 ((2) Select Addition Position). The user adds a processing item to the set item display area 32 by selecting the Insert / Add Processing Item button 38 ((3) Press the Insert / Add Button) ((4) The Processing Item is Added). The user can create a desired image processing flow by repeating this process as needed. Furthermore, during or after creating the processing settings, the user can change the execution order as needed by selecting a processing item in the set item display area 32 and then selecting the Change Execution Order button 39.

[0096] Fig. 10 is a diagram showing an example of an image processing flow created using the screen shown in Fig. 9. The image processing flow 50 shown in Fig. 10 includes a general-purpose processing item 42b for acquiring a target image from the camera 200, an AI processing item 40a, and a general-purpose processing item 42a for labeling.

[0097] The flow creator 11 sets processing conditions for each of the one or more processing items selected as the processing items to be executed.

[0098] 11 is a diagram showing an example of a screen for setting processing conditions for AI processing items. The screen 60 shown in FIG. 11 is used to set the AI ​​model 15. The screen 60 includes setting areas 61 and 62, display areas 63 and 64, and a button 65.

[0099] The setting area 61 is used to set a hardware device that executes the inference process using the AI ​​model 15. The setting area 61 includes an input field 61a for inputting a hardware device. In response to an operation of the input field 61a, the flow creation unit 11 displays a pull-down menu of hardware devices that can execute the inference process and prompts the user to select a device from the pull-down menu. The flow creation unit 11 sets the selected device as the hardware device that executes the inference process using the AI ​​model 15.

[0100] The setting area 62 is used to specify the AI ​​model 15. The setting area 62 includes a radio button 62a and input fields 62b and 62c.

[0101] As described above, the image processing unit 12 includes, for example, the OpenVINO (registered trademark) tool. AI model formats that can be executed in the OpenVINO tool include the IR (Intermediate Representation) format and the ONNX (Open Neural Network Exchange) format. Therefore, the user converts the format of the AI ​​model 15 that the user independently created using the learning device 300 into the IR format or the ONNX format in advance. The radio button 62a is used to select the format of the AI ​​model 15. Note that if the format of the AI ​​model created using the learning device 300 matches or is compatible with the format of the AI ​​model that the image processing unit 12 can read, the user does not need to convert the format of the AI ​​model 15.

[0102] The input fields 62b and 62c accept the specification of the AI ​​model 15. Specifically, the file path of the AI ​​model 15 is input into the input fields 62b and 62c. The IR format AI model 15 is composed of an xml format file and a bin format file. The input field 62b is used to specify the xml format file. The input field 62c is used to specify the bin format file.

[0103] The display area 63 is used to display the input image conditions of the AI ​​model that can be used in the corresponding AI processing item and the output (inference result) conditions of the AI ​​model. The input image format includes the batch number "n", the number of channels "c", the image height "h", and the image width "w".

[0104] The conditions for the inference result of the AI ​​model that outputs an image include the number of batches "n", the number of channels "c", the image height "h", and the image width "w". The AI ​​model that outputs an image includes, for example, the AI ​​model 15A shown in FIG.

[0105] The conditions for the inference results of the AI ​​model that does not output images include the number of batches “n” and the number of channels “c.” The AI ​​model that does not output images includes, for example, the AI ​​model 15B shown in FIG.

[0106] The display area 64 is used to display the time required to load the AI ​​model 15.

[0107] The button 65 is used to start reading the AI ​​model 15. In response to clicking the button 65, the flow creation unit 11 starts reading the AI ​​model 15 according to the file paths entered in the input fields 62b and 62c.

[0108] The format of the AI ​​model 15 may vary depending on the AI ​​processing item. For example, the AI ​​processing item 40a corresponding to the image processing "segmentation" uses an AI model 15 (e.g., AI model 15A shown in FIG. 4) that outputs an output image in which the value of each pixel indicates the probability of belonging to a specific class. Therefore, the AI ​​processing item 40a cannot use an AI model 15 that does not output an image (e.g., AI model 15B shown in FIG. 5). Therefore, the flow creation unit 11 may determine whether the AI ​​model specified in the input fields 62b and 62c is compatible with the AI ​​processing item. Specifically, the flow creation unit 11 compares the format of the output data of the AI ​​model specified in the input fields 62b and 62c with the conditions for the output (inference result) of the AI ​​model usable in the AI ​​processing item (hereinafter referred to as "output conditions"). The flow creation unit 11 determines that the specified AI model is compatible with the AI ​​processing item if the format of the output data of the specified AI model satisfies the output conditions. If the format of the output data of the specified AI model does not satisfy the output conditions, the flow creation unit 11 determines that the specified AI model does not conform to the AI ​​processing item.

[0109] The flow creation unit 11 may output an error message when it determines that the specified AI model does not match the AI ​​processing item.

[0110] 12 is a diagram showing another example of a screen for setting processing conditions for an AI processing item. A screen 60A shown in FIG. 12 is used to set conditions (hereinafter referred to as "determination conditions") for determining whether the appearance of an object is good or bad based on the inference results of AI model 15. The determination conditions are set, for example, for AI processing item 40c that uses AI model 15C that performs classification.

[0111] 12, the screen 60A includes an input field 66 for inputting a lower limit value and an input field 67 for inputting an upper limit value for each class. As described above, the AI ​​model 15C outputs the belonging probability of each class based on the characteristics of the input image. The belonging probability has a value between 0 and 1. Therefore, values ​​between 0 and 1 are input into the input fields 66 and 67. The user inputs the lower limit value and the upper limit value of the range that the belonging probability (score) can take when the appearance of the object is good into the input fields 66 and 67, respectively.

[0112] Alternatively, the flow creation unit 11 may set the conditions that a character string can take as judgment conditions for an AI processing item that uses an AI model 15B (see Figure 5) that outputs a character string as an inference result in response to user input.

[0113] Furthermore, the flow creation unit 11 sets a parameter set 16 corresponding to the AI ​​model 15 set in the AI ​​processing item in accordance with the user's input. The user simply inputs the parameter set 16 provided in step S7 of FIG. 2 .

[0114] <Examples of Item Execution Unit> Next, first to fifth examples of the internal configuration of the item execution unit 14 will be described with reference to FIGS.

[0115] (First Example) Fig. 13 is a diagram showing a first example of the internal configuration of an item execution unit 14A that executes AI processing items. Fig. 13 shows the internal configuration of an item execution unit 14A.

[0116] 13, the item execution unit 14A includes an AI model reading unit 71, an image conversion unit 72, and an inference unit 73. The AI ​​model reading unit 71 and the inference unit 73 are realized by, for example, the OpenVINO (registered trademark) tool.

[0117] The AI ​​model reading unit 71 reads the file of a specified AI model 15. For example, the AI ​​model reading unit 71 reads an AI model 15A (see FIG. 4) that outputs an output image in which the value of each pixel indicates the probability of belonging to a specific class as an inference result. Alternatively, the AI ​​model reading unit 71 may read an AI model 15B (see FIG. 5) that outputs a character string appearing in an input image as an inference result. Alternatively, the AI ​​model reading unit 71 may read an AI model 15C that performs classification or an AI model 15E that performs object recognition. Alternatively, the AI ​​model reading unit 71 may read an AI model 15F (see FIG. 6) that outputs the center coordinates, width, and height of a bounding box surrounding an object appearing in an input image, the object type label, and a score as an inference result.

[0118] The AI ​​model reading unit 71 reads a file in an AI model format that can be executed in the OpenVINO tool. However, if a file having a format different from the AI ​​model format that can be executed in the OpenVINO tool is specified, the AI ​​model reading unit 71 may convert the format of the specified file. The AI ​​model reading unit 71 may convert the format of the AI ​​model using a "model optimizer" provided by the OpenVINO tool.

[0119] The image conversion unit 72 converts the format of the target image into a format suitable for the AI ​​model 15 based on the first parameter included in the parameter set 16.

[0120] Specifically, the image conversion unit 72 converts the size (height and width) of the target image into the size indicated by the first parameter.

[0121] The image conversion unit 72 also converts the format of the value of each pixel of the target image. For example, the image conversion unit 72 converts the value of each pixel of the target image from an integer type (e.g., uchar type or byte type) to a floating-point type (e.g., float type or double type) indicated by the first parameter.

[0122] Furthermore, the image conversion unit 72 performs the same standardization process on the target image as the standardization process performed on the training images when training the AI ​​model 15. That is, the image conversion unit 72 performs the standardization process on the target image using the mean value μ and standard deviation σ included in the first parameter. As described above, the standardization process is a process of dividing the deviation of the luminance value of each pixel from the mean value μ by the standard deviation σ.

[0123] The inference unit 73 performs inference using the AI ​​model 15 read by the AI ​​model reading unit 71. The inference unit 73 inputs the target image converted by the image conversion unit 72 to the AI ​​model 15 and obtains an inference result. The item execution unit 14A outputs the inference result to the outside.

[0124] (Second Example) Fig. 14 is a diagram showing a second example of the internal configuration of an item execution unit. Fig. 14 shows the internal configuration of an item execution unit 14B that executes AI processing items. As shown in Fig. 14, the item execution unit 14B differs from the item execution unit 14A shown in Fig. 13 in that it includes a determination unit 74.

[0125] The determining unit 74 determines whether or not the inference result obtained by the inference unit 73 satisfies a determination condition. The determination condition is set in advance by the flow creating unit 11.

[0126] The item execution unit 14B outputs the determination result by the determination unit 74 to the outside. The determination result is an example of "output information" in the present disclosure. The item execution unit 14B may output the inference result obtained by the inference unit 73 to the outside in addition to the determination result.

[0127] When the AI ​​model reading unit 71 reads the AI ​​model 15A (see FIG. 4 ), for example, a condition that the probability of belonging to the "defect" class does not exceed a predetermined upper or lower limit is set as a judgment condition. Alternatively, a condition that the area of ​​the region where the probability of belonging to the "defect" class exceeds a first threshold is set as a judgment condition is set as a judgment condition. In this case, a user of the image processing device 100 can recognize whether the appearance of the object is good or bad by checking the judgment result.

[0128] Alternatively, when the AI ​​model reading unit 71 reads the AI ​​model 15B (see FIG. 5), for example, possible conditions for the character string are set as the judgment conditions. In this case, the user of the image processing device 100 can recognize whether or not there is an abnormality in the character string by checking the judgment result.

[0129] Alternatively, when the AI ​​model reading unit 71 reads the AI ​​model 15C that performs classification, for example, a condition that the probability of belonging to the "defect" class is less than a threshold is set as a judgment condition. In this case, the user of the image processing device 100 can recognize the presence or absence of a defect in the object by checking the judgment result.

[0130] (Third Example) Fig. 15 is a diagram showing a third example of the internal configuration of an item execution unit. Fig. 15 shows the internal configuration of an item execution unit 14C that executes AI processing items. As shown in Fig. 15, the item execution unit 14C differs from the item execution unit 14A shown in Fig. 13 in that it includes a normalization unit 75 and an image conversion unit 76.

[0131] The AI ​​model reading unit 71 of the item execution unit 14C reads the AI ​​model 15 that outputs an image as an inference result. For example, the AI ​​model reading unit 71 reads the AI ​​model 15A (see FIG. 4) that outputs an output image in which the value of each pixel indicates the probability that the pixel belongs to a specific class.

[0132] The normalization unit 75 performs normalization processing on the image that is the inference result obtained by the inference unit 73, based on the second parameter included in the parameter set 16. As described above, the second parameter represents the difference between the optimal threshold Th optimized for classification into a specific class and a predetermined value (typically 0.5).

[0133] The normalization unit 75 normalizes the value of each pixel of the image, for example, according to a method based on the Min-Max method. Specifically, the normalization unit 75 normalizes the value val of each pixel according to the following equation (1): val' = {(val - threshold) / (max - min)} + 0.5 (1) where threshold is a value represented by the second parameter. max is the maximum value of the pixel. min is the minimum value of the pixel. The pixel value indicates the probability of belonging, and can take a value between 0 and 1. Therefore, a fixed value of "1" may be used as (max - min).

[0134] By performing normalization according to the above formula (1), a predetermined value (typically 0.5) can be used as the optimal threshold for determining classification into a particular class in the normalized image.

[0135] In addition, the normalization unit 75 may perform normalization using a cumulative distribution function (CDF) (see "Mike Van Ness, Madeleine Udell, "CDF Normalization for Controlling the Distribution of Hidden Nodes", I (Still) Can't Believe It's Not Better Workshop at NeurIPS 2021 (October 19, 2021)" (Non-Patent Document 1)).

[0136] The image conversion unit 76 converts the value of each pixel of the image normalized by the normalization unit 75 from a floating-point type (e.g., float type or double type) to an integer type (e.g., uchar type or byte type). That is, the image conversion unit 76 converts the value of each pixel to an integer. In addition to converting to an integer type, the image conversion unit 76 also performs a magnification conversion from the range (0 to 1) changed by the above-mentioned standardization process to a range that can be represented as an image (e.g., 0 to 255).

[0137] The item execution unit 14C outputs to the outside the image (hereinafter referred to as "processed output image") converted by the image conversion unit 76. The processed output image is an example of "output information" in the present disclosure.

[0138] (Fourth Example) Fig. 16 is a diagram showing a fourth example of the internal configuration of an item execution unit. Fig. 16 shows the internal configuration of an item execution unit 14D that executes AI processing items. As shown in Fig. 16, the item execution unit 14D differs from the item execution unit 14C shown in Fig. 15 in that it includes a difference image generation unit 77.

[0139] The AI ​​model reading unit 71 of the item execution unit 14D reads the AI ​​model 15D that performs autoencoding. As described above, the AI ​​model 15D encodes an input image into a low-dimensional latent representation, and outputs a reconstructed image obtained by decoding the latent representation as an inference result.

[0140] The difference image generating unit 77 generates a difference image between the restored image, which is the inference result obtained by the inference unit 73 , and the target image converted by the image converting unit 72 .

[0141] The normalization unit 75 performs normalization processing on the difference image based on a second parameter included in the parameter set 16. The second parameter represents the difference between an optimal threshold Th2 optimized for detecting an abnormal portion from the difference image and a predetermined value (typically 0.5). As a result, the predetermined value (typically 0.5) can be used as the optimal threshold for detecting an abnormal portion in the normalized difference image.

[0142] The image conversion unit 76 converts the value of each pixel of the difference image normalized by the normalization unit 75 from floating-point type (e.g., float type or double type) to integer type (e.g., uchar type or byte type). In addition to converting to integer type, the image conversion unit 76 also performs a magnification conversion from the value range (0 to 1) changed by the above-mentioned standardization process to a range that can be represented as an image (e.g., 0 to 255).

[0143] The item execution unit 14D outputs to the outside the image (hereinafter referred to as "processed difference image") converted by the image conversion unit 76. The processed difference image is an example of "output information" in the present disclosure.

[0144] In the fourth example of the internal configuration of the item execution unit 14D, the normalization unit 75 may be omitted. In this case, the item execution unit 14D may perform integer type conversion and magnification conversion on the restored image generated by the inference unit 73, and then generate a difference image between the integer-converted restored image and the target image.

[0145] (Fifth Example) Fig. 17 is a diagram showing a fifth example of the internal configuration of an item execution unit. Fig. 17 shows the internal configuration of an item execution unit 14E corresponding to the general-purpose processing item "labeling." As shown in Fig. 17, the item execution unit 14E includes a binarization processing unit 81, a blobbing unit 82, and a determination unit 83.

[0146] The binarization processing unit 81 binarizes the input image. Specifically, the binarization processing unit 81 converts the value of each pixel of the input image into 0 or 1 based on a preset threshold value.

[0147] The blobbing unit 82 assigns a label to a pixel cluster consisting of consecutive pixels with a value of "1" (or "0").

[0148] The determination unit 83 determines whether the pixel block labeled by the blobbing unit 82 satisfies a determination condition. The determination condition is set in advance by the flow creation unit 11. For example, the determination condition is that the area of ​​the pixel block is equal to or less than a threshold value. The item execution unit 14E outputs the determination result by the determination unit 83 to the outside.

[0149] (Application examples of the item execution units of the first to fifth examples) A ​​user of the image processing device 100 creates the image processing flow shown in Fig. 10 using the screen 30 shown in Fig. 9. In this case, the image processing unit 12 includes an item execution unit 14C corresponding to the AI ​​processing item 40a included in the image processing flow shown in Fig. 10, and an item execution unit 14E corresponding to the general-purpose processing item 42b. The item execution unit 14E performs labeling on the processed output image output from the item execution unit 14C.

[0150] The normalization unit 75 of the item execution unit 14C performs normalization processing on the image, which is the inference result obtained by the inference unit 73, according to the above formula (1). The difference between the optimal threshold Th1 determined by the evaluation tool 304 and a predetermined value (typically 0.5) is substituted for "threshold" in formula (1). Therefore, the predetermined value (typically 0.5) is the optimal threshold for determining whether the normalized image is classified into a specific class. Therefore, the user simply sets the value (typically 127 or 128) obtained by converting the predetermined value (typically 0.5) into an integer as the threshold for the binarization processing unit 81 of the item execution unit 14E. This allows the user to easily set an optimal threshold for the binarization processing unit 81. In other words, the user does not need to adjust the threshold for the binarization processing unit 81.

[0151] A user of the image processing device 100 may create an image processing flow including the AI ​​processing item 40d and the general-purpose processing item 42b of "labeling" using the screen 30 shown in Fig. 9. In this case, the image processing unit 12 includes an item execution unit 14D corresponding to the AI ​​processing item 40d and an item execution unit 14E corresponding to the general-purpose processing item 42b. The item execution unit 14E performs labeling on the processed difference image output from the item execution unit 14D.

[0152] The normalization unit 75 of the item execution unit 14D performs normalization processing on the difference image generated by the difference image generation unit 77 according to the above equation (1). The difference between the optimal threshold Th2 determined by the evaluation tool 304 and a predetermined value (typically 0.5) is substituted for "threshold" in equation (1). Therefore, in the normalized difference image, the predetermined value (typically 0.5) is the optimal threshold for detecting abnormalities in the difference image. Therefore, the user simply sets the value (typically 127 or 128) obtained by converting the predetermined value (typically 0.5) into an integer as the threshold for the binarization processing unit 81 of the item execution unit 14E. This allows the user to easily set an optimal threshold for the binarization processing unit 81. In other words, the user does not need to adjust the threshold for the binarization processing unit 81.

[0153] A user of the image processing device 100 may use the screen 30 shown in FIG. 9 to create an image processing flow including the AI ​​processing item 40c and the general-purpose processing item 42 of "threshold judgment." In this case, the image processing unit 12 includes an item execution unit 14A corresponding to the AI ​​processing item 40c and an item execution unit 14 corresponding to the general-purpose processing item 42 of "threshold judgment." The item execution unit 14 corresponding to the general-purpose processing item 42 of "threshold judgment" determines whether the inference result output from the item execution unit 14A satisfies a predetermined judgment condition and outputs the judgment result. In this case, for example, a condition that the membership probability of a specific class is within a set range is set as the judgment condition.

[0154] A user of the image processing device 100 may use the screen 30 shown in FIG. 9 to create an image processing flow including the AI ​​processing item 40e and the general-purpose processing item 42 for "threshold judgment." In this case, the image processing unit 12 includes an item execution unit 14A corresponding to the AI ​​processing item 40e and an item execution unit 14 corresponding to the general-purpose processing item 42 for "threshold judgment." The item execution unit 14 corresponding to the general-purpose processing item 42 for "threshold judgment" determines whether the inference result output from the item execution unit 14A satisfies a predetermined judgment condition and outputs the judgment result. In this case, the judgment condition may be, for example, a condition related to the size or position of a rectangular area in which the probability of belonging to a specific class exceeds a threshold.

[0155] <Operation Procedure of Image Processing Device> Fig. 18 is a diagram showing an example of the operation procedure of the image processing device. First, the inspector prepares the AI ​​model 15 (step S11). Specifically, the inspector stores a file of the AI ​​model 15 in the image processing device 100. Furthermore, the inspector prepares the parameter set 16 (step S12).

[0156] Next, the inspector sets the file path of the AI ​​model 15 in the image processing device 100 (step S13).

[0157] Next, the inspector instructs the image processing device 100 to read the AI ​​model 15 (step S14). As a result, the image processing device 100 reads the file of the AI ​​model 15 based on the set file path.

[0158] The image processing device 100 notifies that the reading of the file of the AI ​​model 15 has been completed (step S15).

[0159] Next, the inspector sets the parameter set 16 in the image processing device 100 (step S16).

[0160] Next, the inspector instructs the image processing device 100 to execute image processing (step S17).

[0161] The CPU 110 of the image processing device 100 executes the designated image processing (step S18). Specifically, the CPU 110 executes image processing according to a pre-created image processing flow.

[0162] Next, the CPU 110 outputs the results of the image processing (step S19), allowing the inspector to check the output results and determine whether the appearance of the object is good or bad.

[0163] 18, in step S18, the CPU 110 executes image processing including at least one AI processing item. Therefore, step S18 includes steps S18a to S18c corresponding to each of the at least one AI processing item.

[0164] In step S18a, the CPU 110 acquires a target image. For example, the CPU 110 acquires the target image from the camera 200. Alternatively, the CPU 110 acquires, as the target image, an image obtained by executing a processing item preceding the AI ​​processing item.

[0165] In step S18b, the CPU 110 converts the target image based on the first parameter included in the parameter set 16 so that it satisfies the input image format of the AI ​​model 15.

[0166] In step S18c, the CPU 110 executes an inference process using the AI ​​model 15. Specifically, the CPU 110 inputs the converted target image into the AI ​​model 15, thereby obtaining an inference result of the AI ​​model 15.

[0167] <Advantages> As described above, the image processing device 100 according to this embodiment includes a flow creation unit 11 that creates an image processing flow by selecting one or more processing items to be executed from among a number of processing items and setting the execution order of the one or more processing items, and an image processing unit 12 that executes the image processing flow. The multiple processing items include at least one AI processing item 40. Each of the at least one AI processing item 40 defines acquisition of an externally created AI model 15 and output of at least one of an inference result obtained by inputting a target image into the AI ​​model 15 and output information generated from the inference result.

[0168] This allows a user of the image processing device 100 to have the image processing device 100 execute an image processing flow that executes multiple processing items, including AI processing items that use an AI model that the user has created, in the desired order.

[0169] Each of at least one AI processing item 40 defines obtaining a first parameter indicating an input image format of the AI ​​model 15, converting a target image using the first parameter to satisfy the input image format, and inputting the converted target image into the AI ​​model 15.

[0170] At least one AI processing item 40 includes, for example, an AI processing item 40a that uses an AI model 15 that outputs an output image in which each pixel has a floating-point value as an inference result. The AI ​​processing item 40a is an example of a “first AI processing item” in the present disclosure.

[0171] The AI ​​processing item 40a defines the execution of one or more image processing operations on an output image and the output image after the one or more image processing operations are performed as output information. The one or more image processing operations include an integerization operation that converts the values ​​of each pixel into integers and a normalization operation that normalizes the values ​​of each pixel. This makes the output information from the item execution unit 14C that executes the AI ​​processing item 40a easier to use in subsequent processing.

[0172] For example, the value of each pixel in the output image represents the probability of belonging to a specific class. The AI ​​processing item 40a defines obtaining a second parameter representing the difference between an optimal threshold Th1 optimized for determining classification into a specific class and a predetermined value, and performing a normalization process using the second parameter. As a result, the predetermined value can be used as the optimal threshold for determining classification into a specific class in the normalized output image.

[0173] At least one of the AI ​​processing items 40 includes, for example, an AI processing item 40d that uses an AI model 15D that outputs a restored image as an inference result by performing an autoencoder on a target image. The AI ​​processing item 40d is an example of a “second AI processing item” in the present disclosure.

[0174] The AI ​​processing item 40d defines the steps of generating a difference image between the converted target image and the restored image, and outputting the processed difference image obtained by performing one or more image processing operations on the difference image as output information. The one or more image processing operations include an integerization operation that converts the values ​​of each pixel into integers, and a normalization operation that normalizes the values ​​of each pixel. This makes the output information from the item execution unit 14D that executes the AI ​​processing item 40d easier to use in subsequent processing.

[0175] The AI ​​processing item 40d defines obtaining a second parameter representing the difference between an optimal threshold Th2 optimized for detecting an abnormal portion from the difference image and a predetermined value, and performing a normalization process using the second parameter, so that the predetermined value can be used as the optimal threshold for detecting an abnormal portion in the normalized difference image.

[0176] The flow creation unit 11 accepts the designation of the AI ​​model 15 in response to the AI ​​processing item 40 being set as one or more processing items to be executed, and determines whether the designated AI model 15 is compatible with the AI ​​processing item 40. By checking the determination result by the flow creation unit 11, the user of the image processing device 100 can thereby recognize that an AI model that is not suitable for the AI ​​processing item has been mistakenly designated.

[0177] <Modification> In the item execution units 14C and 14D, one of the normalization unit 75 and the image conversion unit 76 may be omitted.

[0178] FIG. 19 is a diagram showing a modified example of the item execution unit shown in FIG. 15. As shown in FIG. 19, the item execution unit 14F differs from the item execution unit 14C shown in FIG. 15 in that it includes an image conversion unit 78. The item execution unit 14F uses an AI model 15 (for example, AI model 15A) that outputs an output image as an inference result. The AI ​​model 15A is an example of the "third AI model" of the present disclosure.

[0179] The image conversion unit 78 performs image processing for display on the image of the inference result obtained from the inference unit 73. For example, the image conversion unit 78 performs image processing to add a specific color to an area where the probability (score) of belonging to a specific class exceeds a preset threshold. The item execution unit 14F displays the image that has been image processed by the image conversion unit 78 (corresponding to the "second processed image" in the present disclosure) on the display 150.

[0180] Fig. 20 is a block diagram showing the functional configuration of an image processing device according to a modified example. As shown in Fig. 19, the image processing device 100 according to the modified example differs from the image processing device 100 shown in Fig. 8 in that it includes an update unit 17.

[0181] The update unit 17 updates the AI ​​model 15 to a re-learned AI model. The re-learned AI model is obtained by re-learning the AI ​​model using the target image.

[0182] Specifically, the update unit 17 accumulates target images input to the AI ​​model 15. The update unit 17 may upload the target images to a server on a cloud system, or may upload the target images to an on-premise hard disk drive. The update unit 17 may upload the target image to a storage destination every time the target image is input to the AI ​​model 15, or may upload multiple target images input to the AI ​​model 15 over a certain period of time to a storage destination all at once.

[0183] The AI ​​model creator performs annotation on the accumulated target images. As a result, correct answer data is assigned to each target image. The AI ​​model creator sets the target images as re-learning images and creates a re-learning dataset including the re-learning images and correct answer data. The AI ​​model creator inputs the re-learning dataset into the learning device 300. As a result, the learning device 300 uses the re-learning dataset to re-learn the AI ​​model 15.

[0184] The update unit 17 periodically acquires the re-trained AI model or acquires the parameter set corresponding to the re-trained AI model in response to a user instruction. At this time, the update unit 17 also acquires the parameter set corresponding to the re-trained AI model. The update unit 17 updates the AI ​​model 15 and the parameter set 16 to the re-trained AI model and the latest parameter set. As a result, the item execution unit 14, which executes the AI ​​processing item, performs inference processing using the re-trained AI model. Furthermore, the item execution unit 14 performs conversion processing and normalization processing of the target image based on the latest parameter set.

[0185] §3 Supplementary Note As described above, the present embodiment includes the following disclosure.

[0186] (Configuration 1) An image processing device (100) comprising: a flow creation unit (11, 110) that creates an image processing flow by selecting one or more processing items to be executed from a plurality of processing items and setting the execution order of the one or more processing items; and an image processing unit (12, 110) that executes the image processing flow, wherein the plurality of processing items include at least one AI processing item (40), and each of the at least one AI processing item (40) defines: acquiring an AI model (15) created externally; and outputting at least one of an inference result obtained by inputting a target image into the AI ​​model (15) and output information generated from the inference result.

[0187] (Configuration 2) The image processing device (100) according to Configuration 1, wherein each of the at least one AI processing item (40) defines: obtaining a first parameter indicating an input image format of the AI ​​model (15); converting the target image using the first parameter to satisfy the input image format; and inputting the converted target image into the AI ​​model (15).

[0188] (Configuration 3) The at least one AI processing item (40) includes a first AI processing item (40a) that uses a first AI model (15A) that outputs an output image in which each pixel has a floating-point value as the inference result, and the first AI processing item (40a) defines: performing one or more image processes on the output image; and outputting the output image after the one or more image processes have been performed as the output information, and the one or more image processes include at least one of a first process that converts the value of each pixel to an integer and a second process that normalizes the value of each pixel. This is the image processing device (100) described in Configuration 1.

[0189] (Configuration 4) The image processing device (100) according to Configuration 3, wherein the value of each pixel of the output image represents the probability of belonging to a specific class, and the first AI processing item (40a) defines: obtaining a second parameter representing the difference between a threshold optimized for determining classification into the specific class and a predetermined value; and performing the second processing using the second parameter.

[0190] (Configuration 5) The at least one AI processing item (40) includes a second AI processing item (40d) that uses a second AI model (15D) that performs an autoencoder on the target image to output a restored image as the inference result, and the second AI processing item defines: generating a difference image between the target image and the restored image; and outputting the difference image or a processed image obtained by performing one or more image processing operations on the difference image as the output information, the image processing device (100) described in Configuration 1.

[0191] (Configuration 6) The image processing device (100) according to Configuration 5, wherein the one or more image processes include at least one of a first process for converting a value of each pixel to an integer and a second process for normalizing a value of each pixel.

[0192] (Configuration 7) The image processing device (100) according to Configuration 6, wherein the second AI processing item (40d) defines: acquiring a second parameter representing the difference between a threshold optimized for detecting an abnormal portion from the difference image and a predetermined value; and executing the second processing using the second parameter.

[0193] (Configuration 8) The image processing device (100) according to any one of configurations 1 to 7, wherein the flow creation unit (11, 110) accepts designation of an AI model (15) to be used in the target AI processing item in response to a target AI processing item being selected as the one or more processing items among the at least one AI processing item, and determines whether the designated AI model is compatible with the target AI processing item.

[0194] (Configuration 9) An image processing device (100) described in any of configurations 1 to 8, wherein the at least one AI processing item (40) includes a third AI processing item using a third AI model that outputs an output image as the inference result, and the third AI processing item defines: generating a first processed image obtained by performing a first image processing on the output image and a second processed image obtained by performing a second image processing on the output image; displaying the second processed image on a display (150); and outputting the first processed image to a subsequent processing item among the one or more processing items.

[0195] (Configuration 10) The image processing device (100) according to any one of configurations 1 to 9, further comprising an update unit (17, 110) that updates the AI ​​model (15) to a re-learned AI model, and the re-learned AI model is obtained by re-learning the AI ​​model (15) using the target image.

[0196] (Configuration 11) An image processing method comprising: a processor (110) creating an image processing flow by selecting one or more processing items to be executed from a plurality of processing items and setting an execution order of the one or more processing items; and the processor (110) executing the image processing flow, wherein the plurality of processing items include at least one AI processing item, and each of the at least one AI processing item defines: acquiring an AI model (15) created externally; and outputting at least one of an inference result obtained by inputting a target image into the AI ​​model (15) and output information generated from the inference result.

[0197] (Configuration 12) A program (115) for causing a computer (110) to execute an image processing method, the image processing method including: creating an image processing flow by selecting one or more processing items to be executed from a plurality of processing items and setting an execution order of the one or more processing items; and executing the image processing flow, the plurality of processing items including at least one AI processing item, each of which defines: acquiring an AI model (15) created externally; and outputting at least one of an inference result obtained by inputting a target image into the AI ​​model (15) and output information generated from the inference result.

[0198] Although the embodiments of the present invention have been described, the embodiments disclosed herein should be considered to be illustrative and not restrictive in all respects. The scope of the present invention is defined by the claims, and it is intended to include all modifications within the meaning and scope of the claims.

[0199] 11 Flow creation unit, 12 Image processing unit, 13 Memory unit, 14, 14A to 14F Item execution unit, 15, 15A to 15F AI model, 16 Parameter set, 17 Update unit, 20, 30, 60, 60A Screen, 32 Set item display area, 34 Processing item selection area, 36 Camera image display area, 38 Insert / add button, 39 Order change button, 40, 40a to 40e AI processing item, 42, 42a, 42b Processing item, 50 Image processing flow, 61, 62 Setting area, 61a, 62b, 62c, 66, 67 Input field, 62a Radio button, 63, 64 Display area, 65 Button, 71 AI model reading unit, 72, 76, 78 Image conversion unit, 73 Inference unit, 74, 83 Determination unit, 75 Normalization unit, 77 Difference image generation unit, 81 Binarization processing unit, 82 Blobbing unit, 90A, 90B, 90F Training dataset, 91 Object, 92A, 92B, 92F Learning image, 93 Defect area, 94A, 94B, 94F Correct answer data, 96A, 96B, 96F Input image, 98A Output image, 98B Character string, 98F Inference result, 100 Image processing device, 106 Memory card, 110 CPU, 112 Main memory, 114 Hard disk, 115 Program, 116 Camera interface, 116a Image buffer, 118 Input interface, 120 Display controller, 124 Communication interface, 126 Data reader / writer, 128 Bus, 150 Display, 160 Input device, 200 Camera, 300 Learning device, 302 Learning tool, 304 Evaluation tool, M user.

Claims

1. An image processing device comprising: a flow creation unit that creates an image processing flow by selecting one or more processing items to be executed from a plurality of processing items and setting an execution order of the one or more processing items; and an image processing unit that executes the image processing flow, wherein the plurality of processing items include at least one AI processing item, and each of the at least one AI processing item defines: acquiring an AI model created externally; and outputting at least one of an inference result obtained by inputting a target image into the AI ​​model and output information generated from the inference result.

2. The image processing device of claim 1, wherein each of the at least one AI processing item defines: obtaining a first parameter indicating an input image format of the AI ​​model; converting the target image using the first parameter to satisfy the input image format; and inputting the converted target image into the AI ​​model.

3. The image processing device of claim 1, wherein the at least one AI processing item includes a first AI processing item using a first AI model that outputs an output image in which each pixel has a floating-point value as the inference result, the first AI processing item defines: performing one or more image processing operations on the output image; and outputting the output image after the one or more image processing operations have been performed as the output information, and the one or more image processing operations include at least one of a first processing operation that converts the value of each pixel to an integer and a second processing operation that normalizes the value of each pixel.

4. An image processing device as described in claim 3, wherein the value of each pixel of the output image represents the probability of belonging to a specific class, and the first AI processing item defines: obtaining a second parameter representing the difference between a threshold optimized for determining classification into the specific class and a predetermined value; and performing the second processing using the second parameter.

5. The image processing device of claim 1, wherein the at least one AI processing item includes a second AI processing item using a second AI model that performs an autoencoder on the target image to output a restored image as the inference result, and the second AI processing item defines: generating a difference image between the target image and the restored image; and outputting the difference image or a processed image obtained by performing one or more image processing operations on the difference image as the output information.

6. The image processing device according to claim 5, wherein the one or more image processes include at least one of a first process for converting the value of each pixel into an integer, and a second process for normalizing the value of each pixel.

7. The image processing device described in claim 6, wherein the second AI processing item defines: obtaining a second parameter representing the difference between a threshold optimized for detecting abnormal portions from the difference image and a predetermined value; and executing the second processing using the second parameter.

8. An image processing device as claimed in any one of claims 1 to 7, wherein the flow creation unit accepts designation of an AI model to be used in a target AI processing item among the at least one AI processing item in response to the target AI processing item being selected as the one or more processing items, and determines whether the designated AI model is compatible with the target AI processing item.

9. An image processing device as described in any one of claims 1 to 8, wherein the at least one AI processing item includes a third AI processing item using a third AI model that outputs an output image as the inference result, and the third AI processing item defines: generating a first processed image obtained by performing a first image processing on the output image and a second processed image obtained by performing a second image processing on the output image; displaying the second processed image on a display; and outputting the first processed image to a subsequent processing item of the one or more processing items.

10. An image processing device according to any one of claims 1 to 9, further comprising an update unit that updates the AI ​​model to a re-trained AI model, wherein the re-trained AI model is obtained by re-training the AI ​​model using the target image.

11. An image processing method comprising: a processor creating an image processing flow by selecting one or more processing items to be executed from a plurality of processing items and setting an execution order of the one or more processing items; and the processor executing the image processing flow, wherein the plurality of processing items include at least one AI processing item, each of the at least one AI processing item defining: acquiring an AI model created externally; and outputting at least one of an inference result obtained by inputting a target image into the AI ​​model and output information generated from the inference result.

12. A program for causing a computer to execute an image processing method, the image processing method including: creating an image processing flow by selecting one or more processing items to be executed from a plurality of processing items and setting an execution order of the one or more processing items; and executing the image processing flow, the plurality of processing items including at least one AI processing item, each of the at least one AI processing item defining: acquiring an AI model created externally; and outputting at least one of an inference result obtained by inputting a target image into the AI ​​model and output information generated from the inference result.

Citation Information

Patent Citations

  • Image processing system, information processing apparatus, information processing method, and information processing program

    JP2018124607A

  • Ai platform system

    JP2022191637A

Cited By

  • New media image processing method and system based on AI large model

    CN120563656A