Image processing apparatus and image processing method
The image processing device improves the detection rate of small abnormalities by generating and using partial images from the original image for learning and inference, effectively addressing the low detection rate issue in conventional systems.
Patent Information
- Application Number
- JP2025020983
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-03-04
AI Technical Summary
Conventional image processing devices using machine learning models have a low detection rate for small abnormalities in images.
An image processing device that generates learning data using multiple partial images from an original image and performs inference using these partial images, improving the detection rate of small abnormalities.
The proposed solution significantly enhances the detection rate of small abnormalities in images by dispersing the positions of abnormal parts and increasing their proportion in the image, leading to improved classification accuracy.
Smart Images

Figure 2025081415000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an image processing device and an image processing method. [Background technology]
[0002] Image recognition (classification) using machine learning models has become more widely used as its performance improves. It is also known that in order to improve classification accuracy, a final classification is determined based on the results of separate classifications using different machine learning models (Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2020-112926 A Summary of the Invention [Problem to be solved by the invention]
[0004] On the other hand, conventional image processing devices that use machine learning models to detect abnormalities contained in images of an object to be inspected have a problem in that they have a low detection rate for small abnormalities.
[0005] Therefore, one of the objectives of the present invention is to improve the detection rate of small abnormalities in an image processing device and image processing method that detects abnormalities contained in an image of an inspection target using a machine learning model. [Means for solving the problem]
[0006] The above-mentioned object can be achieved by an image processing device comprising a generation means for generating learning data for a learning model from a first original image of the inspection object, a learning means for learning the learning model using the learning data, and an inference means for performing inference processing on input data generated from a second original image of the inspection object using the learned learning model, wherein the generation means generates a plurality of partial images from the first original image and generates the learning data using the plurality of partial images, and the inference means uses the partial images of the second original image as input data. Effect of the Invention
[0007] According to the present invention, in an image processing device and an image processing method that detect abnormal portions contained in an image of an inspection target using a machine learning model, it is possible to improve the detection rate of small-sized abnormal portions. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 shows a configuration of an image processing system 100. [Diagram 2] A block diagram showing the configuration of a cloud server 200 and an edge server 300. [Diagram 3] FIG. 5A is a diagram showing an example of the appearance of a smartphone 500, and FIG. 5B is a diagram showing an example of the configuration. [Figure 4] 6A and 6B are diagrams showing an example of the appearance of a printer 600, and FIG. 6C is a diagram showing an example of the configuration. [Diagram 5] FIG. 1 shows a software configuration of a processing system 100. [Figure 6] A conceptual diagram showing the input / output structure when using the learning model 252 and the trained model 352. [Figure 7] FIG. 13 is a diagram showing the contents of pre-learning processing in the first embodiment. [Figure 8] FIG. 1 shows a configuration during learning and inference in the first embodiment. [Figure 9] FIG. 11 is a diagram showing the contents of test image processing in the first embodiment. [Figure 10] FIG. 11 is a diagram showing the contents of pre-learning processing in the second embodiment. [Figure 11] Diagram showing how to split an image [Figure 12] FIG. 11 is a diagram showing the contents of test image processing in the second embodiment. [Figure 13] FIG. 13 shows a configuration during learning and inference in the third embodiment. [Figure 14] FIG. 13 shows a configuration during learning and inference in the third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] The present invention will be described in detail below based on its exemplary embodiments with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. In addition, although multiple features are described in the embodiments, not all of them are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numbers are used for the same or similar configurations, and duplicated explanations are omitted.
[0010] In addition, the following describes an embodiment of the invention in an image processing system in which a machine learning model is arranged in an external device of an image processing device, the image processing device provides images to the external device, and the external device learns the machine learning model and performs recognition (classification) using the learned model. However, the image processing device may have the machine learning model and implement the invention.
[0011] <First embodiment> (Image Processing System Configuration) FIG. 1 is a block diagram showing an example of the overall configuration of an image processing system 100 according to an embodiment of the present invention. The image processing system 100 has a configuration in which a cloud server 200, an edge server 300, and a device 400 are connected to each other so as to be able to communicate with each other. Here, a configuration example in which the cloud server 200 is on the Internet, and the edge server 300 and the device 400 are on a local area network (LAN) 102 will be described. However, the cloud server 200, the edge server 300, and the device 400 can be connected in any configuration. In addition, it is not essential to separate the cloud server 200 and the edge server 300, and the functions of both may be implemented by one server. Furthermore, the functions implemented by the cloud server 200 and the edge server 300 may be implemented by the device 400.
[0012] The device 400 is a general term for electronic devices capable of communicating with the cloud server 200 and the edge server 300. In FIG. 1, a digital camera 402, a client terminal 401, a smartphone 500, and a printer 600 are shown as examples of the device 400. However, the device 400 may include any electronic device that has a communication interface and can supply image data. The client terminal 401 is a computer device such as a personal computer, a tablet terminal, or a game console. In the following, the configuration and operation common to the digital camera 402, the client terminal 401, the smartphone 500, and the printer 600 will be described as the configuration and operation of the device 400.
[0013] The router 103 connecting the Internet 104 and the LAN 102 may have a wireless LAN access point function. In this case, the device 400 can be connected to the LAN 102 through an access point provided by the router 103. For example, it is also possible to configure the printer 600 and the client terminal 401 to be connected to the LAN 102 by wire, and the smartphone 500 and the digital camera 402 to be connected to the LAN 102 by wireless. The device 400 and the edge server 300 can communicate with the cloud server 200 via the Internet 104 connected via the router 103.
[0014] The edge server 300 and the device 400 can communicate with each other via the LAN 102. The devices 400 can also communicate with each other via the LAN 102. The smartphone 500 and the printer 600 can communicate with each other via short-range wireless communication 101 in addition to via the LAN 102. The short-range wireless communication 101 may be wireless communication conforming to, for example, the Bluetooth (registered trademark) standard or the NFC standard. The smartphone 500 is also connected to a mobile phone network 105, and can also communicate with the cloud server 200 via the mobile phone network 105.
[0015] 1 is an example, and may have a different configuration. For example, a device other than the router 103 may function as an access point. Also, the edge server 300 and the device 400 may be connected in a form different from the LAN 102. For example, various forms are possible, such as wireless access technology classified as LPWA (Low Power, Wide Area), wireless connection using ZigBee, Bluetooth, infrared communication, short-range wireless communication 101, and wired connection using USB.
[0016] (Server configuration) 2 is a block diagram showing an example of the configuration of the cloud server 200 and the edge server 300. The configuration of the cloud server 200 will be described below, but it is assumed that the edge server 300 has the same functions.
[0017] The cloud server 200 includes a main board 210 that controls the entire device, a network connection unit 201 , and a hard disk unit 202 .
[0018] A CPU 211 arranged on the main board 210 operates according to a control program stored in a program memory 213 (ROM) connected via an internal bus 212, and settings and variables stored in a data memory 214 (RAM). The CPU 211 controls the operation of the server 200 by executing the program.
[0019] The CPU 211 controls the network connection unit 201 via a network control circuit 215 to communicate with other devices via a network such as the Internet 104 or a LAN 102. The CPU 211 can also write data to a hard disk unit 202 connected via a hard disk control circuit 216 and read data from the hard disk unit 202.
[0020] The hard disk unit 202 stores an operating system that is loaded into the program memory 213 and executed by the CPU 211, control software for the server 200, application software, various data, and the like.
[0021] A GPU 217 is connected to the internal bus 212 of the main board 210. The GPU 217 can execute various calculation processes instead of the CPU 211. Since the GPU 217 can execute parallel processing at high speed, it can execute calculations for learning a neural network using a method such as deep learning and calculations for inference using a learned model more efficiently than the CPU 211. In this embodiment, the processing of a learning unit 251 described later is executed using the GPU 217 in addition to the CPU 211. Specifically, the CPU 211 and the GPU 217 cooperate to execute calculations, thereby implementing and learning a machine learning model. The learning unit 251 may be executed using only one of the CPU 211 and the GPU 217. Similarly to the learning unit 251, the inference unit 351 can also be executed using the GPU 217.
[0022] In this embodiment, the edge server 300 has the same configuration as the cloud server 200, but may have a different configuration. For example, the main board 210 of the edge server 300 does not need to be equipped with the GPU 217. In addition, the cloud server 200 and the edge server 300 may have different performances even if they have the same configuration name.
[0023] (Appearance of Smartphone 500) 3(a) is a diagram showing an example of the appearance of the display surface of the smartphone 500 seen from the front, and an example of a communication unit of the smartphone 500. The smartphone 500 is a general term for a mobile phone that has a touch display, a camera, a function for connecting to a data network such as the Internet, and is capable of executing various applications on an OS.
[0024] The short-distance wireless communication unit 501 can communicate with a short-distance wireless communication unit of another device within the communication range. The wireless LAN unit 502 can communicate with a wireless LAN access point or a wireless LAN unit of another device within the communication range. The line connection unit 503 can connect to a mobile phone network and communicate. These communication units are stored in the housing of the smartphone 500 and communicate through an antenna provided on the surface of the housing or the like.
[0025] The touch display 504 is an LCD or organic EL display panel equipped with a touch panel. The surface on which the display screen of the touch display 504 exists is defined as the front surface of the smartphone 500. When a touch operation on the touch display 504 is detected, the smartphone 500 interprets the touch operation according to the display content of the touch display 504 and executes various operations. The power button 505 is a button for turning the power of the smartphone 500 on and off.
[0026] (Smartphone configuration) 3B is a block diagram showing an example of the configuration of the smartphone 500. The smartphone 500 includes a main board 510 that controls the entire device, a wireless LAN unit 502, a short-range wireless communication unit 501, and a line connection unit 503.
[0027] A CPU 511 arranged on the main board 510 operates according to a control program stored in a program memory 513 (ROM) connected via an internal bus 512 and settings and variables stored in a data memory 514 (RAM). The CPU 511 controls the operation of the smartphone 500 by executing the program.
[0028] The CPU 511 can communicate with a wireless LAN access point or a wireless LAN unit of another device within the communication range by controlling the wireless LAN unit 502 via the wireless LAN control circuit 515. The CPU 511 can detect other short-distance wireless communication terminals within the communication range and transmit and receive data with other short-distance wireless communication terminals by controlling the short-distance wireless communication unit 501 via the short-distance wireless communication control circuit 516. The CPU 511 can also connect to the mobile phone network 105 and transmit and receive voice and data by controlling the line connection unit 503 via the line control circuit 517. The CPU 511 can control the display content of the touch display 504 and detect touch operations by controlling the operation unit control circuit 518.
[0029] CPU 511 can capture still images and videos by controlling camera unit 519. CPU 511 stores image data obtained by capturing images in image memory 520 in data memory 514. CPU 511 can also store image data acquired from the outside via a mobile phone line, LAN 102, and short-range wireless communication 101 in image memory 520, and transmit image data stored in image memory 520 to the outside.
[0030] The non-volatile memory 521 retains data even when the power is turned off, and therefore stores user data such as contacts, communication history, image data to be saved, and application software.
[0031] (Printer appearance) FIG. 4 is a diagram showing an example of the appearance of a printer 600. In this embodiment, the printer 600 is a printer equipped with a scanner, called a complex machine or multifunction printer (MFP). FIG. 4(a) is a perspective view showing an example of the appearance of the printer 600. A manuscript table 601 is made of a transparent material such as glass, and is a place where a manuscript to be read is placed. A manuscript table pressure plate 602 can be opened and closed, and when closed, it presses the manuscript table 601 and blocks light from the manuscript table 601. Recording media of various sizes set in a print paper insertion port 603 are transported one by one during printing, pass through the printing unit, and are discharged from a print paper discharge port 604.
[0032] 4B shows an example of the appearance of the top surface of the printer 600 and a schematic diagram of a communication unit provided in the printer 600. An operation panel 605 and a short-range wireless communication unit 606 are provided on the top surface of the platen 602. The short-range wireless communication unit 606 can communicate with a short-range wireless communication unit of another device within the communication range. In addition, a wireless LAN antenna 607 is connected to a wireless LAN unit (not shown) and enables the printer 600 to communicate with a wireless LAN access point or a wireless LAN unit of another device within the communication range.
[0033] (Printer configuration) 4C is a block diagram showing an example of the configuration of the printer 600. The printer 600 has a main board 610 that controls the entire device, a wireless LAN unit 608, and a short-range wireless communication unit 606.
[0034] A CPU 611 arranged on the main board 610 operates according to a control program stored in a program memory 613 (ROM) connected via an internal bus 612, and settings and variables stored in a data memory 614 (RAM). The CPU 611 controls the operation of the printer 600 by executing the program.
[0035] The CPU 611 controls a scanner unit 615 to read an original document and stores the read image data in an image memory 616 in a data memory 614. The CPU 611 also controls a printing unit 617 to print an image in the image memory 616 in the data memory 614 onto a recording medium.
[0036] The CPU 611 can communicate with a wireless LAN access point or a wireless LAN unit of another device within the communication range by controlling the wireless LAN unit 608 via the wireless LAN control circuit 618. The CPU 611 can detect other short-range wireless communication terminals within the communication range and transmit and receive data to and from other short-range wireless communication terminals by controlling the short-range wireless communication unit 606 via the short-range wireless communication control circuit 620.
[0037] The CPU 611 controls the operation unit control circuit 621 to display the status of the printer 600, a menu screen, and the like on the operation panel 605, and to detect operations on the operation panel 605. The operation panel 605 is equipped with a backlight, and the CPU 611 can control the turning on and off of the backlight via the operation unit control circuit 621.
[0038] (Software configuration) Fig. 5 is a diagram showing an example of the software configuration of the image processing system 100. For ease of explanation and understanding, Fig. 5 shows only software related to learning and inference processing, which is necessary for explaining this embodiment, among the software running on the image processing system 100. For example, an operating system, middleware, applications for maintenance, and the like are not shown.
[0039] Cloud server 200 has a learning data generation unit 250, a learning unit 251, and a learning model 252. The learning data generation unit 250 is a module that generates learning data for learning model 252 from data received from the outside. The learning data has input data X and teacher data T indicating a correct answer of a learning result for the input data X. The learning data generation unit 250 supplies the generated learning data to the learning unit 251.
[0040] The learning unit 251 is a program module that executes learning of the learning model 252 using the learning data generated by the learning data generation unit 250. The learning model 252 accumulates the results of learning performed by the learning unit 251. Here, the learning model 252 is implemented using a neural network. Learning of the learning model 252 is executed by optimizing weighting parameters between each node of the neural network using the learning data by a known method.
[0041] The learning model 252 (trained model) for which parameter optimization (learning) has been completed is supplied from the cloud server 200 to the edge server 300, and is held in the edge server 300 as a trained model 352. The entire learning model 252 may be supplied to the edge server 300, or only a portion necessary for inference processing in the edge server 300 may be supplied to the edge server 300. The edge server 300 can use the trained model 352 to perform inference processing such as classification of input data and prediction of numerical values based on the input data (regression).
[0042] The edge server 300 has a data collection and provision unit 350, an inference unit 351, and a trained model 352. The data collection and provision unit 350 is a module that transmits data received from the device 400 and data collected by the edge server 300 itself to the cloud server 200 as data for generating training data. The inference unit 351 is a program module that provides input data based on the data received from the device 400 to the trained model 352 to execute inference processing, and returns the output of the trained model 352 to the device 400.
[0043] The device 400 includes an application unit 450 and a data transmission / reception unit 451. The application unit 450 is a module that realizes various functions in the device 400. Here, it is assumed that an application module included in the application unit 450 uses a trained model 352 that the edge server 300 has.
[0044] The data transmitter / receiver unit 451 transmits data used for learning the learning model 252, among the data acquired by the device 400, to the data collector / provider unit 350 of the edge server 300. The data transmitter / receiver unit 451 transmits data used for inference processing, among the data acquired or generated by the device 400, to the inference unit 351 of the edge server 300. When the data transmitter / receiver unit 451 receives a result of the inference processing from the inference unit 351 of the edge server 300, it supplies the result to the application module that requested the inference processing.
[0045] In the present embodiment, the learning model 252 learned in the cloud server 200 is supplied to the edge server 300 and used for inference processing in the edge server 300. However, the location where the learning model is provided and the location where processing using the learned model is performed may be changed. For example, the learning model may be implemented in the device 400, and learning of the learning model and inference processing using the learning model may also be performed in the device 400. For example, it is possible to determine whether or not to place the learning model in the device 400 based on the relationship between the processing speed and power consumption required for calculations related to the learning model and the hardware resources of the device 400. Then, when it is not possible to place the learning model in the device 400 or placing it is not desirable, the learning model is placed in an external device.
[0046] Furthermore, when placing a learning model in an external device, placing the learning model in an external device on the same network as device 400 can reduce the time required to obtain the results of the inference process compared to placing the learning model in an external device on a different network.
[0047] In this embodiment, learning using a large amount of input data is performed by the cloud server 200, which has a higher processing capacity than the edge server 300, and inference processing is performed by the edge server 300. By performing the inference processing by the edge server 300, the communication time required for the inference processing can be shortened.
[0048] When learning and inference processing are performed by different entities, a configuration suitable for each processing can be adopted, so that resources can be saved and a configuration capable of executing at a higher speed can be used. Note that the location where the learning model is provided and the location where processing using the trained model is performed may be dynamically changed depending on, for example, the state of the network. For example, inference processing is usually performed by the edge server 300, but when the load on the edge server 300 is high, inference processing may be performed by the cloud server 200.
[0049] (Learning model) FIG. 6 is a diagram showing a schematic diagram of the learning process of the learning model 252 and the inference process using the learned model 352. 6(a) is a diagram showing a schematic diagram of input / output data of the learning model 252 in the learning process and a learning method. Input data X 801 is supplied to the input layer of the learning model 252. Details of the input data X 801 will be described later.
[0050] Output data Y 803 is output as a result of processing input data X 801 by the learning model 252, which is a machine learning model. During learning, teacher data T 802 is given as correct answer data for processing the input data X 801. Therefore, by providing the output data Y 803 and the teacher data T 802 to a loss function 804, a deviation amount L 805 of the processing result relative to the correct answer (teacher data) is obtained. For a large amount of learning data, the connection weighting coefficients between the nodes of the neural network constituting the learning model 252 are updated so that the deviation amount L 805 approaches 0. The error backpropagation method is an example of a method for optimizing the connection weighting coefficients between the nodes of the neural network so that the deviation amount L 805 becomes small.
[0051] Specific examples of algorithms for machine learning include nearest neighbor methods, naive Bayes methods, decision trees, and support vector machines. Deep learning and deep metric learning are also known, which use neural networks to generate features and connection weighting coefficients for learning. In this embodiment, among these known algorithms, any of them can be appropriately used in consideration of the application of machine learning. There are no particular limitations on the method of implementing the learning model 252. The learning model 252 can be implemented, for example, by a convolutional neural network (CNN), a recurrent neural network (RNN), an autoencoder, a generative adversarial network (GAN), or the like.
[0052] 6(b) is a diagram showing input and output data of the trained model 352 in the inference process. The input data X 811 is supplied to the input layer of the trained model 352. Details of the input data X 811 will be described later. The input data X 811 has the same format as the input data X 801 used during training, but there is no corresponding teacher data.
[0053] Output data Y 813) is output as a result of processing input data X 811 by the trained model 352. In the inference process, the output data Y 813 is returned to the device 400 as a processing result. The trained model 352 may be implemented by a neural network having the same configuration as the training model 252, or only a part of the training model 252 necessary for the inference process may be included as the trained model 352. By making the trained model 352 have a smaller configuration than the training model 252, it is possible to reduce the amount of data in the trained model 352 and shorten the calculation time during the inference process.
[0054] Fig. 7 shows a specific example of learning data to be applied to the learning model 252 when detecting anomalies from an inspection image of a semiconductor substrate using the image processing system 100. Fig. 7(a) shows an example of original image (first original image) data 900 used for learning data. The original image (original image data) 900 is image data obtained by photographing the surface of a semiconductor substrate, and is photographed so that an abnormal part 905 is located in the center of the image.
[0055] The resolution of the original image data 900 is 224 x 224 pixels, and the horizontal and vertical dimensions of the abnormal part 905 are several pixels to several tens of pixels. In the semiconductor process, even extremely small abnormalities on the substrate can cause defects. Semiconductor substrates may also be warped, or have irregularities on the surface due to abnormal parts. For this reason, a lens with a large focal depth (for example, 5 μm or more) is used to capture the inspection image so that a clear image can be obtained over the entire capture range.
[0056] In this embodiment, the original image data 900 is not used as it is for learning, but a plurality of partial images (partial image data) 901 extracted from the original image data 900 are used for learning. Specifically, data of four partial images obtained by dividing the original image into two equal parts in the horizontal and vertical directions, and one partial image extracted from the center part of the original image and having the same size as these four partial images, are used for learning. The center of the partial image extracted from the center part is equal to the center of the original image. In FIG. 7(b), this image extraction process is described as pre-processing. Moreover, the four partial images obtained by equal division are described as the lower left, upper left, lower right, and upper right, and the one partial image extracted from the center part is described as the center.
[0057] In this way, five partial image data 901 of 112×112 pixels are generated from the original image data 900 of 224×224 pixels. By such pre-processing, the partial images at the lower left, upper left, lower right, and upper right become images in which abnormal parts exist at positions away from the center of the image. By generating partial images, the positions of abnormal parts in the image can be dispersed. In addition, the proportion (area ratio) of abnormal parts in the image can be increased. Furthermore, learning can be performed with more images than when the entire area of the original image is used for learning all at once. Specifically, in the example of FIG. 7, learning can be performed using images five times the number of original images.
[0058] Here, input data is classified into data containing an anomaly and data not containing an anomaly by inference processing using the trained model 352. Therefore, correct classification for each partial image is prepared as training data. Specifically, as shown in FIG. 7(c), an image 902 containing an anomaly to be recognized is classified as class A, and an image 903 containing no anomaly to be recognized (normal) is classified as class B, and training data is generated by a human visually checking and classifying partial images (data) 901.
[0059] In this embodiment, for ease of explanation and understanding, the input data is classified into two classes by the inference process, but it may be classified into three or more classes. The classes may be defined based on a criterion other than the presence or absence of an abnormality to be recognized. The training data may be generated by a method other than visual inspection. For example, a simple learning model may be created based on an image confirmed visually to generate training data for a partial image.
[0060] Fig. 8 is a diagram showing a schematic flow of the learning process and inference process and related data in the image processing system 100 in this embodiment. An image 902 of class A in which an anomaly to be recognized exists and an image 903 of class B in which an anomaly to be recognized exists (normal), as explained with reference to Fig. 7, are supplied from the device 400 to the learning data generating unit 250 via the edge server 300. In addition, when the teacher data is generated based on visual inspection, the teacher data is also supplied from the device 400 to the learning data generating unit 250.
[0061] The learning data generation unit 250 generates input data 801 and teacher data 802 as learning data from an image 902 of class A and an image 903 of class B, and supplies them to the learning unit 251. When teacher data is provided together with the images, the learning data generation unit 250 may use the supplied teacher data. The learning unit 251 uses the learning data to train the learning model 252, and completes the training when the deviation amount L indicated by the loss function becomes less than a predetermined threshold. At this point, the learning model 252 becomes a trained model. Then, all or a part of the learning model 252 is supplied to the edge server 300 and stored as a trained model 352.
[0062] Thereafter, the test image 904 is supplied as input data 811 to the trained model 352, and the determined class is obtained as output data. Note that the test image 904 has the same size (112×112 pixels) as the input data 801 used for training, but the method of generating it from the original image is different.
[0063] 9(a) shows an example of an original image 900 for determining the presence or absence of an abnormality using the trained model 352. The data transmitter / receiver 451 of the device 400 transmits the original image 900 (second original image) to the data collector / provider 350 and the inference unit 351 of the edge server 300. The inference unit 351 generates a test image 904 from the original image 900 and supplies it to the trained model 352 as input data 811.
[0064] 9(b), for each of the original images 900, the inference unit 351 generates a test image 904 of 112×112 pixels by trimming the center of the image so that the center of the test image 904 coincides with the center of the original image 900. As in image 3, due to trimming, the position of an abnormality in the test image 904 may be in the periphery of the image. However, since the trained model 352 has been trained using input data including an image in which an abnormality exists in the periphery, it can make a highly accurate judgment even for a test image such as image 3.
[0065] The machine learning configuration of the edge server 300 and the cloud server 200 that performs the process shown in FIG. 8 can be implemented using, for example, Keras as a deep learning library and TensorFlow as a backend of Keras. However, other backends may be used. Also, it can be implemented using other known machine learning frameworks, whether open source or commercially available, such as TensorFlow, Caffe, Chainer, Pytorch, HALCON, VisionPro Vidi, etc. Also, it can be implemented without using a ready-made framework.
[0066] The trained model according to this embodiment was compared with a trained model trained by a conventional method that uses the original image as is, and the judgment accuracy was compared. As a result, it was confirmed that the trained model according to this embodiment had a higher judgment accuracy.
[0067] ●(Second embodiment) Next, a second embodiment of the present invention will be described. Note that since this embodiment can be implemented by the image processing system 100 described in the first embodiment, a description of the contents common to the first embodiment will be omitted.
[0068] This embodiment differs from the first embodiment in that the abnormal part is present in the center of the photographed image of the subject in that the abnormal part may be present at any position in the photographed image of the subject. Here, an MRI image is used as an example of a photographed image of the subject in which the position of the abnormal part is uncertain. FIG. 10(a) shows an example of an original image 900. In this embodiment, the resolution of the original image 900 is 600×600 pixels, and the horizontal and vertical sizes of the abnormal part are several tens of pixels.
[0069] In the first embodiment, the abnormal part was photographed so as to be located in the center of the original image. Therefore, when generating learning images from the original image 900, the original image 900 was divided equally in the vertical and horizontal directions into four partial images, and one partial image was extracted from the center.
[0070] In contrast, in this embodiment, since abnormal parts exist at various positions in the original image, extraction is not performed from the center part. Therefore, as shown in Fig. 10(b), the original image 900 is divided to generate partial images 901. Fig. 10(b) shows an example in which the original image 900 is divided into two equal parts in the horizontal and vertical directions to generate four partial images 901 of 300 x 300 pixels.
[0071] Note that the original image 900 may be divided in either the horizontal or vertical direction to generate the partial images 901. Fig. 11(a) shows an example in which the original image 900 is divided into three only in the vertical direction to generate the partial images 901, and Fig. 11(b) shows an example in which the original image 900 is divided into three only in the horizontal direction to generate the partial images 901.
[0072] In this embodiment, too, a plurality of partial images are generated from the original image 900 to be used as learning images, thereby increasing the proportion of the detection target (e.g., abnormal part) in the entire image. Also, it is possible to efficiently learn about the detection target whose position in the image is uncertain. As a result, it is possible to improve the recognition accuracy of the detection target.
[0073] Here, input data is classified into data containing an anomaly and data not containing an anomaly by inference processing using the trained model 352. Therefore, correct classification for each partial image 901 is prepared as training data. Specifically, as shown in FIG. 10(c), an image 902 containing an anomaly to be recognized is classified as class A, and an image 903 containing no anomaly to be recognized (normal) is classified as class B, and the training data is generated by having a human visually check and classify the partial images 901.
[0074] In this embodiment, for ease of explanation and understanding, the input data is classified into two classes by the inference process, but it may be classified into three or more classes. The classes may be defined based on a criterion other than the presence or absence of an abnormality to be recognized. The training data may be generated by a method other than visual inspection. For example, a simple learning model may be created based on an image confirmed visually to generate training data for a partial image.
[0075] Using the input data (partial image 901) generated in this manner and the teacher data, the learning model 252 is trained in the same manner as in the first embodiment, and is supplied to the edge server 300 as the trained model 352.
[0076] Thereafter, the test image 904 is supplied to the trained model 352 as input data 811, and the determined class is obtained as output data. In this embodiment, the test image 904 is generated in the same manner as the training image (partial image 901).
[0077] 12(a) shows an example of an original image 900 for determining the presence or absence of an abnormality using the trained model 352. As shown in FIG. 12(b), the inference unit 351 generates partial images by dividing each of the original images 900 into two equal parts in the horizontal and vertical directions as test images 904. The inference unit 351 supplies the generated test images 904 to the trained model 352 as input data 811.
[0078] In this embodiment, unlike the first embodiment, multiple test images are generated from one original image. Therefore, if one or more of the multiple test images generated from the same original image are classified into class A by the trained model 352, it is determined that an anomaly has been detected in the original image.
[0079] The trained model according to this embodiment was compared with a trained model trained by a conventional method that uses the original image as is, and the judgment accuracy was compared. As a result, it was confirmed that the trained model according to this embodiment had a higher judgment accuracy.
[0080] ●(Third embodiment) Next, a third embodiment of the present invention will be described. Since this embodiment can be implemented by the image processing system 100 described in the first embodiment, the description of the contents common to the first embodiment will be omitted. This embodiment uses the same MRI images as the second embodiment. The following description will focus on the parts that are different from the second embodiment.
[0081] Fig. 13 is a diagram that illustrates the flow of the learning process and the inference process in the image processing system 100 in this embodiment, and related data, similar to Fig. 8. As shown in Fig. 13, in this embodiment, a separate learning model is used for each region obtained by dividing the original image.
[0082] 14(a) and (b), the process up to generation of partial image 901 from original image 900 is the same as in the second embodiment. Thereafter, for each partial image 901 at the same position (each region of the original image), image 902 having an abnormality to be recognized is classified as class A, and image 903 having no abnormality to be recognized (normal) is classified as class B, thereby generating teacher data. Specifically, as shown in FIG 14(c), upper left partial image 907, lower left partial image 908, upper right partial image 909, and lower right partial image 910 are classified as class A and class B, thereby generating teacher data.
[0083] The data transmitter / receiver 451 of the device 400 supplies an image 902 of class A and an image 903 of class B, and further, teacher data as necessary, for each region obtained by dividing the original image, to the learning data generation unit 250. The learning data generation unit 250 generates input data 801 and teacher data 802 as learning data for each region obtained by dividing the original image, and supplies them to the learning unit 251.
[0084] The learning unit 251 uses the learning data to learn each learning model 252, and completes the learning when the deviation amount L indicated by the loss function becomes less than a predetermined threshold. At this point, the learning model 252 becomes a trained model. Then, all or a part of the learning model 252 is supplied to the edge server 300 and stored as a trained model 352. Note that at least one of the learning data generation unit 250 and the learning unit 251 may be provided for each type of partial image, similar to the learning model 252.
[0085] The generation of the test image 904 and the input data 811 is similar to that in the second embodiment. In the second embodiment, the input data 811 is input to one trained model 352. In the present embodiment, four pieces of input data 811 corresponding to the lower left, upper left, lower right, and upper right test images 904 are input to the corresponding trained models 352.
[0086] From the four trained models 352, a judgment result for each of the four regions into which one original image is divided is obtained. The inference unit 351 obtains a judgment result for the original image from the four judgment results. For example, if even one of the four judgment results indicates that an abnormality exists, the inference unit 351 judges that an abnormality exists in the original image.
[0087] The present embodiment also provides the same effects as those of the second embodiment. In addition, by using a learning model for each region obtained by dividing the original image, noise between regions is reduced compared to the case where one learning model is used, and the determination accuracy is improved.
[0088] The trained model according to this embodiment was compared with a trained model trained by a conventional method that uses the original image as is, and the judgment accuracy was compared. As a result, it was confirmed that the trained model according to this embodiment had a higher judgment accuracy.
[0089] (Other embodiments) Although the above embodiment describes a configuration using one type of learning model, multiple learning models with different image sizes and configurations may be used in parallel to perform learning and inference. The final inference result can be obtained by, for example, ensemble judgment.
[0090] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions. [Explanation of symbols]
[0091] 100... processing system, 101... short-range wireless communication, 102... LAN, 103... router, 104... Internet, 105... mobile phone network, 200... cloud server, 300... edge server, 400... device
Claims
1. A generating means for generating learning data for a learning model from a first original image obtained by capturing an inspection object; A learning means for learning the learning model using the learning data; and an inference means for performing an inference process on input data generated from a second original image obtained by photographing an object to be inspected, using the learned learning model, the generating means generates a plurality of partial images from the first original image, and generates the learning data using the plurality of partial images; The inference means uses a partial image of the second original image as the input data.
13. An image processing device comprising:
2. the generating means generates a plurality of partial images by dividing the first original image and a partial image by extracting a central portion of the first original image; the inference means uses a partial image obtained by extracting a central portion of the second original image as the input data; The image processing device according to claim 1 .
3. The learning model is used to detect an abnormal portion from the first original image and the second original image; The image processing apparatus according to claim 2 , wherein the first original image and the second original image are images captured such that the abnormal portion is present in a central portion.
4. The generating means generates a plurality of partial images by dividing the first original image, the inference means uses, as the input data, a plurality of partial images obtained by dividing the second original image in the same manner as the first original image; The image processing device according to claim 1 .
5. The learning model is used to detect an abnormal portion from the first original image and the second original image; The image processing apparatus according to claim 4 , wherein the abnormal portion can be present at any position in the first original image and the second original image.
6. The image processing device according to claim 4 , wherein the learning model is provided for each of the partial images.
7. The image processing device has a plurality of devices communicably connected to each other, The learning means and the inference means are provided in separate devices. The image processing device according to claim 1 .
8. the learning means supplies a part or all of the learning model for which the learning has been completed to a device in which the inference means is provided; The inference means uses the learning model supplied from the learning means as the learned learning model. The image processing device according to claim 7.
9. 9. The image processing apparatus according to claim 7, wherein a device that supplies the first original image and the second original image is a device separate from a device in which the learning means and the inference means are provided.
10. An image processing method executed by an image processing device, comprising: A generation step of generating learning data for a learning model from a first original image obtained by capturing an inspection object; a learning step of learning the learning model using the learning data; An inference process of performing an inference process on input data generated from a second original image obtained by photographing an object to be inspected, using the learned learning model; The generating step includes: generating a plurality of sub-images from the first original image; generating the learning data using the plurality of partial images; In the inference step, a partial image of the second original image is used as the input data.
13. An image processing method comprising:
11. A program for causing a computer to function as each of the means included in the image processing device according to any one of claims 1 to 6.
Citation Information
Patent Citations
Symbol recognition device, and sign recognition apparatus for vehicle
JP2015092305A
Learning model creation device, type determination system, and learning model creation method
JP2020087211A
Image processing method, image processing apparatus, image pickup apparatus, and storage medium
US10354369B2
Image processing method, image processing apparatus, image capture apparatus, image processing program, and storage medium
WO2018037521A1
Image recognition system and image recognition method capable of suppressing false recognition
JP2020112926A