Image generation device, training device, image processing device, image generation method, training method, and image processing method
Patent Information
- Application Number
- JP2024546769
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2023-08-04
- Filing Date
- 2023-08-04
- Publication Date
- 2025-05-22
AI Technical Summary
Machine learning models for image recognition in non-destructive industrial product inspection face challenges in covering the diversity of background information due to the complexity of radiographic images, leading to a lack of effective training data, especially when detecting various defects and object characteristics.
An image generation device processes a first image to create multiple frequency-processed images emphasizing different frequency components, which are then combined and normalized to generate additional training data, expanding the variations of image sets used for machine learning, and a learning device trains a mathematical model using these enhanced data sets to improve defect detection accuracy.
The expanded training data set enhances the accuracy of defect detection models by emphasizing high-frequency defect images and adjusting background information, leading to improved detection of defects within radiographic images.
Abstract
Description
Image generation device, learning device, image processing device, image generation method, learning method, and image processing method
[0001] The disclosed technology relates to an image generation device, a learning device, an image processing device, an image generation method, a learning method, and an image processing method.
[0002] The following techniques are known as techniques related to machine learning of mathematical models used in image processing. For example, International Publication No. 2020 / 017211 describes a medical image learning device that includes an image acquisition unit that acquires a first image and a second image that have different spectral distributions from each other, an image processing unit that performs image processing on the first image to generate a third image, and a learning unit that uses the first to third images to train a recognizer to be applied to automatic recognition. The image processing unit generates a third image from the first image by performing at least one of image processing that suppresses signals in bands included in the spectral distribution of the first image that have characteristics different from those of the corresponding bands in the second image, and image processing that emphasizes signals in bands included in the first image that have characteristics identical to or similar to those of the corresponding bands in the second image.
[0003] International Publication No. 2020 / 175446 describes a learning method for machine learning a generative model that estimates, from a first image, a second image that includes image information with a higher resolution than the first image. In this method, a first training image that includes first resolution information with a lower resolution than the second image and a second training image that includes second resolution information with a higher resolution than the first training image and serves as a correct image corresponding to the first training image are used as training data.
[0004] In nondestructive inspection of industrial products using radiographic images, attempts have been made to detect defects inside products using mathematical models such as convolutional neural networks (CNNs) built using machine learning. In machine learning, a lack of training data is typically a problem given the complexity of the problem to be solved. To obtain as much variation as possible from a small amount of data, so-called "data augmentation" is often performed. For example, when training data is image data, adding processed images to the training data can improve the accuracy of the mathematical model being trained by increasing or decreasing the brightness (pixel value), changing the contrast, scaling, changing the blending ratio of color channels in color images, rotating, or inverting the original image.
[0005] In machine learning of mathematical models used for image recognition, even the common data augmentation methods described above may not be able to fully compensate for insufficient real data. In particular, radiological images used in non-destructive testing of industrial products contain a wide variety of background information, including defects to be detected, as well as undulations due to thickness variations in the object (product) itself and the shape characteristics of the object (product) itself. As a result, the variety of images that must be prepared as training data is enormous, and it is not easy to prepare an image set that encompasses the diversity of background information.
[0006] The disclosed technology has been made in consideration of the above points, and aims to expand the variety of image sets used as training data in machine learning.
[0007] An image generating device according to the disclosed technology includes at least one first processor that acquires a first image, generates a plurality of frequency-processed images by emphasizing or extracting different frequency components from the first image, performs different arithmetic processing on each of the plurality of frequency-processed images, generates at least one second image by combining the frequency components of the plurality of frequency-processed images that have been subjected to arithmetic processing, and outputs the first image and the second image as training data to be used in machine learning of a mathematical model that performs a predetermined inference on an input image.
[0008] The first processor may perform, as the arithmetic processing, a process of applying different weighting factors to the plurality of frequency-processed images. The first processor may perform, as the arithmetic processing, a process of generating a first frequency-processed image including relatively low frequency components and a second frequency-processed image including relatively high frequency components, and applying a relatively small weighting factor to the first frequency-processed image and a relatively large weighting factor to the second frequency-processed image. The first processor may perform, as the arithmetic processing, a process of magnifying, by different magnification factors, a difference between the average luminance value for each pixel of the plurality of frequency-processed images and the average luminance value for each pixel.
[0009] The first processor may generate a first frequency processed image by applying a filter process to the first image, and generate a second frequency processed image by subtracting a frequency component of the first frequency processed image from the first image. The first processor may perform a normalization process to normalize brightness of the first image and the second image. The first image may be a radiographic image including an image of a specific structural portion of the object in a specific frequency region, and the mathematical model may be a model that detects the image of the specific structural portion included in the radiographic image.
[0010] A learning device according to the disclosed technology includes at least one second processor. The second processor trains a mathematical model using the first image and the second image provided by the image generating device as training data. The second processor may perform a normalization process to normalize the luminance of the first image and the second image.
[0011] An image processing device according to the disclosed technology includes at least one third processor. The third processor uses the mathematical model trained by the learning device to output a detection result of an image of a specific structural portion of an object for an input image. The third processor may acquire the input image, generate a frequency-processed image by emphasizing or extracting specific frequency components for the input image, input the input image to the mathematical model to obtain a first inference result, input the frequency-processed image to the mathematical model to obtain a second inference result, and output a detection result by comprehensively evaluating the first inference result and the second inference result.
[0012] The image generation method according to the disclosed technology involves acquiring a first image, generating a plurality of frequency-processed images by emphasizing or extracting different frequency components from the first image, performing different arithmetic processing on each of the plurality of frequency-processed images, generating at least one second image by combining the frequency components of the plurality of frequency-processed images that have been subjected to arithmetic processing, and outputting the first image and the second image as training data to be used in machine learning of a mathematical model that performs a predetermined inference on an input image, by at least one first processor possessed by the image generation device.
[0013] The learning method according to the disclosed technology is such that at least one second processor included in the learning device executes a process of learning a mathematical model using the first image and the second image provided using the above-mentioned image generation method as training data.
[0014] The image processing method according to the disclosed technology is such that at least one third processor included in the image processing device uses the mathematical model learned by the above-mentioned learning method to execute a process of outputting a detection result of an image of a specific structural part of an object for an input image.
[0015] The disclosed technology makes it possible to expand the variety of image sets used as training data in machine learning.
[0016] 1 is a diagram illustrating an example of the configuration of an image processing system according to an embodiment of the disclosed technology. 2 is a diagram illustrating an example of the hardware configuration of an image generating device according to an embodiment of the disclosed technology. 3 is a functional block diagram illustrating an example of the functional configuration of an image generating device according to an embodiment of the disclosed technology. 4 is a diagram illustrating an example of the flow of processing in an image generating device according to an embodiment of the disclosed technology. 5 is a flowchart illustrating an example of the flow of image generation processing according to an embodiment of the disclosed technology. 6 is a diagram illustrating an example of the hardware configuration of a learning device according to an embodiment of the disclosed technology. 7 is a functional block diagram illustrating an example of the functional configuration of a learning device according to an embodiment of the disclosed technology. 8 is a diagram illustrating an example of the hardware configuration of an image processing device according to an embodiment of the disclosed technology. 9 is a functional block diagram illustrating an example of the functional configuration of an image processing device according to an embodiment of the disclosed technology. 10 is a diagram illustrating an example of the flow of processing in an image processing device according to an embodiment of the disclosed technology. 11 is a flowchart illustrating an example of the flow of defect detection processing according to an embodiment of the disclosed technology.
[0017] Hereinafter, an example of an embodiment of the disclosed technology will be described with reference to the drawings. In each drawing, the same or equivalent components and parts are given the same reference numerals, and redundant description will be omitted.
[0018] FIG. 1 is a diagram illustrating an example of the configuration of an image processing system 1 according to an embodiment of the disclosed technology. The image processing system 1 includes an image generation device 10, a learning device 20, and an image processing device 30. The image generation device 10 generates a second image by performing frequency processing on a first image, and outputs the first and second images as training data to be used in machine learning of a mathematical model that performs predetermined inference on an input image. The learning device 20 trains the mathematical model using the first and second images provided by the image generation device 10 as training data. The image processing device 30 uses the mathematical model trained by the learning device 20 to output a detection result of an image of a specific structural part of an object from the input image.
[0019] The image generation device 10, the learning device 20, and the image processing device 30 will be described in detail below. The following description will be given taking as an example a case where the images handled by the image processing system 1 are radiographic images acquired during non-destructive testing of industrial products. These radiographic images may include images of defects (flaws) occurring inside the object. The following description will be given taking as an example a case where the mathematical model handled by the image processing system 1 is a defect detection model that detects defects (flaws) contained in the radiographic images, and the image processing device 30 detects the defects (flaws) contained in the radiographic images using the defect detection model. The following description will also be given taking as an example a case where the image generation device 10, the learning device 20, and the image processing device 30 are configured as separate computers.
[0020] [Image Generation Device] Fig. 2 is a diagram showing an example of the hardware configuration of the image generation device 10. The image generation device 10 includes a CPU (Central Processing Unit) 101, a RAM (Random Access Memory) 102, a non-volatile memory 103, an input device 104, a display 105, and a network interface 106. These hardware components are connected to a bus 107. The display 105 is, for example, a liquid crystal display. The input device 104 includes, for example, a keyboard and a mouse, and may also include a proximity input device such as a touch panel display, and a voice input device such as a microphone. The network interface 106 is an interface for connecting the image generation device 10 to a network.
[0021] The nonvolatile memory 103 is a nonvolatile storage medium such as a hard disk or flash memory. An image generation program 110 is stored in the nonvolatile memory 103. The RAM 102 is a work memory for the CPU 101 to execute processing. The CPU 101 loads the image generation program 110 stored in the nonvolatile memory 103 into the RAM 102 and executes processing in accordance with the image generation program 110. The CPU 101 is an example of a "first processor" in the disclosed technology.
[0022] Fig. 3 is a functional block diagram showing an example of the functional configuration of the image generating device 10. When the CPU 101 executes an image generating program 110, the image generating device 10 functions as a first image acquisition unit 11, a frequency processing unit 12, an arithmetic processing unit 13, a synthesis processing unit 14, and an image output unit 15. Fig. 4 is a diagram showing an example of the processing flow in the image generating device 10. Below, the functions of each of the functional units constituting the image generating device 10 will be described with reference to Fig. 4.
[0023] The first image acquisition unit 11 acquires a first image 41. The first image 41 is a radiographic image acquired during non-destructive testing of industrial products. The first image 41 includes an image of a defect (flaw) occurring inside the object (product). Generally, an image can be decomposed into spatial frequency components for interpretation. The image of the defect (flaw) exists in a relatively high frequency region of the first image 41. For example, if the resolution of the first image 41 is 25 to 100 μm / px, the image of the defect (flaw) will be at least 1 to 2 px. Note that the first image 41 may be an original image or an original image that has undergone preprocessing other than the frequency processing described below, such as enlargement, reduction, rotation, and gradation processing. The first image acquisition unit 11 may perform this preprocessing.
[0024] The frequency processing unit 12 generates a first frequency-processed image 42 by performing first frequency processing 51 on the first image 41. The first frequency processing 51 is processing for extracting low-frequency components from the first image 41 or processing for enhancing the low-frequency components of the first image 41. The first frequency processing 51 may be, for example, filtering using a low-pass filter. The first frequency processing 52 is also known as a noise removal technique. It can also be implemented using a median filter, an averaging filter, a Gaussian filter, or a bilateral filter. In the first frequency-processed image 42, images of defects (flaws) contained in the first image 41 are almost completely removed, while background images such as undulations due to thickness changes in the object (product) itself and shape features of the object (product) itself remain.
[0025] The frequency processing unit 12 further performs second frequency processing 52 on the first image 41 to generate a second frequency-processed image 43. The second frequency processing 52 is a process for extracting high-frequency components from the first image 41 or a process for enhancing the high-frequency components of the first image 41. The second frequency processing 52 may be, for example, a filter process using a high-pass filter. The second frequency-processed image 43 can also be generated by subtracting the frequency components of the first frequency-processed image 42 from the first image 41. In the second frequency-processed image 43, background images such as undulations due to thickness changes in the object (product) itself and shape features of the object (product) itself are almost completely removed, while images of defects (flaws) included in the first image 41 remain. In other words, the images of defects (flaws) are more clearly depicted in the second frequency-processed image 43.
[0026] The arithmetic processing unit 13 performs a first arithmetic processing 53 on the first frequency processed image 42, and performs a second arithmetic processing 54 on the second frequency processed image 43, which is different from the first arithmetic processing 53. For example, the arithmetic processing unit 13 assigns a relatively small weighting coefficient w 1 (≧0) is added to the second frequency-processed image 43 as the first calculation process 53. 2 (>w 1 ) may be performed as the second calculation process 54. That is, the calculation processing unit 13 calculates the luminance value of each pixel of the first frequency processed image 42 by P1 x,y When we say, w 1 ×P1 x,y is output as the first frequency processed image 42A after the first arithmetic processing. In addition, the arithmetic processing unit 13 outputs the luminance value of each pixel of the second frequency processed image 43 as P2 x,y When we say, w 2 ×P2 x,y is output as the second frequency processed image 43A after the second arithmetic processing. Note that the weighting coefficient w 1 The calculation processing unit 13 may set the weighting coefficient w 1 and w 2By changing the combination of the first and second frequency processed images 42A and 43A after the arithmetic processing, a plurality of image pairs each consisting of the first and second frequency processed images 42A and 43A may be generated.
[0027] The calculation processing unit 13 may perform a process of contrast enhancement with different intensities on the first frequency processed image 42 and the second frequency processed image 43, instead of or in addition to the process of adding weighting coefficients with different magnitudes to the first frequency processed image 42 and the second frequency processed image 43. Specifically, the first calculation processing 53 performs a process of contrast enhancement with different intensities on the first frequency processed image 42 and the second frequency processed image 43, by multiplying the difference between the average luminance value of each pixel of the first frequency processed image 42 by a magnification m 1 (>0) is performed as the first calculation process 53, and the difference between the average value of the luminance value of each pixel and the second frequency processed image 43 is calculated as m 1 Magnification m greater than 2 (>m 1 ) may be performed as the second calculation process 54. 1 and m 2 By changing the combination of the first and second frequency processed images 42A and 43A after the arithmetic processing, a plurality of image pairs each consisting of the first and second frequency processed images 42A and 43A may be generated.
[0028] The synthesis processing unit 14 generates at least one second image 44 by performing synthesis processing 55, which synthesizes the frequency components of the first frequency processed image 42A that has been subjected to the first arithmetic processing 53 and the frequency components of the second frequency processed image 43A that has been subjected to the second arithmetic processing 54. For example, the arithmetic processing unit 13 assigns weighting coefficients w 1 , w 2 When the process of adding w is performed, the first frequency processed image 42 and the second frequency processed image 43 are subjected to weighted addition processing by the synthesis process 55 to generate the second image 44. 1 <w 2 By doing so, it is possible to obtain a second image 44 in which the image of a defect (flaw) existing in the high frequency region is more emphasized. 1By setting r to zero, the second frequency processed image 43 essentially becomes the second image 44, and the image of the defect (flaw) is depicted more clearly in the second image 44. When a plurality of image pairs each consisting of the first frequency processed image 42A and the second frequency processed image 43A after the arithmetic processing are generated, the synthesis processing unit 14 performs synthesis processing 55 on each of the plurality of image pairs. As a result, a plurality of second images 44 are generated from one first image 41.
[0029] The image output unit 15 outputs the first image 41 acquired by the first image acquisition unit 11 and at least one second image 44 generated based on the first image 41 as training data to be used in machine learning of a mathematical model that performs predetermined inference on the input image. The image output unit 15 may output the first image 41 and the second image 44 after performing a normalization process to normalize the luminance of each pixel on these images. The normalization process is expressed by, for example, the following equation (1). In equation (1), X norm is the brightness value after normalization. X is the brightness value before normalization. X max is the maximum value of brightness (e.g., 255). min is the minimum value of luminance (for example, 0). By performing the normalization process, it is possible to prevent the luminance distribution from significantly deviating between the first image 41 and the second image 44. norm = (X-X min ) / (X max -X min ) ... (1)
[0030] 5 is a flowchart showing an example of the flow of image generation processing performed by the CPU 101 executing the image generation program 110. The image generation program 110 is executed when, for example, the user operates the input device 104 to instruct the start of processing.
[0031] In step S1, the first image acquisition unit 11 acquires a first image 41. The first image 41 may be an original image or an original image that has been subjected to preprocessing other than frequency processing, such as enlargement, reduction, rotation, and gradation processing. The first image acquisition unit 11 may perform this preprocessing.
[0032] In step S2, the frequency processing unit 12 generates a first frequency processed image 42 and a second frequency processed image 43 by emphasizing or extracting different frequency components from the first image 41 acquired in step S1. The first frequency processed image 42 mainly contains low frequency components.
[0033] In step S3, the calculation processing unit 13 performs different calculation processes on the first frequency processed image 42 and the second frequency processed image 43 generated in step S2. For example, the calculation processing unit 13 assigns a relatively small weighting coefficient w 1 (≧0) is added to the second frequency-processed image 43 as the first calculation process 53. 2 (>w 1 ) is added as the second calculation process 54. 1 and w 2 By changing the combination of the first and second frequency processed images 42A and 43A after the arithmetic processing, a plurality of image pairs each consisting of the first and second frequency processed images 42A and 43A may be generated.
[0034] In step S4, the synthesis processing unit 14 synthesizes the frequency components of the first frequency processed image 42A and the second frequency processed image 43A that have been subjected to the arithmetic processing in step S3, thereby generating at least one second image 44. When multiple image sets each consisting of the first frequency processed image 42A and the second frequency processed image 43A after the arithmetic processing have been generated, the synthesis processing unit 14 performs synthesis processing 55 on each of the image sets. As a result, multiple second images 44 are generated.
[0035] In step S5, the image output unit 15 outputs the first image 41 acquired in step S1 and at least one second image 44 generated in step S4 as training data to be used in machine learning of a mathematical model that performs predetermined inference on the input image. The image output unit 15 may output the first image 41 and the second image 44 after performing a normalization process that normalizes the luminance of each pixel on these images.
[0036] As described above, the image generating device 10 according to an embodiment of the disclosed technology acquires a first image 41 and generates a plurality of frequency-processed images by emphasizing or extracting different frequency components from the first image 41. The image generating device 10 performs different arithmetic processing on each of the plurality of frequency-processed images and generates at least one second image 44 by combining the frequency components of the plurality of frequency-processed images that have been subjected to arithmetic processing. The image generating device 10 outputs the first image 41 and the second image 44 as training data to be used in machine learning of a mathematical model that performs predetermined inference on an input image.
[0037] In machine learning for mathematical models of image recognition, common data augmentation techniques, such as increasing or decreasing the brightness (pixel value), changing the contrast, scaling, changing the blending ratio of color channels in color images, rotating, or inverting, often fail to fully compensate for the lack of real data. In particular, radiological images used in nondestructive testing of industrial products contain a wide variety of background information, including defects to be detected, as well as undulations due to thickness variations in the object (product) itself and the shape features of the object (product) itself. As a result, the variety of images that must be prepared as training data is enormous, and it is not easy to prepare an image set that encompasses the diversity of background information.
[0038] According to the image generating device 10 relating to an embodiment of the disclosed technology, a second image 44 is generated based on a first image 41, which is an original image, and the first image 41 and the second image 44 are provided as training data, so that the training data can be expanded.
[0039] The second image 44 is an image obtained by performing different arithmetic processing on the first frequency-processed image 42 and the second frequency-processed image 43 and synthesizing these frequency components. For example, by varying the combination of weighting factors used in the arithmetic processing, it is possible to generate multiple second images 44. This enables further data expansion and adds diversity to the training data. In this way, the image generating device 10 according to an embodiment of the disclosed technology makes it possible to expand the variety of image sets used as training data in machine learning.
[0040] Furthermore, since the second image 44 is an image obtained by combining the frequency components of the first frequency processed image 42 and the second frequency processed image 43, it is possible to independently adjust the appearance of the image of the defect (flaw) present in the high frequency region and the image of the background present in the low frequency region. For example, if a relatively small weighting coefficient w 1 is added to the second frequency-processed image 43, which mainly contains high-frequency components, and a relatively large weighting coefficient w 2 By performing an arithmetic process of adding w to the frequency components of these images and synthesizing them, it is possible to obtain a second image 44 in which the image of a defect (flaw) existing in the high frequency region is more emphasized. 1 By setting the second frequency processed image 43 to zero, the second image 44 is essentially formed, and the image of the defect (flaw) is depicted more clearly in the second image 44. In machine learning of a defect detection model that detects defects (flaws) contained in an input image, by using an image in which the defect (flaw) to be detected is emphasized as training data, the image features of the defect (flaw) can be more clearly taught to the defect detection model. This makes it possible to improve the accuracy of defect detection in the defect detection model.
[0041] In the present embodiment, the case where the frequency processing unit 12 generates two frequency-processed images (the first frequency-processed image 42 and the second frequency-processed image 43) has been exemplified, but the disclosed technology is not limited to this example. The frequency processing unit 12 may generate three or more frequency-processed images by emphasizing or extracting different frequency components from the first image 41. This further increases the variety of arithmetic processing in the arithmetic processing unit 13, enabling the generation of more diverse second images 44.
[0042] [Learning Device] Fig. 6 is a diagram showing an example of the hardware configuration of the learning device 20. The hardware configuration of the learning device 20 is similar to the hardware configuration of the image generating device 10. The learning device 20 includes a CPU 201, a RAM 202, a non-volatile memory 203, an input device 204, a display 205, and a network interface 206. These pieces of hardware are connected to a bus 207. The CPU 201 is an example of a "second processor" in the disclosed technology.
[0043] The non-volatile memory 203 stores a learning program 210, a defect detection model 220, and training data 60. The defect detection model 220 is a mathematical model using a CNN constructed to predict class labels for each pixel of an input image. The defect detection model 220 may be, for example, a known encoder-decoder model (ED-CNN: Encoder-Decoder Convolutional Neural Network). The encoder-decoder model is a model consisting of an encoder that extracts features from an image using a convolutional layer and a decoder that outputs a probability map based on the extracted features. The probability map is the result of deriving the probability that each pixel of the input image belongs to a certain class for each pixel. The defect detection model 220 is constructed by machine learning using multiple radiographic images to which correct labels have been assigned as training data. The training data 60 is a data set in which a first image 41, a second image 44, and the correct labels 45 and 46 assigned to these images, which are supplied from the image generating device 10, form one unit. The non-volatile memory 203 stores multiple training data 60. The training data 60 is used for machine learning of the defect detection model 220 .
[0044] 7 is a functional block diagram showing an example of the functional configuration of the learning device 20. In the learning device 20, the CPU 201 executes a learning program 210, thereby functioning as a teacher data acquisition unit 21 and a learning unit 22. The teacher data acquisition unit 21 acquires teacher data 60 stored in the non-volatile memory 203. The learning unit 22 uses the teacher data 60 acquired by the teacher data acquisition unit 21 to train a defect detection model 220 using, for example, an error backpropagation method.
[0045] Note that the normalization process for normalizing the luminance of each pixel for the first image 41 and the second image 44, which are the teacher data 60, may be performed by the teacher data acquisition unit 21 or the learning unit 22 of the learning device 20, instead of by the image output unit 15 of the image generating device 10. Performing the normalization process can prevent a significant deviation in the luminance distribution between the first image 41 and the second image 44. If the luminance distributions of the first image 41 and the second image 44 significantly deviate from each other, the defect detection model 220 may unintentionally distinguish between the first image 41 and the second image 42, potentially failing to extract the features of a defect (flaw) commonly present in these images. Performing the normalization process on the first image 41 and the second image 44 can prevent the defect detection model 220 from unintentionally distinguishing between the first image 41 and the second image 44.
[0046] As described above, the learning device 20 according to an embodiment of the disclosed technology uses the first image 41 and the second image 44 generated by the image generating device 10 as training data to train the defect detection model 220. According to the learning device 20 according to this embodiment, the defect detection model 220 is trained using the training data 60 extended by the second image 44, and therefore the accuracy of defect detection in the defect detection model 220 can be improved.
[0047] Note that each process for generating the second image 44, which is performed in the image generating device 10, may be performed in the learning device 20. Furthermore, each process for generating the second image 44 may be performed in the input layer of the defect detection model 220.
[0048] [Image Processing Device] Fig. 8 is a diagram showing an example of the hardware configuration of the image processing device 30. The hardware configuration of the image processing device 30 is similar to the hardware configuration of the image generation device 10. The image processing device 30 includes a CPU 301, a RAM 302, a non-volatile memory 303, an input device 304, a display 305, and a network interface 306. These pieces of hardware are connected to a bus 307. The CPU 301 is an example of a "third processor" in the disclosed technology. The non-volatile memory 303 stores a defect detection program 310 and a trained defect detection model 320. The defect detection model 320 is a mathematical model trained by the learning device 20.
[0049] Fig. 9 is a functional block diagram showing an example of the functional configuration of the image processing device 30. The image processing device 30 functions as an input image acquisition unit 31, a frequency processing unit 32, an inference result acquisition unit 33, and a detection result output unit 34 when a CPU 301 executes a defect detection program 310. Fig. 10 is a diagram showing an example of the flow of processing in the image processing device 30. Below, the functions of each of the functional units constituting the image processing device 30 will be described with reference to Fig. 10.
[0050] The input image acquisition unit 31 acquires an input image 71 to be input as a processing target for the image processing device 30. The input image 71 is a radiological image similar to the first image 41. The input image 71 may be an original image that has been subjected to preprocessing other than frequency processing, such as enlargement, reduction, rotation, and gradation processing. The input image acquisition unit 31 may perform this preprocessing.
[0051] The frequency processing unit 32 generates a frequency-processed image 72 by performing frequency processing 82 on the input image 71. The frequency processing 82 may be, for example, processing to extract high-frequency components from the input image 71 or processing to emphasize high-frequency components of the input image 71. The frequency processing 82 may be processing to extract or emphasize the same frequency components as the frequency components extracted or emphasized in the second frequency processing 52 performed in the image generating device 10, or processing to extract or emphasize different frequency components. The frequency processing unit 32 may generate multiple frequency-processed images 72 by emphasizing or extracting different frequency components from each other for the input image 71.
[0052] The inference result acquisition unit 33 acquires a first inference result 73 by inputting the input image 71 to the trained defect detection model 320. The inference result acquisition unit 33 further acquires a second inference result 74 by inputting the frequency-processed image 72 to the trained defect detection model 320. When a plurality of frequency-processed images 72 are generated, the inference result acquisition unit 33 acquires a plurality of second inference results 74 corresponding to each of the plurality of frequency-processed images 72. The first inference result 73 and the second inference result 74 may be, for example, a probability map that is a result of deriving, for each pixel in the image, the probability that the pixel in question indicates a defect (flaw).
[0053] The detection result output unit 34 outputs a defect image detection result 75 for the input image 71 by comprehensively evaluating the first inference result 73 and the second inference result 74. For example, the detection result output unit 34 may generate the overall probability map by calculating the average value of the probability map as the first inference result 73 and the probability map as the second inference result 74. Alternatively, the detection result output unit 34 may generate the overall probability map by calculating a weighted average value of the probability map as the first inference result 73 and the probability map as the second inference result 74. Alternatively, the detection result output unit 34 may generate the overall probability map by applying the maximum value to the probability map as the first inference result 73 and the probability map as the second inference result 74. The detection result output unit 34 may derive the detection result 75 by, for example, performing threshold processing on the overall probability map.
[0054] 11 is a flowchart showing an example of the flow of defect detection processing performed by the CPU 201 executing the defect detection program 310. The defect detection program 310 is executed when, for example, the user operates the input device 304 to instruct the start of processing.
[0055] In step S31, the input image acquisition unit 31 acquires an input image 71 to be input as a processing target for the image processing device 30. The input image 71 is a radiological image similar to the first image 41. In step S32, the frequency processing unit 32 generates a frequency-processed image 72 by emphasizing or extracting specific frequency components from the input image 71 acquired in step S31.
[0056] In step S33, the inference result acquisition unit 33 acquires a first inference result 73 by inputting the input image 71 acquired in step S31 into the trained defect detection model 320. In step S34, the inference result acquisition unit 33 acquires a second inference result 74 by inputting the frequency-processed image generated in step S32 into the trained defect detection model 320. In step S35, the detection result output unit 34 comprehensively evaluates the first inference result 73 acquired in step S33 and the second inference result 74 acquired in step S34, and outputs a defect (flaw) image detection result 75 for the input image 71 acquired in step S31.
[0057] As described above, the image processing device 30 according to the embodiment of the disclosed technology uses the defect detection model 320 trained by the learning device 20 to output a defect (flaw) image detection result for the input image 71. According to the image processing device 30 according to the present embodiment, defect (flaw) detection is performed using the defect detection model 220 trained using the teacher data 60 extended by the second image 44, thereby improving the accuracy of defect detection.
[0058] Furthermore, since the detection result 75 is output by comprehensively evaluating the first inference result 73 and the second inference result 74, it is possible to appropriately balance the detection accuracy and detection sensitivity. Note that the defect (flaw) image detection result 75 may be output based on only one of the first inference result 73 and the second inference result 74. When only the first inference result 73 is used, the frequency processing unit 32 and the frequency-processed image 72 are not necessary.
[0059] In the above description, the image generation device 10, the learning device 20, and the image processing device 30 are configured as separate computers, but the image generation device 10, the learning device 20, and the image processing device 30 may be configured as one or two computers. For example, the image generation device 10 and the learning device 20 may be configured as the same computer. Furthermore, the learning device 20 and the image processing device 30 may be configured as the same computer. Furthermore, the image generation device 10, the learning device 20, and the image processing device 30 may be configured as the same computer.
[0060] Furthermore, in the above description, an example has been given in which the images handled by the image processing system 1 are radiographic images acquired in non-destructive testing of industrial products, but the disclosed technology is not limited to this example. The disclosed technology can be applied to any image that includes an image of a specific structural part of an object in a specific frequency range (e.g., an optical image, an MRI (Magnetic Resonance Imaging) image, an SEM (Scanning Electron Microscope) image, an ultrasound image, etc.). Furthermore, in the above description, an example has been given in which the mathematical model to be learned is a defect detection model that detects defects (flaws) contained in a radiographic image, but the disclosed technology is not limited to this example. The disclosed technology can be applied to a mathematical model that detects or recognizes a specific structural part of an object that exists in a specific frequency range in an image.
[0061] In the above embodiment, the following various processors can be used as the hardware structure of processing units that perform various processes, such as the first image acquisition unit 11, frequency processing unit 12, arithmetic processing unit 13, synthesis processing unit 14, image output unit 15, teacher data acquisition unit 21, learning unit 22, input image acquisition unit 31, frequency processing unit 32, inference result acquisition unit 33, and detection result output unit 34. As described above, the various processors include CPUs and GPUs, which are general-purpose processors that execute software (programs) and function as various processing units, as well as dedicated electrical circuits, such as programmable logic devices (PLDs) that are processors whose circuit configuration can be changed after manufacture, such as FPGAs, and application-specific integrated circuits (ASICs), which are processors having a circuit configuration specifically designed to perform specific processes.
[0062] A single processing unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, multiple processing units may be configured with a single processor.
[0063] Examples of configuring multiple processing units with a single processor include: first, a form in which one processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server computers, and this processor functions as multiple processing units; second, a form in which a processor is used to realize the functions of an entire system including multiple processing units with a single IC (Integrated Circuit) chip, as typified by systems on chips (SoCs); thus, various processing units are configured as hardware structures using one or more of the above-mentioned various processors; furthermore, the hardware structures of these various processors can be, more specifically, electrical circuits combining circuit elements such as semiconductor elements.
[0064] Furthermore, in the above embodiment, the various programs are described as being pre-stored (installed) in non-volatile memory, but this is not limiting. The various programs may be provided in a form recorded on a recording medium such as a CD-ROM (Compact Disc Read Only Memory), a DVD-ROM (Digital Versatile Disc Read Only Memory), or a USB (Universal Serial Bus) memory. The various programs may also be downloaded from an external device via a network. In other words, the programs described in this embodiment (i.e., program products) may be provided on a recording medium or may be distributed from an external computer.
[0065] The following supplementary note is further disclosed regarding the above-described embodiments: (Supplementary Note 1) An image generating device having at least one first processor, wherein the first processor acquires a first image, generates a plurality of frequency-processed images by emphasizing or extracting different frequency components from the first image, performs different arithmetic processing on each of the plurality of frequency-processed images, generates at least one second image by combining the frequency components of the plurality of frequency-processed images that have been subjected to the arithmetic processing, and outputs the first image and the second image as training data to be used in machine learning of a mathematical model that performs a predetermined inference on an input image.
[0066] (Supplementary Note 2) The image generating device according to Supplementary Note 1, wherein the first processor performs, as the arithmetic processing, a process of adding weighting coefficients different from one another to the plurality of frequency-processed images.
[0067] (Supplementary Note 3) The image generating device according to Supplementary Note 1 or Supplementary Note 2, wherein the first processor generates a first frequency processed image including relatively low frequency components and a second frequency processed image including relatively high frequency components, and performs the arithmetic processing of adding a relatively small weighting factor to the first frequency processed image and a relatively large weighting factor to the second frequency processed image.
[0068] (Supplementary Note 4) The image generating device according to any one of Supplementary Note 1 to Supplementary Note 3, wherein the first processor performs, as the arithmetic processing, a process of enlarging a difference between an average value of luminance values for each pixel of the plurality of frequency-processed images and the average value of luminance values for each pixel by a magnification factor different from each other.
[0069] (Supplementary Note 5) The image generating device according to any one of Supplementary Note 1 to Supplementary Note 4, wherein the first processor generates a first frequency processed image by performing a filter process on the first image, and generates a second frequency processed image by subtracting a frequency component of the first frequency processed image from the first image.
[0070] (Supplementary Note 6) The image generating device according to any one of Supplementary Notes 1 to 5, wherein the first processor performs a normalization process to normalize luminance of the first image and the second image.
[0071] (Supplementary Note 7) The image generating device according to any one of Supplementary Notes 1 to 6, wherein the first image is a radiographic image including an image of a specific structural part of an object in a specific frequency region, and the mathematical model is a model that detects the image of the specific structural part included in the radiographic image.
[0072] (Supplementary Note 8) A learning device having at least one second processor, wherein the second processor trains the mathematical model using the first image and the second image provided from the image generating device described in any one of Supplementary Note 1 to Supplementary Note 7 as training data.
[0073] (Supplementary Note 9) The learning device according to Supplementary Note 8, wherein the second processor performs a normalization process to normalize luminance of the first image and the second image.
[0074] (Supplementary Note 10) An image processing device having at least one third processor, wherein the third processor uses the mathematical model learned by the learning device described in Supplementary Note 8 or Supplementary Note 9 to output a detection result of an image of a specific structural part of an object for an input image.
[0075] (Supplementary Note 11) The image processing device described in Supplementary Note 10, wherein the third processor acquires the input image, generates a frequency-processed image by emphasizing or extracting specific frequency components from the input image, acquires a first inference result by inputting the input image into the mathematical model, acquires a second inference result by inputting the frequency-processed image into the mathematical model, and outputs the detection result by comprehensively evaluating the first inference result and the second inference result.
[0076] (Supplementary Note 12) An image generation method in which at least one first processor included in an image generation device executes the following processes: acquire a first image; generate a plurality of frequency-processed images by emphasizing or extracting different frequency components from the first image; perform different arithmetic processing on each of the plurality of frequency-processed images; generate at least one second image by combining the frequency components of the plurality of frequency-processed images that have been subjected to the arithmetic processing; and output the first image and the second image as training data to be used in machine learning of a mathematical model that performs a predetermined inference on an input image.
[0077] (Supplementary Note 13) A learning method in which at least one second processor of a learning device executes a process of learning the mathematical model using the first image and the second image provided using the image generation method described in Supplementary Note 12 as training data.
[0078] (Supplementary Note 14) An image processing method in which at least one third processor of an image processing device performs a process of outputting a detection result of an image of a specific structural part of an object for an input image using the mathematical model trained by the learning method described in Supplementary Note 13.
[0079] The disclosure of Japanese Patent Application No. 2022-146397, filed on September 14, 2022, is incorporated herein by reference in its entirety. In addition, all documents, patent applications, and technical standards described herein are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually indicated to be incorporated by reference.
Claims
1. An image generation device having at least one first processor, wherein the first processor acquires a first image, generates a plurality of frequency-processed images by emphasizing or extracting different frequency components from the first image, performs different arithmetic processing on each of the plurality of frequency-processed images, generates at least one second image by combining the frequency components of the plurality of frequency-processed images that have been subjected to the arithmetic processing, and outputs the first image and the second image as training data to be used in machine learning of a mathematical model that performs predetermined inference on an input image.
2. The image generating device according to claim 1, wherein the first processor performs the arithmetic processing of adding different weighting coefficients to the plurality of frequency-processed images.
3. The image generating device according to claim 2, wherein the first processor generates a first frequency processed image containing relatively low frequency components and a second frequency processed image containing relatively high frequency components, and performs the arithmetic processing of adding a relatively small weighting factor to the first frequency processed image and a relatively large weighting factor to the second frequency processed image.
4. The image generating device according to claim 1, wherein the first processor performs the arithmetic processing of magnifying the difference between the average brightness value of each pixel and each of the plurality of frequency-processed images by different magnifications.
5. The image generating device according to claim 1, wherein the first processor generates a first frequency processed image by applying filter processing to the first image, and generates a second frequency processed image by subtracting frequency components of the first frequency processed image from the first image.
6. The image generating device according to claim 1, wherein the first processor performs a normalization process for normalizing the luminance of the first image and the second image.
7. The image generating device according to claim 1, wherein the first image is a radiological image containing an image of a specific structural part of an object in a specific frequency region, and the mathematical model is a model that detects the image of the specific structural part contained in the radiological image.
8. A learning device having at least one second processor, wherein the second processor uses the first image and the second image provided from the image generating device described in any one of claims 1 to 7 as training data to learn the mathematical model.
9. The learning device according to claim 8, wherein the second processor performs a normalization process to normalize the brightness of the first image and the second image.
10. An image processing device having at least one third processor, wherein the third processor uses the mathematical model learned by the learning device described in claim 8 to output a detection result of an image of a specific structural part of an object for an input image.
11. The image processing device described in claim 10, wherein the third processor: acquires the input image; generates a frequency-processed image by emphasizing or extracting specific frequency components from the input image; acquires a first inference result by inputting the input image into the mathematical model; acquires a second inference result by inputting the frequency-processed image into the mathematical model; and outputs the detection result by comprehensively evaluating the first inference result and the second inference result.
12. An image generation method in which at least one first processor of an image generation device executes the following processes: acquire a first image; generate a plurality of frequency-processed images by emphasizing or extracting different frequency components from the first image; perform different arithmetic processing on each of the plurality of frequency-processed images; generate at least one second image by combining the frequency components of the plurality of frequency-processed images that have been subjected to the arithmetic processing; and output the first image and the second image as training data to be used in machine learning of a mathematical model that performs a predetermined inference on an input image.
13. A learning method in which at least one second processor of a learning device executes a process of training the mathematical model using the first image and the second image provided using the image generation method described in claim 12 as training data.
14. An image processing method in which at least one third processor of an image processing device performs a process of outputting a detection result of an image of a specific structural part of an object for an input image using the mathematical model trained by the training method described in claim 13.