Image estimation method, evaluation value estimation method, and image estimation device

The image estimation method enhances estimation performance by standardizing patterns in training images and restoring necessary information, addressing the challenges of limited training data and overfitting in machine learning engines.

JP7765351B2Active Publication Date: 2025-11-06HITACHI LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022098864
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-06-20
Publication Date
2025-11-06
Estimated Expiration
2042-06-20

AI Technical Summary

Technical Problem

Machine learning-based estimation engines face challenges in achieving high estimation performance due to the difficulty in collecting a large number of training images and the risk of overfitting when handling a wide variety of pattern variations.

Method used

An image estimation method that applies a transformation process to training and input images to standardize patterns, reducing the number of training images needed by uniforming them through image processing, and uses inverse transformations to restore necessary information in the output.

Benefits of technology

Improves estimation performance even with a small number of training images by reducing the training load and minimizing overfitting, while maintaining image quality and pattern variation in the output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007765351000001
    Figure 0007765351000001
  • Figure 0007765351000002
    Figure 0007765351000002
  • Figure 0007765351000003
    Figure 0007765351000003
Patent Text Reader

Abstract

To improve estimation performance of an estimation engine even for a less learning image.SOLUTION: An image estimation method includes steps of: acquiring multiple learning images f imaging a target object for a learning; generating multiple conversion images f' by performing a conversion processing on the multiple learning images f; determining an internal parameter of an estimation engine in a machine learning type using the multiple conversion images f'; acquiring an input image h imaging a real target object; generating a converted image h' by performing the conversion processing on the input image h; estimating an estimation image H' by inputting the converted image h' into the trained estimation engine; generating an inverse transformation estimated image H by performing an inverse transformation processing of a part of the conversion processing or all processing of the conversion processing on the estimated image H'; and outputting the inverse conversion estimated image H as an output image. The conversion processing equalizes the pattern of the multiple learning images f related to the conversion item by the image processing to adjust at least one or more conversion items.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image estimation method, an evaluation value estimation method, and an image estimation device. [Background technology]

[0002] In recent years, the performance of machine learning has improved dramatically with the proposal of deep network models such as convolutional neural networks (CNNs). Many image processing methods that utilize estimation engines based on machine learning have been proposed. For example, as an example of application to visual inspection, Patent Document 1 discloses a method for automatically inspecting welds for geometric defects using machine learning. Furthermore, image processing based on machine learning covers a wide range of applications, including semantic segmentation and recognition, image classification, image conversion, and image quality improvement.

[0003] For example, when the objective is to improve image quality by estimating a high-quality image from an input image, the estimation engine updates its internal parameters (such as network weights or biases) so as to reduce the difference between the estimated image output from the estimation engine using an image of the object to be evaluated for learning (learning image) as input and a high-quality image (correct image) that is the correct estimated image that has been taught in advance.

[0004] Even when the purpose is image evaluation, which estimates the evaluation value of an evaluation object from an input image, the learning of the estimation engine, similar to the case of image quality improvement, updates the internal parameters of the estimation engine so as to reduce the difference between an estimated image output from the estimation engine using a learning image as an input and a correct evaluation value (correct evaluation value) that is taught in advance. The evaluation value can be an evaluation result such as the quality level of the evaluation object, the presence or absence of defects, and the degree of abnormality. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] International Publication No. 2020 / 129617 Summary of the Invention [Problem to be solved by the invention]

[0006] In order for the above-mentioned machine learning-based estimation engine to achieve high estimation performance for input images with a wide variety of pattern variations, it is necessary to train it with a large number of images that correspond to the pattern variations.

[0007] However, capturing images is generally a labor-intensive task, and in reality, it can be difficult to collect a large number of training images in advance on a production line, etc. Furthermore, if a large number of pattern variations are taught to an estimation engine, the learning load increases, and there is a risk of overfitting in machine learning. Overfitting is a state in which high performance is achieved for training images, but performance is not achieved for untrained data, resulting in a decline in generalization performance.

[0008] The present invention has been made in view of the above circumstances, and has as its object to improve the estimation performance of an estimation engine even with a small number of training images. [Means for solving the problem]

[0009] The present application includes a number of means for solving at least part of the above problems, examples of which are as follows.

[0010] In order to solve the above problem, an image estimation method according to one embodiment of the present invention is an image estimation method executed by a computer, the image estimation method including: a first training image acquisition step of acquiring a plurality of training images f obtained by capturing an object for training; a first transformed image generation step of applying a transformation process to the plurality of training images f to generate a plurality of transformed images f'; a learning step of determining internal parameters of a machine learning estimation engine using the plurality of transformed images f'; a second image acquisition step of acquiring an input image h obtained by capturing an actual object; a second transformed image generation step of applying the transformation process to the input image h to generate a transformed image h'; an estimated image generation step of inputting the transformed image h' to the trained estimation engine to estimate an estimated image H'; an inverse transformed estimated image generation step of applying an inverse transformation process of some or all of the transformation process to the estimated image H' to generate an inverse transformed estimated image H; and an output step of outputting the inverse transformed estimated image H as an output image, wherein the transformation process uniforms the patterns of the plurality of training images f with respect to the transformation items by image processing that adjusts at least one or more transformation items. [Effects of the Invention]

[0011] According to the present invention, the estimation performance of the estimation engine can be improved even with a small number of training images.

[0012] Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a schematic diagram showing an example of the overall processing sequence performed by the visual inspection system according to the first embodiment. [Figure 2] FIG. 2 is a schematic diagram showing an example of input and output images of the visual inspection system according to the first embodiment. [Figure 3] FIG. 3 is a diagram schematically showing an example in which the conversion items are position, rotation, and magnification. [Figure 4] FIG. 4 is a diagram schematically illustrating an example in which the transformation items are rotation and distortion. [Figure 5A] FIG. 5A is a diagram schematically illustrating an example in which the conversion items are brightness, contrast, and noise. [Figure 5B] FIG. 5B is a diagram schematically illustrating an example in which the conversion items are shading and shadow. [Figure 6] FIG. 6 is a schematic diagram (part 1) showing an example of a learning method for an estimation engine. [Figure 7] FIG. 7 is a schematic diagram (part 2) illustrating an example of a learning method for an estimation engine. [Figure 8A] FIG. 8A is a schematic diagram (part 1) showing an example of a GUI (Graphical User Interface). [Figure 8B] FIG. 8B is a schematic diagram (part 2) showing an example of the GUI. [Figure 9] FIG. 9 is a schematic diagram showing an example of the overall processing sequence performed by the visual inspection system according to the second embodiment. [Figure 10] FIG. 10 is a schematic diagram illustrating an example of a learning method for the estimation engine. [Figure 11] FIG. 11 is a diagram showing the hardware configuration of the appearance inspection system according to each of the first and second embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0014] An embodiment of the present invention will be described below with reference to the drawings. In all drawings used to describe the embodiment, the same components are generally designated by the same reference numerals, and repeated description thereof will be omitted where appropriate. It goes without saying that, in the following embodiments, the components (including element steps, etc.) are not necessarily essential unless otherwise specified or considered to be clearly essential in principle. It goes without saying that the terms "consisting of A," "made of A," "having A," and "including A" do not exclude other elements, unless otherwise specified to include only the relevant element. Similarly, in the following embodiments, when referring to the shape, positional relationship, etc. of components, etc., this includes those that are substantially similar or similar to the shape, etc., unless otherwise specified or considered to be clearly essential in principle.

[0015] First Embodiment <1. Overall processing sequence in the visual inspection system> 1 is a schematic diagram showing an example of an overall processing sequence performed by the visual inspection system according to this embodiment. The visual inspection system according to this embodiment is an example of an image estimation device, and executes an image estimation method. The image estimation method is a computer-executed method, and includes: a first training image acquisition step of acquiring a plurality of training images f obtained by capturing an object for training; a first transformed image generation step of applying a transformation process to the plurality of training images f to generate a plurality of transformed images f'; a learning step of determining internal parameters of a machine learning estimation engine using the plurality of transformed images f'; a second image acquisition step of acquiring an input image h obtained by capturing an actual object; a second transformed image generation step of applying the transformation process to the input image h to generate a transformed image h'; an estimated image generation step of inputting the transformed image h' to the trained estimation engine to estimate an estimated image H'; an inverse transformed estimated image generation step of applying an inverse transformation process of some or all of the transformation process to the estimated image H' to generate an inverse transformed estimated image H; and an output step of outputting the inverse transformed estimated image H as an output image, wherein the transformation process is characterized in that the patterns of the plurality of training images f are made uniform with respect to the transformation items by image processing that adjusts at least one or more transformation items. The estimation engine is also characterized in that it receives an image as input and estimates an enhanced image in which one of an image restored image, a super-resolution image, and a region of interest is enhanced.

[0016] The above features will be described in detail. One of the purposes of this embodiment is to improve image quality by estimating a high-quality image from an input image. Examples of high-quality images include restored images in which blur or noise has been removed, super-resolution images in which image resolution has been improved, and ROI-enhanced images in which a region of interest (ROI) such as a pattern edge or defect is selectively emphasized. In super-resolution processing, the number of pixels of the input image and the output image may differ. In learning the estimation engine, an image of an evaluation object for learning (learning image) is input, and the internal parameters of the estimation engine (network weights, biases, etc.) are updated so as to reduce the difference between an estimated image output from the estimation engine and a high-quality image (correct image), which is a previously taught correct estimated image.

[0017] FIG. 2 is a schematic diagram showing an example of input and output images of the visual inspection system according to this embodiment. The structures and appearances of evaluation targets vary widely in many industrial products, including machinery, metals, chemicals, food, and textiles, and the types of defects that occur therein. FIGS. 2(a1) to 2(f1) are example input images of two gears, one large and one small, photographed as evaluation targets. The defects are, in order, a small foreign object 200, a large foreign object 201, a chip 202, a crack 203, roughness on the gear surface, and a deformed gear contour. Visual inspections are widely performed based on these inspection images to evaluate various aspects of workmanship, such as foreign object adhesion, internal defects and criticality, surface scratches, spots, dirt, shape defects, and assembly defects. Most of these visual inspections are performed visually by inspectors, requiring high-quality images with high defect visibility.

[0018] 2(a2) to (f2) are high-quality images corresponding to FIGS. 2(a1) to (f1), respectively, in which defects such as foreign particles 204 and 205, chips 206, and cracks 207 are emphasized, and the texture of the gear surface and the gear contour (edge) are clearly visible. In the visual inspection system of this embodiment, an image shown in FIGS. 2(a1) to (f1) is input as an input image, and the internal estimation engine is trained to output an image shown in FIGS. 2(a2) to (f2) as an output image.

[0019] The overall processing sequence in this embodiment will be described with reference to Fig. 1. The processing sequence is roughly divided into a learning phase 100 and an operation phase 101.

[0020] In the learning phase 100, the visual inspection system captures an image of an evaluation object 102 for learning purposes to obtain a learning image f (104) (S103). The image is obtained by capturing a digital image of the surface or interior of the evaluation object using an imaging device such as a CCD (Charge Coupled Device) camera, an optical microscope, a charged particle microscope, an ultrasonic inspection device, or an X-ray inspection device. Another example of "acquisition" is simply receiving images captured by another system and storing them in a storage resource of the visual inspection system. The visual inspection system may process captured still images one by one, or may continuously process video images captured at, for example, 30 fps (frames per second) in real time.

[0021] Generally, visual inspection systems use training images f for training to determine the internal parameters of an estimation engine. To obtain an estimation engine with high estimation performance, it is necessary to prepare many training images f corresponding to pattern variations and use them for training. However, collecting a large number of training images can be difficult. To address this issue, data augmentation is known, in which captured images are processed to generate pseudo-images and augment the training images. However, there is a limit to the image variations that can be augmented. Furthermore, if a large number of pattern variations are taught to the estimation engine, the learning load increases, which can lead to overfitting in machine learning.

[0022] This embodiment is characterized by the fact that, rather than comprehensively preparing training images for pattern variations, the image variations present at the time of training are reduced, thereby drastically reducing the number of training images and the training load. In this embodiment, the information contained in an image is considered to be separated into information to be input to a machine learning estimation engine and other information. For example, if the position information of an object to be evaluated on the image is not related to the estimation by the estimation engine, there is no need to include the position information in the training images.

[0023] Such information that does not need to be handled by the estimation engine is subjected to a conversion process (called conversion T) in advance by the visual inspection system (S106) to standardize it among all training images (standardize it to a reference value 108). An image to which conversion T has been applied is called a transformed image. The transformed image of training image f in the training phase 100 is f' (109), and the transformed image of input image h in the operation phase 101, which will be described later, is h' (119).

[0024] Since transformation T can have multiple transformation items, the user selects in advance which transformation items to include in transformation T (S105). Details regarding the types of transformation items will be described later, but for example, if the position of the evaluation object is included as a transformation item, transformation T shifts the evaluation object that exists in various positions in each training image so that it exists in the same position (reference value) in all training images.

[0025] The estimation engine receives as input a transformed image f' (109) in which all evaluation objects are located at the reference value, and estimates a high-quality image in which the evaluation objects are also located at the reference value as an estimated image F' (111). If transformation T is not performed, it would be necessary to prepare and train a group of training images that encompass all differences in the positions of the evaluation objects, but by performing transformation T, it is no longer necessary to include pattern variations related to the positions of the evaluation objects in the training images. The visual inspection system trains the transformed image f' in the estimation engine and determines the internal parameters 117 of the estimation engine (S110).

[0026] In the operation phase 101, the visual inspection system captures an image of an actual evaluation object 102 (S103) and acquires an input image h (118). Furthermore, the visual inspection system applies a transformation T including the transformation items selected in step S105 of the learning phase 100 to the input image h (S106) to obtain a transformed image h' (119). The visual inspection system then inputs the transformed image h' to an estimation engine that uses the internal parameters 117 determined in the learning phase 100, and estimates an estimated image H' (121), which is a high-quality image of the transformed image h' (S120).

[0027] However, there are cases where some or all of the information homogenized by transformation T is necessary in the final output image output by the visual inspection system. Therefore, the user selects the transformation items necessary for the final output from among the transformation items of transformation T (S112), and the visual inspection system applies an inverse transformation U of transformation T for each selected transformation item to the estimated image H' to restore the information before homogenization (S113).

[0028] The image after the inverse transformation is called the inverse transformed estimated image H (122), and this becomes the output image of the visual inspection system. For example, if the estimation engine does not need to handle the position of the evaluation object 102, uniformization is performed using transformation T. However, if the position of the evaluation object 102 is information the user needs for the output image, the uniformed position of the evaluation object in estimated image H' is returned to the state in input image h using transformation U.

[0029] To perform inverse transformation U, when performing transformation T, the visual inspection system saves information about how each transformation item was transformed or what the state was before transformation for each input image h as transformation parameters 123. By performing inverse transformation U based on the transformation parameters 123, the visual inspection system can restore the information removed by transformation T for each image by the inverse transformation U.

[0030] Generally, there is a large amount of variation in images. In this embodiment, by performing a transform T and an inverse transform U before and after processing by the estimation engine, it is possible to partially reduce the load on machine learning without sacrificing the variation in the output image. Because the image variation handled by machine learning itself is reduced, problems related to the cost of collecting training images and the training load can be fundamentally resolved. This allows the estimation performance of the estimation engine to be improved even with a small number of training images.

[0031] <2. Examples of conversion items in conversion T and inverse conversion U> As described above, in this embodiment, the information contained in an image is considered to be divided into information to be input to a machine learning estimation engine and other information, and the latter information is excluded by transformation T. Furthermore, the information excluded by transformation T that is necessary for the final output image is restored by inverse transformation U.

[0032] The transformation T and the inverse transformation U have a plurality of transformation items 107, 114, and what the transformation items 107, 114 in the transformation T and the inverse transformation U should be varies depending on the estimation engine, so any combination is possible. Note that the transformation T and the inverse transformation U may each include a plurality of transformation items. Specific examples of the transformation items 107, 114 will be described below. Specific examples of the transformation items 107, 114 include position, rotation, inversion, magnification, distortion, image brightness, image contrast, noise, shading, and shadow.

[0033] FIG. 3 is a diagram schematically showing an example in which the conversion items are position, rotation, and magnification.

[0034] First, the learning phase 100 will be described. Figures 3(a1) and 3(b1) are schematic diagrams showing an example of two learning images f acquired by capturing images of an evaluation object for learning. The evaluation objects in these two images are two gears, one large and one small, each with the same structure and size, and the same foreign matter attached to them, but there are differences in the positional shift, rotation, and magnification of the evaluation objects on the images. The dotted rectangular frames 300 and 301 surrounding the two gears are auxiliary lines that indicate the differences in the positional shift, rotation, and magnification of the evaluation objects between the images, and are not drawn in the actual captured images.

[0035] If information regarding differences in positional shift, rotation, and magnification is necessary for the estimation engine to estimate a high-quality image, it is not possible to remove this information in advance. On the other hand, if the estimation engine does not need to handle these differences in information, the user selects positional shift, rotation, and magnification as conversion items of transformation T, and the appearance inspection system then applies transformation T to standardize the evaluation object on the image to the same position, orientation, and magnification. Note that the appearance inspection system may automatically select the conversion items of transformation T.

[0036] 3(a2) and 3(b2) are schematic diagrams showing an example of a converted image f' obtained by standardizing the training image f shown in FIGS. 3(a1) and 3(b1) with respect to the above conversion items.

[0037] For example, with regard to rotation, the object to be evaluated is rotated in the converted image f' using the angle at which the centers of the two gears are aligned side-by-side as the reference value 108. Of course, the reference value 108 for each conversion item is not limited to the example shown in the figure and can be set arbitrarily. The reference value may be automatically calculated by the visual inspection system based on the training images. In this case, for example, the average or median of the above-mentioned angles in the multiple training images f may be used as the reference value. Furthermore, the visual inspection system may select one of the multiple training images f and use the above-mentioned angle in the selected training image f as the reference value. In this case, the states of the other training images are adjusted to the state of the selected training image f. Alternatively, the visual inspection system may use a reference value input by the user.

[0038] Furthermore, the visual inspection system may set multiple reference values ​​for one conversion item. In this case, for example, it is possible to adjust the state of each training image to the state that is closest among the multiple reference values. The fewer the number of reference values, the more the variation in the training images is reduced. By having the estimation engine learn using a converted image f' that is standardized for each conversion item, it is possible to reduce the cost of collecting training images and the learning load.

[0039] Next, the operation phase 101 will be described. The two input images h used in the operation phase 101 are assumed to be the same as the images shown in Figures 3(a1) and 3(b1). As in the learning phase 100, the visual inspection system generates a converted image h' by standardizing the object to be evaluated on the image to the same position, orientation, and magnification using transformation T. The converted image h' becomes the images shown in Figures 3(a2) and 3(b2), respectively. Next, the visual inspection system inputs the converted image h' to the estimation engine that has been trained in the learning phase. The estimation engine estimates an estimated image H' (Figures 3(a3) and 3(b3)), which is a high-quality image of the converted image h'. Figures 3(a3) and 3(b3) are schematic diagrams showing an example of the estimated image H'.

[0040] However, estimated image H' loses information about the positional shift, rotation, and magnification differences of the object to be evaluated that were present in input image h. Therefore, the visual inspection system applies inverse transformation U to estimated image H' to generate an inversely transformed estimated image H (FIGS. 3(a4) and (b4)) in which the information removed by transformation T is restored, and this is used as the output image. FIGS. 3(a4) and (b4) are schematic diagrams showing an example of the inversely transformed estimated image H. In inversely transformed estimated image H, the rectangular dotted-line frames 309 and 310 surrounding the two gears are the same as the corresponding dotted-line frames 300 and 301 in the input image, and information about the position, orientation, and magnification, which differ from image to image, has been restored.

[0041] By using magnification as a conversion item, even if the actual size of the evaluation object varies, it is possible to standardize the apparent size of the evaluation object on the image by using the conversion T. Furthermore, although not shown, it is also possible to standardize the appearance by using not only image rotation but also image inversion as a conversion item.

[0042] FIG. 4 is a diagram schematically illustrating an example in which the transformation items are rotation and distortion.

[0043] In operation phase 101, the visual inspection system applies transformation T to the input image h (Fig. 4(a)), converting the orientation and distortion of the object to be evaluated in the image to reference values ​​to generate a transformed image h' (Fig. 4(b)), which is then input to the estimation engine. The estimation engine then estimates an estimated image H' (Fig. 4(c)), which is a high-quality image of the transformed image h'. At this stage, the estimation engine only improves the image quality, so the rectangular dotted-line frames 401 and 402 that surround the two gears, drawn as auxiliary lines, are the same for both the transformed image h' and the estimated image H'.

[0044] As in the example of Figure 3, the visual inspection system can generate an inverse-transformed estimated image H (Figure 4(d)) by applying an inverse transformation U of the transformation T to estimated image H', and use this as the output image. In the inverse-transformed estimated image H shown in Figure 4(d), information about the rotation and distortion of the object to be evaluated that was present in the input image h has been restored, and the rectangular dotted-line frames 400 and 403 that surround the two gears, drawn as auxiliary lines, are the same in both images.

[0045] As a variation of the processing method, the inverse transform U does not need to be a complete inverse transform of the transform T. In other words, the transform items of the transform T and the inverse transform U do not need to match. In the example of FIG. 4(b), the transform items of the transform T are rotation and distortion. However, if you want to include rotation information but not distortion information in the final output image, you can restore only the rotation by using rotation as the transform item of the inverse transform U, and obtain an output image with only the rotation restored and distortion removed (FIG. 4(e)). Because distortion is removed, the dotted frame 404, which is an auxiliary line, becomes the same as the dotted frame 402 in the estimated image H' (FIG. 4(c)) after rotation.

[0046] FIG. 5A is a diagram schematically showing an example in which the conversion items are brightness, contrast, and noise, and FIG. 5B is a diagram schematically showing an example in which the conversion items are shading and shadow.

[0047] The outline of the process other than the conversion items and the meaning of the diagram are the same as in Figures 3 and 4.

[0048] 5A(a1)-(a4) show examples in which the conversion items of the conversion T and the inverse conversion U are brightness and contrast. In this case, the visual inspection system uses the conversion T to equalize the apparent brightness and contrast of the training image or input image (FIG. 5A(a1)) of the object to be evaluated, which vary depending on the imaging conditions, to a reference value, to obtain a converted image h' (FIG. 5A(a2)). For example, processing such as matching the brightness histogram of each image to a reference brightness histogram is conceivable. Alternatively, processing such as converting a color image to a grayscale image is conceivable. The visual inspection system applies the inverse conversion U to an estimated image H' (5A(a3)), which is estimated by an estimation engine using the converted image h' as ​​input, and generates an inversely converted estimated image H (5A(a4)) as the output image.

[0049] Figures 5A(b1) to 5A(b5) show examples in which noise is used as the transformation item for each of the transformations T and the inverse transformation U. In this case, the visual inspection system applies transformation T to the training image or input image (Figure 5A(b1)) to equalize the noise in these images and obtain a transformed image h' (5A(b2)). The noise may be removed as shown in Figure 5A(b2), or the noise level of each image may be adjusted to a reference noise level. If noise is unnecessary in the output image, the noise-removed estimated image H' shown in Figure 5A(b3) may be used as the output image. While noise is generally considered unnecessary in high-quality images, there are cases in which an image appears unnatural if no noise is superimposed on it, or in which differences in the amount of noise superimposed on each image are desired. In such cases, the visual inspection system can use the inverse-transformed estimated image H shown in Figure 5A(b4), in which noise has been restored by inverse transformation U, as the output image.

[0050] Furthermore, as a variation of the processing method, the inverse transformation U does not need to be a complete inverse transformation of the transformation T. In other words, it is possible to change the degree (strength) to which the inverse transformation U is applied to the transformation T. For example, it would be unnatural if noise were completely eliminated from the output image as in FIG. 5A(b3), but if noise equivalent to that of the input image is superimposed on the output image as in FIG. 5A(b4), making it difficult to observe, the visual inspection system can apply a weight to the transformation U to restore a small amount of noise as in FIG. 5A(b5). This is not limited to noise, but the strength can be adjusted for all transformation items.

[0051] Figures 5B(c1)-(c5) show examples in which the transformation item for each of the transformation T and inverse transformation U is shading. In this case, the visual inspection system applies transformation T to the training image or input image (Figure 5B(c1)) to uniformize the shading of these images and obtain a transformed image h' (Figure 5B(c2)). As with the noise cases in Figures 5A(b1)-(b5), the visual inspection system can generate output images from either an image with completely removed shading (Figure 5B(c3)), an image with completely restored shading (Figure 5B(c4)), or an image with half the shading restored (Figure 5B(c5)). Strong shading in the output image can be difficult to observe, but completely removing it can also create unnatural results. For example, in real-time display of high-resolution images of moving images, it would be unnatural if the shading remained completely unchanged when the positional relationship between the object being evaluated and the lighting was shifted. Therefore, the strength of the restoration by the transformation U can be adjusted according to the user's wishes.

[0052] Figures 5B(d1) to (d4) show examples in which the transformation items of transformation T and inverse transformation U are shadows. In this case, the visual inspection system applies transformation T to the training image or input image (Figure 5B(d1)) to remove or equalize the shadows in these images and obtain a transformed image h' (Figure 5B(d2)). If shadows are not necessary in the output image, the visual inspection system may use the estimated image H' from which the shadows have been removed, as shown in Figure 5B(d3), as the output image. Alternatively, the visual inspection system may use the inverse-transformed estimated image H, as shown in Figure 5B(d4), in which noise has been restored by inverse transformation U, as the output image.

[0053] Although specific examples of conversion items have been shown above, the conversion items are not limited to these and can be set for any pattern variation contained in an image. For example, multiple conversion items may be selected from multiple conversion items 107 and 114. The conversion items in conversion T and inverse conversion U vary depending on what pattern variations are handled in the estimation engine and what kind of image the user desires for the main image, so any combination can be set (S105, S112). The conversion items included in inverse conversion U are some or all of the selected items included in conversion T.

[0054] <3. Learning the estimation engine> 6 and 7 are schematic diagrams showing an example of a learning method for the estimation engine. The learning method can be broadly classified into two types (1) and (2) below, with Figs. 6(a1) to (a5) and Figs. 7(a1) to (a5) corresponding to (1), and Figs. 6(b1) to (b5) and Figs. 7(b1) to (b5) corresponding to (2).

[0055] (1) In the learning phase 100, the internal parameters of the image estimation engine are determined based on an estimated image F' estimated by the estimation engine using the transformed image f' as input, and a transformed correct image G' obtained by applying transformation T to the correct image G.

[0056] (2) In the learning phase 100, the internal parameters of the estimation engine are determined based on an inverse-transformed estimated image F, which is obtained by applying an inverse transformation U to an estimated image F' estimated by the estimation engine using the transformed image f' as input, and a ground truth image G for the training image f.

[0057] The learning of the estimation engine is basically performed by updating the internal parameters of the estimation engine, such as the weights and biases of the network, so as to reduce the difference between an estimated image output from the estimation engine using a training image as input and a previously taught correct image. Note that the difference between the estimated image and the correct image may be, for example, the difference in pixel data for each pixel of each image.

[0058] However, in this embodiment, the estimation engine outputs a transformed estimated image, not an estimated image, which cannot be directly compared with the target image. Therefore, in (1), a transformed target image obtained by applying a transformation process to the target image is compared with the transformed estimated image, and in (2), an inverse transformed estimated image obtained by applying an inverse transformation U to the estimated image is compared with the target image.

[0059] As a specific example of (1), in the learning method of FIGS. 6(a1) to (a5), the visual inspection system applies transformation T to training image f (FIG. 6(a1)) to obtain transformed image f' (FIG. 6(a2)). Next, the visual inspection system uses this transformed image f' as input and estimates estimated image F' (FIG. 6(a3)) using an estimation engine. The visual inspection system also applies transformation T to ground truth image G (FIG. 6(a4)) to obtain transformed ground truth image G' (FIG. 6(a5)). The visual inspection system calculates a loss value based on the difference between estimated image F' (FIG. 6(a3)) and transformed ground truth image G' (FIG. 6(a5)), and determines internal parameters of the estimation engine so as to reduce the loss value.

[0060] As a specific example of (2), in the learning method of FIGS. 6(b1) to (b5), the visual inspection system applies transformation T to training image f (FIG. 6(b1)) to obtain transformed image f' (FIG. 6(b2)). The visual inspection system uses this transformed image f' as input and estimates estimated image F' (FIG. 6(b3) or reference numeral 111 in FIG. 1) using an estimation engine. The visual inspection system applies inverse transformation U (step S113 in FIG. 1) to estimated image F' to obtain inversely transformed estimated image F (FIG. 6(b4) or reference numeral 115 in FIG. 1). The visual inspection system calculates a loss value based on the difference between the inversely transformed estimated image F (FIG. 6(b4)) and the ground truth image G (FIG. 6(b5)), and determines the internal parameters of the estimation engine so as to reduce the loss value.

[0061] In the above-described embodiment, learning was performed by minimizing the loss value based on the difference between images. However, learning may also be performed using a Generative Adversarial Network (GAN). A GAN consists of two competing networks called a Generator and a Discriminator. By using a GAN, learning is possible even when the estimated image and the ground truth image are not aligned. Furthermore, as shown in FIG. 7, even if the evaluation targets of the estimated image and the ground truth image are different, the estimated image can be trained to have the same image quality as the ground truth image.

[0062] That is, as a specific example of (1) using GAN, in the learning method of Figures 7(a1) to (a5), the visual inspection system applies transformation T to training image f (Figure 7(a1)) to obtain transformed image f' (Figure 7(a2)). The visual inspection system uses this transformed image f' as input and estimates estimated image F' (Figure 7(a3) or reference numeral 111 in Figure 1) using an estimation engine (Generator). The visual inspection system also applies transformation T to target image G (Figure 7(a4)) to obtain transformed target image G' (Figure 7(a5)). The visual inspection system uses a Discriminator to determine whether estimated image F' (Figure 7(a3)) is the same type as transformed target image G' (Figure 7(a5)), and if it is determined that they are not the same type, it increases the loss value. The generator is trained to make the estimated image F' the same type of image as the converted ground truth image G', and the discriminator is trained to detect that the estimated image F' is not the same type as the converted ground truth image G'. By repeating this adversarial training, the estimation accuracy of the estimation engine is improved.

[0063] As a specific example of (2) using GAN, in the learning method of Figures 7(b1) to (b5), a visual inspection system applies transformation T to training image f (Figure 7(b1)) to obtain transformed image f' (Figure 7(b2)). The visual inspection system uses this transformed image f' as input and estimates estimated image F' (Figure 7(b3) or reference numeral 111 in Figure 1) using an estimation engine (Generator). The visual inspection system applies inverse transformation U (step S113 in Figure 1) to estimated image F' to obtain inversely transformed estimated image F (Figure 7(b4) or reference numeral 115 in Figure 1). The visual inspection system uses a Discriminator to determine whether the inversely transformed estimated image F (Figure 7(b4)) is the same type as the ground truth image G (Figure 7(b5)), and if it is determined that they are not the same type, it increases the loss value. The generator is trained to make the inverse-transformed estimated image F the same type of image as the ground truth image G, and the discriminator is trained to detect that the inverse-transformed estimated image F is not the same type of image as the ground truth image G. By repeating this adversarial learning, the estimation accuracy of the estimation engine is improved.

[0064] <4.GUI (Graphical User Interface)> The visual inspection system according to this embodiment is characterized by having a GUI that accepts designation from the user of a learning method for the estimation engine, including conversion items for the conversion processing and conversion items for the inverse conversion processing.

[0065] 8A and 8B are schematic diagrams showing examples of GUIs. During training of the estimation engine, the user can select, using radio buttons 800, whether to compare the estimated image F' with the converted correct image G' as shown in FIGS. 6(a1)-(a5) or 7(a1)-(a5), or to compare the inverse-transformed estimated image F with the correct image G as shown in FIGS. 6(b1)-(b5) or 7(b1)-(b5). The user can select, using radio buttons 801, whether to use the "difference between comparison images" (FIG. 6) or "GAN" (FIG. 7) as the image comparison method. By specifying the ID of the training data using a pull-down menu 802, the user can display some or all of the training image f, converted image f', estimated image F', inverse-transformed estimated image F, correct image G, and converted correct image G' associated with that ID in the image display area 803. In this way, by displaying not only the training image f but also the transformed image f', the estimated image F', the inverse transformed estimated image F, the correct image G, and the transformed correct image G', the user can know what processing has been performed.

[0066] The illustration shows a group of images in the learning phase 100, but some or all of the input image h, transformed image h', estimated image H', and inverse transformed estimated image (output image) H in the operation phase 101 can also be displayed in a similar manner. The transformation items and transformation methods for transformation T and inverse transformation U can be specified in transformation specification area 804. The transformation items for transformation T can be specified using check boxes 805, and the method for setting the reference values ​​for the specified transformation items can be specified using radio buttons 806.

[0067] The items that can be specified using radio button 806 include "Automatically calculate reference value," "Match to image with ID XX," and "Specify reference value." If the reference value is a position and "Automatically calculate reference value" is checked, the visual inspection system will calculate, for example, the average position of the center positions of multiple input images f as the reference value. Also, if "Match to image with ID XX" is checked, the visual inspection system will calculate, as the reference value, the center position of the input image f associated with the specified ID "XX." If "Specify reference value" is checked, the visual inspection system will display a screen prompting the user to input a reference value, and will adopt the value entered by the user on that screen as the reference value.

[0068] Furthermore, the user can specify the conversion items of the inverse conversion U using check boxes 807, and can specify the strength for each specified conversion item in an strength specification area 808. The visual inspection system accepts the specified strengths and executes the inverse conversion U at those strengths.

[0069] Second Embodiment <5. Estimating evaluation values ​​using visual inspection systems> In the first embodiment, the objective of the visual inspection system was image quality improvement by estimating a high-quality image from an input image, but the present invention is not limited to this. In the second embodiment, the objective of the visual inspection system is image evaluation by estimating an evaluation value from an input image. Examples of evaluation values ​​include the quality level of the object to be evaluated, the presence or absence of defects, the degree of abnormality, and the criticality of the object, the evaluation value in the case of visual inspection is an area label, and the evaluation value in the case of defect detection is a defective area.

[0070] The visual inspection system according to this embodiment is an example of an evaluation value estimation device and executes an evaluation value estimation method. The evaluation value estimation method is a computer-executable method including: a first learning image acquisition step of acquiring multiple learning images f of a learning object; a first converted image generation step of applying a conversion process to the multiple learning images f to generate multiple converted images f'; a learning step of determining internal parameters of a machine learning estimation engine using the multiple converted images f'; a second image acquisition step of acquiring an input image h of an actual object; a second converted image generation step of applying the conversion process to the input image h to generate a converted image h'; an estimated evaluation value calculation step of inputting the converted image h' to the trained estimation engine and estimating an estimated evaluation value R'; a second estimated evaluation value calculation step of estimating a second estimated evaluation value R based on the values ​​of the input image h and the estimated evaluation value R' for some or all of the conversion processes; and an output step of outputting the second estimated evaluation value R, wherein the conversion process is characterized by adjusting at least one conversion parameter to uniformize the patterns of the multiple learning images f with respect to the conversion parameters through image processing.

[0071] These features will be described in detail. FIG. 9 is a schematic diagram showing an example of the overall processing sequence performed by the visual inspection system according to this embodiment. The processing sequence is broadly divided into a learning phase 900 and an operation phase 901. In the learning phase 900, the visual inspection system captures an evaluation object 902 for learning purposes to obtain a learning image f (904) (S903). As with the first embodiment shown in FIG. 1, this embodiment distinguishes between information contained in an image that should be input to a machine learning estimation engine and other information. Information that does not need to be handled by the estimation engine is subjected to a conversion process (referred to as conversion T) in advance by the visual inspection system (S906) to standardize it among all learning images (standardize it to a reference value 908). An image to which conversion T has been applied is called a converted image. The converted image of learning image f in the learning phase is f' (909), and the converted image of input image h in the operation phase, which will be described later, is h' (919).

[0072] Because transformation T can have multiple transformation items 907, the user selects in advance which transformation items 907 to include in transformation T (S905). For example, if the position of the evaluation object is to be included in the transformation items 907, the evaluation object, which exists at various positions in each training image, is shifted so that it is at the same position (reference value) in all training images. The estimation engine uses a transformed image f' (909) in which all evaluation objects are located at the reference value as input and estimates an estimated evaluation value P' (911). If transformation T is not performed, it would be necessary to prepare and train a group of training images that encompass the differences in the positions of the evaluation object. However, by performing transformation T, it is no longer necessary to include pattern variations regarding the position of the evaluation object in the training images. The visual inspection system trains a machine learning estimation engine with the transformed image f' and determines the internal parameters 917 of the estimation engine (S910).

[0073] In the operation phase 901, the visual inspection system captures an image of an actual evaluation object 902 (S903) and acquires an input image h (918). Furthermore, the visual inspection system applies a transformation T including the transformation items selected in step S905 of the learning phase 900 to the input image h (S906) to obtain a transformed image h' (919). The visual inspection system then inputs the transformed image h' to a machine learning estimation engine that uses the internal parameters 917 determined in the learning phase 900, and estimates an estimated evaluation value R' (921) of the transformed image h' (S920).

[0074] However, some or all of the information standardized by transformation T may be necessary for estimating the final output value of the visual inspection system. Therefore, the visual inspection system stores, for each input image h, information on how each transformation item 907 was performed or the state before transformation as transformation parameters 923. The user selects transformation items 914 required for the final output from among the transformation items of transformation T (S912), and a second estimated evaluation value R (922) is estimated by a second estimation engine from the transformation parameters 923 and estimated evaluation value R' (921) for each selected transformation item (S913). This second estimated evaluation value R becomes the output value of the image processing system.

[0075] For example, consider a case where the conversion item 907 of the conversion T is the position of the evaluation object. If it is not necessary to handle the position of the evaluation object in step S920, where the evaluation value is estimated by the machine learning estimation engine, the visual inspection system performs uniformization using the conversion T. If the estimated evaluation value is the degree of abnormality of the evaluation object and the criterion for determining the degree of abnormality handled by the estimation engine is a change in contrast in the input image of the evaluation object, then it is not necessary for the image input to the estimation engine to include positional variations.

[0076] On the other hand, the final anomaly determination may take into account the position information excluded by the conversion T. That is, if a large positional deviation of the evaluation object in addition to a change in contrast is also deemed to be an anomaly, the appearance inspection system estimates a second estimated evaluation value R (922) using a second estimation engine (S913) based on the estimated evaluation value R' (911) related to contrast handled in step S920 for estimating the evaluation value using the estimation engine and the positional information included in the conversion parameters 923, and sets this as the final output value. The second estimated evaluation value R can reflect both anomalies, such as a change in contrast and a positional deviation.

[0077] On the other hand, by temporarily excluding information related to some of the conversion items 907 using the conversion T, it is possible to reduce the number of training images and the training load in machine learning. This makes it possible to improve the estimation performance of the estimation engine even with a small number of training images.

[0078] 6. Learning the estimation engine Fig. 10 is a schematic diagram showing an example of a learning method for an estimation engine. The learning method can be broadly classified into two types (1) and (2) below. Fig. 10(a1) to (a5) correspond to (1), and Fig. 10(b1) to (b5) correspond to (2).

[0079] (1) In the learning phase 900, the transformed image f' is used as input to estimate an estimated evaluation value P' in an estimation engine. A correct answer S' of the estimated evaluation value is estimated from the second correct answer evaluation value S and the transformation parameters. The internal parameters of the image estimation engine are determined based on the estimated evaluation value P' and the correct answer S' of the estimated evaluation value.

[0080] (2) In the learning phase 900, the estimation engine estimates an evaluation value P' using the converted image f' as input. The second estimation engine estimates a second estimated evaluation value P (915) using the evaluation value P' and the conversion parameters 923 as input. The internal parameters of the estimation engine are determined based on the second estimated evaluation value P (915) and the correct answer value S of the second estimated evaluation value.

[0081] The estimation engine is basically trained by updating the internal parameters of the estimation engine, such as the network weights and biases, so as to reduce the difference between the estimated evaluation value output from the estimation engine using a training image as input and the correct evaluation value (the correct value of the second estimated evaluation value) that is previously taught.

[0082] However, in this embodiment, the estimation engine outputs an evaluation value for the converted image f', not an evaluation value for the training image f instructed by the user, and therefore cannot directly compare it with the correct value of the second estimated evaluation value provided by the user. Therefore, in (1), the correct value S' of the estimated evaluation value estimated from the correct value S of the second estimated evaluation value is compared with the estimated evaluation value P', and in (2), the second estimated evaluation value P estimated from the estimated evaluation value P' is compared with the correct value S of the second estimated evaluation value.

[0083] As a specific example of (1), in the learning method of FIGS. 10(a1) to (a5), the visual inspection system applies transformation T to training image f (FIG. 10(a1)) to obtain transformed image f' (FIG. 10(a2)). Here, as an example, the transformation item of transformation T is the texture of the gear surface, which is homogenized among the training images. In FIG. 10(a2), the texture is removed and converted to a plain pattern. The visual inspection system inputs this transformed image f' and uses an estimation engine to estimate an estimated evaluation value P' (FIG. 10(a3) or reference numeral 911 in FIG. 9). The evaluation value may be, for example, the degree of abnormality of the evaluation object (gear). Because the texture information has been homogenized, in step S910 the estimation engine estimates the degree of abnormality using a judgment criterion other than texture. For example, because a foreign substance 1000 is attached to transformed image f' (FIG. 10(a2)), the degree of abnormality is estimated to be 45%.

[0084] On the other hand, the user assigns a correct value (correct value S of the second estimated evaluation value) to the visual inspection system for the training image f before conversion (Fig. 10(a1)). If the user's criteria for determining the degree of abnormality include not only the presence or absence of foreign matter but also texture (roughness of the gear surface), the abnormality level will be high at 80% (Fig. 10(a4)), because training image f has foreign matter attached and the gear surface is rough. The conversion parameter (texture information in this case) is added to this correct value S of the second estimated evaluation value (80% abnormality level) to estimate the abnormality level (correct value S' of the estimated evaluation value) by excluding the abnormality level due to texture alone from the above abnormality level. If the increase in the abnormality level due to roughness of the gear surface is subtracted and the abnormality level due to foreign matter alone is estimated, it will be, for example, 50% (Fig. 10(a5)). The visual inspection device calculates a loss value based on the difference between the estimated evaluation value P' (Fig. 10(a3)) and the correct value S' (Fig. 10(a5)) of the estimated evaluation value, and determines the internal parameters of the estimation engine so that the loss value becomes small.

[0085] As a specific example of (2), in the learning method of Figures 10(b1) to (b5), the visual inspection system applies transformation T to training image f (Figure 10(b1)) to obtain transformed image f' (Figure 10(b2)). The visual inspection system uses this transformed image f' as input and estimates an estimated evaluation value P' (Figure 10(b3) or reference numeral 911 in Figure 9) using an estimation engine. As in the case of (1) above, if the transformation item of transformation T is the texture of the gear surface and the evaluation value is the degree of abnormality of the evaluation object (gear), the estimated evaluation value P' does not take into account the degree of abnormality based on the texture.

[0086] Therefore, in step S913, the visual inspection system estimates a second estimated evaluation value P (FIG. 10(b4)) using a second estimation engine based on the estimated evaluation value P' and the transformation parameters 923 (texture information in this case). This enables the user to compare this with the correct value (correct value S of the second estimated evaluation value) in FIG. 10(b5) assigned to the training image f (FIG. 10(b1)) before transformation. That is, the visual inspection system calculates a loss value based on the difference between the second estimated evaluation value P (FIG. 10(b4)) and the correct value S of the second estimated evaluation value (FIG. 10(b5)), and determines the internal parameters of the estimation engine so as to reduce the loss value.

[0087] The second embodiment differs from the first embodiment in that what is estimated is an evaluation value instead of an image, but the conversion items, learning method, GUI, etc. are the same as those of the first embodiment.

[0088] Furthermore, various machine learning engines can be used as the estimation engine in each of the first and second embodiments, including, for example, deep neural networks such as convolutional neural networks (CNNs), support vector machines (SVMs), support vector regressions (SVRs), and k-nearest neighbors (k-NNs). These estimation engines can handle image estimation, region segmentation, classification problems, and regression problems.

[0089] <7. Hardware configuration of image processing system> FIG. 11 is a diagram showing the hardware configuration of the appearance inspection system according to each of the first and second embodiments.

[0090] The visual inspection system 1 includes the above-mentioned image capturing device 1106 and computer 1100. An example of the image capturing device 1106 has already been described.

[0091] The computer 1100 is hardware that executes the image estimation method of the first embodiment and the evaluation value estimation method of the second embodiment, and includes the following:

[0092] *Processor 1101: Examples of the processor 1101 include a CPU (Central Processing Unit), a GPU (Graphical Processing Unit), and an FPGA (Field Programmable Gate Array), but other hardware may be used as long as it can process the image processing method.

[0093] *Storage resource 1102: Examples of the storage resource 1102 include non-volatile memory such as RAM (Random Access Memory), ROM (Read Only Memory), HDD (Hard Disk Drive), and flash memory. The storage resource may store a program that causes the processor 1101 to execute the image estimation method and the evaluation value estimation method described in each of the above embodiments.

[0094] * GUI device 1103: Examples of the GUI device 1103 include a display and a projector, but other hardware may be used as long as it can display a GUI.

[0095] *Input device 1104: Examples of the input device 1104 include a keyboard, a mouse, and a touch panel, but other devices may be used as long as they can accept operations from the user. Also, the input device 1104 and the GUI device 1103 may be integrated into one piece of hardware.

[0096] *Communication interface device 1105: Examples of the communication interface device 1105 include USB (Universal Serial Bus), Ethernet, and Wi-Fi. Any other interface device may be used as long as it can directly receive images from the imaging device 1106 or allows the user to transmit the images to the computer 1100. Furthermore, a portable non-volatile storage medium (not shown) on which the images are stored may be connected to the communication interface device 1105, and the images may be stored in the computer 1100. Examples of such portable non-volatile storage media include flash memory, DVD (Digital Versatile Disc), CD-ROM (Compact Disc), and Blu-ray disc.

[0097] The above is the hardware configuration of the computer 1100. Note that the visual inspection system may include a plurality of computers 1100 and a plurality of image capturing devices 1106.

[0098] The above-mentioned program may be stored in the computer 1100 via the following path: *The program is stored in a portable nonvolatile storage medium, and the medium is connected to the communication interface device 1105 to distribute the program to the computer 1100.

[0099] *The program is distributed to the computer 1100 by a program distribution server. The program distribution server has a storage resource that stores the program, a processor that performs distribution processing to distribute the program, and a communication interface device that can communicate with the communication interface device 1105 of the computer 1100.

[0100] As mentioned above, the embodiments described above do not limit the scope of the invention as claimed, and not all of the elements and combinations thereof described in the embodiments are necessarily essential to the solution of the invention.

[0101] Although each embodiment deals with a single input image as input information, each embodiment can also be applied to cases where there are multiple input images and multiple types of output images or estimated evaluation values. In this case, the estimation engine becomes another input or another output.

[0102] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0103] The present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to an embodiment including all of the described components. Furthermore, part of the configuration of one embodiment can be replaced with the configuration of another embodiment, and the configuration of another embodiment can be added to the configuration of one embodiment. Furthermore, part of the configuration of each embodiment can be added to, deleted from, or replaced with another configuration.

[0104] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. Furthermore, the above-described configurations, functions, etc. may be implemented in software by a processor interpreting and executing a program that implements each function. Information such as the program, decision table, and files that implement each function can be stored in memory, a storage device such as an HDD or SSD, or a recording medium such as an IC (Integrated Circuit) card, an SD (Secure Digital) card, or a DVD. Furthermore, the control lines and information lines shown are those considered necessary for explanation, and do not necessarily represent all control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]

[0105] 100, 900...Learning phase, 101, 901...Operation phase, 102, 902...Evaluation object, 104, 904...Learning image, 107, 114, 907, 914...Conversion item, 108, 908...Reference value, 117, 917...Internal parameters, 109, 909...Transformed image, 111...Estimated image, 115...Inverse transformed estimated image, 118, 918...Input image, 119, 919...Transformed image, 121...Estimated image, 122...Inverse transformed estimated image, 123, 923...Conversion parameters, 200...Foreign object, 201...Foreign object, 203...Crack, 204, 205, 1000... Foreign matter, 207...crack, 300, 301, 309, 310, 400-404...dotted frame, 800, 801...radio button, 802...pull-down menu, 803...image display area, 804...conversion specification area, 805, 807...check box, 806...radio button, 808...intensity specification area, 911, 921...estimated evaluation value, 915, 922...second estimated evaluation value, 1100...computer, 1101...processor, 1102...storage resource, 1103...GUI device, 1104...input device, 1105...communication interface device, 1106...imaging device.

Claims

1. 1. A computer-implemented image estimation method, comprising: a first learning image acquisition step of acquiring a plurality of learning images f obtained by capturing an object for learning; a receiving step of receiving from a user a specification of an item that is not handled by the estimation engine among the conversion items of the conversion processing and the conversion items of the inverse conversion processing; a first converted image generating step of generating a plurality of converted images f′ by performing the conversion process of uniforming the plurality of learning images f with respect to the items designated by the user to a reference value that can be arbitrarily set; a learning step of determining internal parameters of a machine learning estimation engine using the plurality of transformed images f′; a second image acquisition step of acquiring an input image h of an actual object; a second transformed image generating step of generating a transformed image h' by performing the transformation process on the input image h; an estimated image generation step of inputting the transformed image h' to the trained estimation engine to estimate an estimated image H'; an inverse transformed estimated image generating step of generating an inverse transformed estimated image H by performing the inverse transform process of a part or all of the transform process on the estimated image H'; and an output step of outputting the inverse transformed estimated image H as an output image. Image estimation methods.

2. 2. The image estimation method according to claim 1, The estimation engine receives an image as an input and estimates one of an image restored image, a super-resolution image, and an enhanced image in which a region of interest is enhanced. Image estimation methods.

3. 2. The image estimation method according to claim 1, The conversion items include at least one of image brightness, image contrast, image noise, image shading, and shade. Image estimation methods.

4. 2. The image estimation method according to claim 1, the computer determines a reference value for the conversion item from the plurality of learning images f; The equalization process is performed by converting the values ​​of the conversion items in each of the learning images f to the reference values. Image estimation methods.

5. 2. The image estimation method according to claim 1, the computer accepts a designation of a degree of the inverse conversion process for each of the conversion items, and performs the inverse conversion process for each of the conversion items at the designated degree; Image estimation methods.

6. 2. The image estimation method according to claim 1, In the learning step, the computer determines internal parameters of the estimation engine based on an estimated image F' estimated by the estimation engine using the transformed image f' as an input and a transformed correct image G' obtained by performing the transformation process on a correct image G. Image estimation methods.

7. 2. The image estimation method according to claim 1, In the learning step, the computer determines internal parameters of the estimation engine based on an inverse-transformed estimated image F obtained by performing an inverse transformation process on an estimated image F' estimated by the estimation engine using the transformed image f' as an input, and a correct image G for the training image f. Image estimation methods.

8. 2. The image estimation method according to claim 1, a display step of displaying a GUI (Graphical User Interface) that accepts designation of the conversion items of the conversion processing and the conversion items of the inverse conversion processing from the user; Image estimation methods.

9. A computer-implemented rating value estimation method, comprising: a first learning image acquisition step of acquiring a plurality of learning images f obtained by capturing an object for learning; a first converted image generating step of generating a plurality of converted images f′ by performing a conversion process on the plurality of training images f; a learning step of determining internal parameters of a machine learning estimation engine using the plurality of transformed images f′; a second image acquisition step of acquiring an input image h of an actual object; a second transformed image generating step of generating a transformed image h' by performing the transformation process on the input image h; an estimated evaluation value calculation step of inputting the converted image h′ to the trained estimation engine to estimate an estimated evaluation value R′; a second estimated evaluation value calculation step of estimating a second estimated evaluation value R based on the value of the input image h and the estimated evaluation value R′ for a part or all of the conversion processing; an output step of outputting the second estimated evaluation value R as an output value; the conversion processing is image processing for adjusting at least one or more conversion items to uniformize the patterns of the plurality of learning images f with respect to the conversion items. Evaluation value estimation method.

10. An image estimation apparatus comprising a processor, The processor: a first learning image acquisition step of acquiring a plurality of learning images f obtained by capturing an object for learning; a receiving step of receiving from a user a specification of an item that is not handled by the estimation engine among the conversion items of the conversion processing and the conversion items of the inverse conversion processing; a first converted image generating step of generating a plurality of converted images f′ by performing the conversion process of uniforming the plurality of learning images f with respect to the items designated by the user to a reference value that can be arbitrarily set; a learning step of determining internal parameters of a machine learning estimation engine using the plurality of transformed images f′; a second image acquisition step of acquiring an input image h of an actual object; a second transformed image generating step of generating a transformed image h' by performing the transformation process on the input image h; an estimated image generation step of inputting the transformed image h' to the trained estimation engine to estimate an estimated image H'; an inverse transformed estimated image generating step of generating an inverse transformed estimated image H by performing the inverse transform process of a part or all of the transform process on the estimated image H'; and outputting the inverse transformed estimated image H as an output image. Image estimation device.

11. The image estimation device according to claim 10, The estimation engine receives an image as an input and estimates a restored image, a super-resolution image, and an enhanced image in which a region of interest is enhanced. Image estimation device.

12. The image estimation device according to claim 10, The transformation items include at least one of image brightness, image contrast, image noise, image shading, and shade. Image estimation device.

13. The image estimation device according to claim 10, the processor determines a reference value for the conversion item from the plurality of training images f; The equalization process is performed by converting the values ​​of the conversion items in each of the learning images f to the reference values. Image estimation device.

14. The image estimation device according to claim 10, the processor accepts a designation of a degree of the inverse conversion process for each of the conversion items, and performs the inverse conversion process for each of the conversion items at the designated degree; Image estimation device.

15. The image estimation device according to claim 10, the processor performs control to display a GUI that accepts designation of the conversion items of the conversion processing and the conversion items of the inverse conversion processing from the user. Image estimation device.

Citation Information

Patent Citations

  • Image processor

    JP2004062719A

  • Super-resolution device and program

    JP2016115313A

  • Information prediction system, information prediction method and information prediction program

    JP2016177829A

  • Visual inspection device, method for improving accuracy of determination for existence / nonexistence of shape failure of welding portion and kind thereof using same, welding system, and work welding method using same

    WO2020129617A1