Image generation method for machine learning, machine learning method, and image processing device for endoscope

JPWO2024201984A5Active Publication Date: 2025-12-11OLYMPUS MEDICAL SYST CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025509574
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2025-12-11
Estimated Expiration
2043-03-31

AI Technical Summary

Technical Problem

Machine learning models used for super-resolution image generation can produce incorrect inferences, particularly in medical images, due to the inclusion of high-frequency components not present in the training images, leading to false patterns such as incorrectly generated blood vessels.

Method used

A method that reduces high-frequency components above the Nyquist frequency of the training image by generating a correct image candidate with reduced frequency components, using an optical low-pass filter or image processing to create a correct image for training, thereby preventing erroneous inferences.

Benefits of technology

This approach reduces the likelihood of generating incorrect output images by ensuring that the machine learning model only learns from existing frequency information, improving the accuracy and reliability of image inference, especially in medical imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2024201984000001
    Figure 2024201984000001
  • Figure 2024201984000002
    Figure 2024201984000002
Patent Text Reader

Abstract

This image generation method for machine learning generates correct answer images (IC) by performing reduction processing on a component of at least a partial frequency band within a frequency band that is higher than the Nyquist frequency (ftn) of training images (IT) with respect to correct answer image candidates (ICC) having a higher resolution than the training image (IT). The correct answer images (IC) and the training images (IT) are used in pairs in machine learning for increasing the resolution of an input image.
Need to check novelty before this filing date? Find Prior Art

Description

Machine learning image generation method, machine learning method, and endoscope image processing device

[0001] The present invention relates to a machine learning image generation method for generating a correct image from correct image candidates with a higher resolution than a training image, a machine learning method for training a machine learning model using training images and correct images, and an endoscopic image processing device that uses the trained machine learning model.

[0002] Technologies for improving the spatial resolution of an image include super-resolution technology and edge enhancement technology. Super-resolution technology increases the number of pixels that make up an image, thereby increasing its resolution compared to the original image. Edge enhancement technology enhances edges by amplifying specific spatial frequency components in the image.

[0003] Techniques for super-resolution using machine learning models have been proposed. The machine learning model performs inference on an input image, creates a post-inference image with higher resolution, and uses this as the output image.

[0004] For example, a machine learning model performs super-resolution machine learning using training data including pairs of training images and correct images with a higher resolution than the training images. The machine learning model inputs training images, performs inference on the training images, and creates output images. The machine learning model then compares the output image with the correct image and performs learning by adjusting parameters so as to reduce the difference between the output image and the correct image. The inference performance of a machine learning model depends on the type of model used and the training data used for machine learning.

[0005] For example, Japanese Patent Application Laid-Open Publication No. 2022-70035 describes a method for creating a machine learning model with increased robustness against noise by creating learning data using a medical input image that has undergone noise reduction as the correct image and an image obtained by reducing the correct image to a lower resolution as the training image.

[0006] However, a model that performs machine learning using training images and correct images with a higher resolution than the training images may generate an output image that includes high-frequency components that are not included in the training images. Since the high-frequency components that are not included in the training images are information that did not originally exist in the training images, the inference results are not guaranteed. Therefore, false patterns may occur due to incorrect inference. For example, in medical images, there is a risk that an image that appears to contain blood vessels in locations where there are no blood vessels actually may be generated.

[0007] The present invention has been made in consideration of the above circumstances, and aims to provide a machine learning image generation method, a machine learning method, and an endoscopic image processing device that can reduce the possibility that a machine learning model will generate an erroneous inferred output image.

[0008] A method for generating images for machine learning according to one aspect of the present invention performs a reduction process on components in at least some of the frequency bands higher than the Nyquist frequency of a candidate correct image that has a higher resolution than a training image that is paired with the correct image, and generates the correct image for machine learning to improve the resolution of an input image.

[0009] A machine learning method according to one aspect of the present invention trains a machine learning model using a candidate correct answer image having a higher resolution than a training image paired with the correct answer image, the candidate correct answer image being generated by performing a reduction process on components in at least a portion of a frequency band higher than the Nyquist frequency of the training image, or the candidate correct answer image being captured by an imaging device equipped with an optical low-pass filter that reduces components in at least a portion of a frequency band higher than the Nyquist frequency of the training image, and the training image.

[0010] An endoscopic image processing device according to one aspect of the present invention includes: a correct image generated by performing a reduction process on components in at least a portion of a frequency band higher than the Nyquist frequency of a training image that is a candidate correct image paired with the correct image, the training image; or a correct image captured by an imaging device equipped with an optical low-pass filter that reduces components in at least a portion of a frequency band higher than the Nyquist frequency of the training image; and a machine learning model connection unit connectable to a machine learning model trained using the training image and the correct image; and a processor, wherein the processor inputs a received endoscopic image to the machine learning model and causes the machine learning model to output the endoscopic image with improved resolution.

[0011] FIG. 1 is a diagram showing a general flow when a machine learning model performs learning to improve image resolution, becomes a trained model, and performs inference, according to each embodiment of the present invention. FIG. 1 is a diagram showing an example of a machine learning model configured using a neural network, according to each embodiment. FIG. 2 is a diagram showing an example of a machine learning model configured using a training image and a correct image, according to each embodiment. FIG. 3 is a diagram showing an example of a machine learning model learning using a correct image generated from a correct image candidate, according to each embodiment. FIG. 4 is a diagram showing first to fourth examples of frequency characteristics of a correct image generated from a correct image candidate, according to each embodiment. FIG. 5 is a graph showing a fifth example of frequency characteristics of a correct image generated from a correct image candidate, according to each embodiment. FIG. 6 is a graph showing a sixth example of frequency characteristics of a correct image generated from a correct image candidate, according to each embodiment. FIG. 7 is a block diagram showing an example of the configuration of a machine learning image generation device that generates a correct image by image processing according to a first embodiment of the present invention. FIG. 8 is a block diagram showing an example of the configuration of a machine learning image generation device that generates a correct image using an optical low pass filter according to a second embodiment of the present invention. FIG. 9 is a block diagram showing a first example of the configuration of a machine learning device according to a third embodiment of the present invention. FIG. 10 is a block diagram showing a second configuration example of a machine learning device in a fourth embodiment of the present invention. FIG. 11 is a block diagram showing a third configuration example of a machine learning device in a fifth embodiment of the present invention. FIG. 12 is a block diagram showing a fourth configuration example of a machine learning device in a sixth embodiment of the present invention. FIG. 13 is a block diagram showing a fifth configuration example of a machine learning device in a seventh embodiment of the present invention. FIG. 14 is a block diagram showing a configuration example of a machine learning device that acquires original images from an endoscope system in an eighth embodiment of the present invention. FIG. 15 is a block diagram showing a configuration example in which a trained model trained by a machine learning device is applied to an endoscope system in a ninth embodiment of the present invention.

[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings, but the present invention is not limited to the embodiments described below.

[0013] In the drawings, the same or corresponding elements are appropriately designated by the same reference numerals. It should be noted that the drawings are schematic, and that the length relationships, length ratios, and quantities of elements within a single drawing may differ from reality in order to simplify the explanation. Furthermore, there may be portions in which the length relationships and ratios differ between multiple drawings.

[0014] 1 to 16 show embodiments of the present invention. FIG. 1 is a diagram showing a general flow of inference when a machine learning model M undergoes learning to improve image spatial resolution using a machine learning method to become a trained model M. For simplicity, in the specification and drawings, the "machine learning model" may be simply referred to as the "model."

[0015] The machine learning model M is a mathematical model such as a neural network (NN). Fig. 2 is a diagram showing an example of the machine learning model M configured by a neural network according to each embodiment.

[0016] The neural network includes an input layer IL, a hidden layer HL, and an output layer OL. Each layer includes multiple neurons (units / nodes) ne, and each neuron ne in one layer is connected to each neuron ne in the next layer by an edge (line) ed, forming a network structure.

[0017] The machine learning model M may be a deep neural network (DNN) with multiple hidden layers HL, with a total depth of four or more layers. Alternatively, the machine learning model M may be a convolution neural network (CNN), an R-CNN (Regions with CNN features) using a CNN, or a fully convolutional network (FCN). Note that the machine learning is not limited to deep learning, and various other well-known learning methods may also be used.

[0018] The machine learning model M performs machine learning using training data LDS in a learning process LP that applies a machine learning method. The training data LDS is a training image dataset that includes pairs of training images IT and correct images IC. The training data LDS including the correct images IC is also called teacher data. The machine learning model M inputs the training images IT, forward propagates the input training image IT data from the input layer IL to the hidden layer HL to the output layer OL, and generates an output image.

[0019] Then, the machine learning model M calculates the difference (loss value) between the output image and the correct image IC using a loss function, and in backpropagation, adjusts parameters using an optimization algorithm to minimize the loss value and performs learning.

[0020] The machine learning model M, which has been trained through the learning process LP, performs inference on the input image I1 in the inference process IP, and outputs the inferred image I2. The trained machine learning model M is composed of, for example, a combination of an AI program (algorithm) and parameters optimized by learning.

[0021] 3 is a diagram showing examples of frequency components of training images IT and correct images IC according to each embodiment. Note that in this specification, frequency refers to spatial frequency.

[0022] Column A of Figure 3 shows an example graph of the frequency components of the training image IT and the reference image IC when a reference image IC with a higher resolution than the training image IT is used. When the reference image IC has a higher resolution than the training image IT, the reference image IC generally contains components (high-frequency components) in a frequency band (greater than f1 and less than f2) higher than the frequency band of the training image IT (0 or more and less than f1), as shown in the area surrounded by a two-dot chain line in Column A of Figure 3. Here, f1 is the upper limit frequency at which the training image IT has components, and f2 is the upper limit frequency at which the reference image IC has components.

[0023] If a combination of training image IT and ground truth image IC as shown in column A of Figure 3 is used, the machine learning model M will be trained to generate an output image in the frequency band from 0 to f2. However, since the high-frequency components in the frequency band from f1 to f2 are information that is not originally present in the training image IT, the inference results of the trained model cannot be guaranteed. As a result, the trained model may make an incorrect inference and generate a false pattern, for example, in a medical image, generating an image in which blood vessels appear to be present in a location where they are not actually present.

[0024] Column B of Figure 3 shows an example graph of the frequency components of the training image IT, the candidate correct image ICC, and the correct image IC when the correct image IC is generated from the candidate correct image ICC, which has a higher resolution than the training image IT.

[0025] The correct image candidate ICC shown in column B of Fig. 3 has the same frequency characteristics as the correct image IC shown in column A of Fig. 3. The correct image IC is generated by reducing the components of the correct image candidate ICC at frequencies higher than frequency f1 without reducing the components of the correct image candidate ICC at frequencies equal to or lower than frequency f1. Here, "reduce" means making the component amount of the correct image IC lower than the component amount of the correct image candidate ICC at the same frequency, and also includes setting the component amount to zero.

[0026] FIG. 4 is a diagram showing an example in which the machine learning model M learns using a correct image IC generated from a correct image candidate ICC according to each embodiment.

[0027] In the learning process LP, the machine learning model M uses the training image IT and a correct image IC generated by reducing the frequency components higher than frequency f1 from the correct image candidate ICC for learning. This reduces erroneous inferences of frequency bands that do not originally exist in the training image IT when inference is performed using the trained model M. For example, it is possible to suppress erroneous inferences in which an image that appears to contain blood vessels is generated in a position where there are no blood vessels.

[0028] FIG. 5 is a table showing first to fourth examples of frequency characteristics of a correct image IC generated from a correct image candidate ICC according to each embodiment.

[0029] With reference to each column in FIG. 5 , the training image IT is composed of a plurality of pixels arranged at a first pixel pitch, which is the sampling distance. The sampling frequency of the training image IT is the reciprocal of the sampling distance. The Nyquist frequency ftn of the training image IT is half the value of the sampling frequency. The training image IT has components in a frequency band equal to or lower than the Nyquist frequency ftn, and has no components in a frequency band higher than the Nyquist frequency ftn (referred to as a high frequency band, as appropriate).

[0030] The correct image candidate ICC is composed of a plurality of pixels arranged at a second pixel pitch smaller than the first pixel pitch, and the second pixel pitch is the sampling distance. Therefore, the correct image candidate ICC generally has components up to a frequency band higher than the Nyquist frequency ftn of the training image IT.

[0031] Furthermore, for each of the correct images IC shown in each column of FIG. 5, the amount of components in the frequency band below the Nyquist frequency ftn is not changed from that of the correct image candidate ICC.

[0032] The first example of the correct image IC shown in column A of FIG. 5 has frequency characteristics in which the amount of components of the correct image candidate ICC is rapidly and smoothly reduced when the frequency exceeds the Nyquist frequency ftn (becomes higher than the Nyquist frequency ftn), and the amount of components becomes zero at a frequency higher than the Nyquist frequency ftn.

[0033] The second example of the correct image IC shown in column B of Figure 5 has frequency characteristics in which the amount of components of the correct image candidate ICC in the high frequency band is generally reduced. Therefore, although the correct image IC also has components in the high frequency band, the amount of components of the correct image IC in the high frequency band is smaller than the amount of components of the correct image candidate ICC.

[0034] In the third example of the correct image IC shown in column C of Fig. 5, when the frequency exceeds the Nyquist frequency ftn (becomes higher than the Nyquist frequency ftn), the component amount of the correct image candidate ICC is discontinuously reduced to a value close to 0. The correct image IC has frequency characteristics with a component amount close to 0 in the high frequency band.

[0035] In the fourth example of the correct image IC shown in column D of FIG. 5, the component amounts are discontinuously set to 0 when the frequency exceeds the Nyquist frequency ftn (becomes higher than the Nyquist frequency ftn), and all component amounts in the high-frequency band are set to 0.

[0036] Incidentally, image resolution can be divided into relative resolution, whose value changes depending on the size of the paper surface of the printed material on which the image is displayed or the size of the display screen (i.e., depending on the number of pixels per unit length), and absolute resolution, which is determined by the number of pixels that make up the image. In this specification, resolution refers to absolute resolution.

[0037] To give one numerical example, the correct image candidate ICC is composed of 1280 pixels in width and 720 pixels in height, and the training image IT is composed of 640 pixels in width and 360 pixels in height.

[0038] 5A to 5C have components in a frequency band higher than the Nyquist frequency ftn. Therefore, in the above numerical example, the correct image IC is composed of, for example, 1280 horizontal and 720 vertical pixels, similar to the correct image candidate ICC.

[0039] On the other hand, the correct image IC shown in column D of Figure 5 does not have components in the frequency band higher than the Nyquist frequency ftn. Therefore, in the above numerical example, the correct image IC may be configured with 1280 horizontal and 720 vertical pixels, like the correct image candidate ICC, or may be configured with 640 horizontal and 360 vertical pixels, like the training image IT. In the latter case, the correct image IC can be generated by reducing the correct image candidate ICC.

[0040] In this way, the correct image IC does not need to have a higher resolution than the training image IT, as the correct image candidate ICC does, and may have the same resolution as the training image IT.

[0041] Note that the correct image IC shown in column A of Figure 5 has components up to a frequency slightly higher than the Nyquist frequency ftn, but the frequency band on the high-frequency side is narrower than that of the correct image candidate ICC. Therefore, in the above numerical example, the correct image IC may have a pixel configuration intermediate between 640 pixels wide by 360 pixels high and 1280 pixels wide by 720 pixels high. In this case, the correct image IC can also be generated by reducing the correct image candidate ICC.

[0042] The relationship between the resolution of the correct image IC and the training image IT is as described above, but the relationship between the spatial resolution is as follows.

[0043] When the component amounts of the correct image IC and the training image IT are compared in the frequency band below the Nyquist frequency ftn, the component amount of the correct image IC is greater than that of the training image IT. In particular, the component amount of the correct image IC is greater than that of the training image IT on the high frequency side in the frequency band below the Nyquist frequency ftn.

[0044] Therefore, even if the resolution of the correct image IC and the training image IT are the same, the spatial resolution of the correct image IC is higher than that of the training image IT. Generally, as the resolution of an image increases, the contrast of the image also increases. Therefore, the contrast of the correct image IC is higher than that of the training image IT.

[0045] FIG. 6 is a graph showing a fifth example of the frequency characteristics of a correct image IC generated from a correct image candidate ICC according to each embodiment.

[0046] In each example shown in Fig. 5, the correct image IC has a reduced amount of components compared to the correct image candidate ICC in frequency bands higher than the Nyquist frequency ftn, but the amount of components is not reduced in frequency bands equal to or lower than the Nyquist frequency ftn. In contrast, in the fifth example shown in Fig. 6, components are reduced in some frequency bands even in frequency bands equal to or lower than the Nyquist frequency ftn.

[0047] 6 does not have components in some of the higher frequency bands within the frequency band below the Nyquist frequency ftn and above half the Nyquist frequency ftn. In this case, errors may occur in inference using the trained model even in the frequency band below the Nyquist frequency ftn where the training image IT does not have information.

[0048] For this reason, in the fifth example shown in Fig. 6, processing is performed to reduce components in at least a portion of the frequency bands (frequency bands indicated by outlined double-headed arrows in Fig. 6) that are equal to or greater than half the Nyquist frequency ftn and equal to or less than the Nyquist frequency fcn of the correct image candidate ICC. In this way, components in the frequency bands equal to or less than the Nyquist frequency ftn may also be reduced or removed as necessary.

[0049] 6, the training image IT has components in the frequency band below frequency f3 indicated by the two-dot chain line, but has no components in the frequency band above frequency f3. Frequency f3 is equal to or greater than half the Nyquist frequency ftn (ftn / 2) but is lower than the Nyquist frequency ftn.

[0050] Therefore, the correct image IC is generated by reducing the components of the correct image candidate ICC in the frequency band higher than frequency f3 and equal to or lower than the Nyquist frequency fcn of the correct image candidate ICC. In the fifth example shown in Figure 6, similar to the second example shown in column B of Figure 5, the correct image IC has frequency characteristics in which the amount of components of the correct image candidate ICC is generally reduced in the frequency band higher than frequency f3.

[0051] FIG. 7 is a graph showing a sixth example of the frequency characteristics of a correct image IC generated from a correct image candidate ICC according to each embodiment.

[0052] The sixth example shown in Figure 7 is an example in which the frequency characteristics of the training image IT and the correct image candidate ICC are the same as those of the fifth example shown in Figure 6, but the frequency characteristics of the correct image IC are different.

[0053] That is, the correct image IC of the sixth example shown in Figure 7 is similar to the fourth example shown in column D of Figure 5 in that the component amounts are discontinuously set to 0 when the frequency exceeds f3 (becomes higher than frequency f3), and all component amounts in the frequency band higher than frequency f3 are set to 0.

[0054] 6 and 7, the correct image IC may be generated by reducing the amount of components of the correct image candidate ICC in a frequency band higher than the upper limit frequency f3 of the frequency band in which the training image IT has components. The upper limit frequency f3 can be determined, for example, by frequency analysis of the training image IT.

[0055] Furthermore, for example, when (1 / 2)≦α<1, frequency analysis may be omitted and the amount of components of the correct image candidate ICC in the frequency band above α×ftn may be reduced to generate the correct image IC.

[0056] In this case, if α=(½), the amount of components of the correct image candidate ICC in the frequency band equal to or greater than half the Nyquist frequency ftn (ftn / 2) is reduced to generate the correct image IC. When α=(½), erroneous inference can be prevented in many practical cases. [First embodiment]

[0057] 8 is a block diagram showing an example of the configuration of a machine learning image generating device 10 that generates a correct image IC by image processing in the first embodiment. The machine learning image generating device 10 shown in FIG. 4 generates a correct image IC from a correct image candidate ICC by a machine learning image generating method using image processing.

[0058] The machine learning image generation device 10 includes a frequency component adjustment unit 11. The frequency component adjustment unit 11 receives information on the Nyquist frequency ftn of the training image IT and a candidate correct image ICC having a higher resolution than the training image IT. The frequency component adjustment unit 11 performs a reduction process on components in at least a portion of the frequency band higher than the Nyquist frequency ftn of the candidate correct image ICC, thereby generating a correct image IC for use in machine learning that improves the resolution of the image.

[0059] The reduction process by the frequency component adjuster 11 may be a process of reducing components in all frequency bands higher than the Nyquist frequency ftn, as shown in columns B and C of FIG.

[0060] The reduction process by the frequency component adjustment unit 11 may also be a process of reducing all frequency band components higher than the Nyquist frequency ftn to zero (reducing them to zero) as shown in column D of Fig. 5. In other words, the reduction process includes a removal process of reducing frequency components to zero.

[0061] The frequency component adjustment unit 11 may reduce the components of all frequency bands higher than the Nyquist frequency ftn by downscaling the candidate correct image ICC. Here, the downscaling process is a process of changing the pixel configuration of the candidate correct image ICC so as to reduce the number of pixels. The candidate correct image IC generated by downscaling the candidate correct image ICC has a resolution lower than that of the candidate correct image ICC but equal to or higher than that of the training image IT.

[0062] In the above-mentioned numerical example in which the correct image candidate ICC is composed of 1280 pixels horizontally and 720 pixels vertically, and the training image IT is composed of 640 pixels horizontally and 360 pixels vertically, the reduction process is a process of reducing the correct image IC to a pixel configuration of 640 pixels or more but less than 1280 pixels horizontally and 360 pixels or more but less than 720 pixels vertically.

[0063] In the reduction process, it is preferable not to reduce the frequency components of the correct image candidate ICC below the Nyquist frequency ftn, so that the resolution of the correct image IC does not decrease below the resolution of the correct image candidate ICC in the frequency band below the Nyquist frequency ftn.

[0064] Not only is the resolution of the correct image IC not reduced below that of the correct image candidate ICC, but further processing may be performed to improve the resolution of the correct image IC above that of the correct image candidate ICC in the frequency band below the Nyquist frequency ftn.

[0065] The reduction process by the frequency component adjustment unit 11 may be a process of reducing the correct image candidate ICC so that the resolution becomes the same as the resolution of the training image IT.

[0066] As described above, the reduction process by the frequency component adjustment unit 11 is not limited to a frequency band higher than the Nyquist frequency ftn. In other words, the reduction process by the frequency component adjustment unit 11 may further include a process of reducing components in at least a part of a frequency band that is equal to or greater than half the Nyquist frequency ftn and is equal to or less than the Nyquist frequency ftn.

[0067] The reduction process by the frequency component adjuster 11 may be performed by applying low-pass filtering as image processing to the correct image candidate ICC.

[0068] Alternatively, the reduction process by the frequency component adjustment unit 11 may be performed by both low-pass filtering and downscaling. For example, the reduction process by the frequency component adjustment unit 11 may be performed by applying low-pass filtering to the correct image candidate ICC and then downscaling the correct image candidate ICC after the low-pass filtering.

[0069] According to the first embodiment, the frequency component adjustment unit 11 generates a correct image IC by reducing components in at least a part of the frequency band higher than the Nyquist frequency ftn for the correct image candidate ICC. This reduces the possibility that a machine learning model trained using the training images IT and the correct image IC will generate an erroneous inferred output image.

[0070] In particular, by setting all frequency band components higher than the Nyquist frequency ftn to zero, it is possible to significantly reduce erroneous inferences such as the occurrence of false patterns.

[0071] Furthermore, when reduction processing is performed using at least one of downscaling and low-pass filtering, the processing load can be significantly reduced and the processing speed can be improved compared to, for example, performing Fourier analysis to reduce components in a frequency band higher than a specific frequency (or equal to or higher than a specific frequency) and then performing an inverse Fourier transform to restore the components.

[0072] In addition, by preventing the resolution of the correct image IC from being lower than the resolution of the correct image candidate ICC in the frequency band equal to or lower than the Nyquist frequency ftn, the machine learning model can learn to generate an inferred image with higher resolution.

[0073] 9 is a block diagram showing an example configuration of a machine-learning image generating device 10 that generates a target image IC using an optical low-pass filter 17 in the second embodiment. In the second embodiment, parts that are the same as those in the first embodiment are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate. In the second embodiment, differences from the first embodiment will be mainly described.

[0074] In the first embodiment, a supervised image IC is generated from a supervised image candidate ICC by image processing. In contrast, a machine-learning image generating device 10 according to a second embodiment shown in Fig. 9 generates a supervised image IC by a machine-learning image generating method in which an image sensor 18 captures a subject light beam that has passed through an optical low-pass filter 17.

[0075] The imaging device 15 that generates the correct image IC is an imaging system that includes, for example, an imaging optical system 16 , an optical low-pass filter 17 , and an imaging element 18 .

[0076] The imaging optical system 16 condenses subject light beams and forms an optical image of the subject on the imaging surface of the imaging element 18. The imaging optical system 16 generally includes multiple optical lenses and an optical diaphragm, although the imaging optical system 16 may have other configurations.

[0077] The optical low pass filter 17 is disposed on the path of the subject light beam, and cuts out high-frequency components of the optical image while passing low-frequency components. The optical low pass filter 17 is generally disposed on the imaging surface of the image sensor 18. However, the optical low pass filter 17 may be disposed in another position.

[0078] The optical low-pass filter 17 reduces high-frequency components in at least a part of the frequency band in the optical image that is higher than the Nyquist frequency ftn of the training image IT. As a specific example, the optical low-pass filter 17 reduces frequency components in the optical image that are equal to or higher than half the Nyquist frequency ftn.

[0079] The image sensor 18 photoelectrically converts the optical image of the subject formed by the imaging optical system 16 via the optical low-pass filter 17 and outputs a signal. The image corresponding to the output signal becomes a correct image IC in which components in at least a part of the frequency band higher than the Nyquist frequency ftn are reduced.

[0080] According to the second embodiment, it is possible to generate an appropriate correct image IC for learning that can reduce inference errors by capturing an optical image of a subject that is formed via the optical low pass filter 17. In this case, there is no need for image processing such as that in the first embodiment, which involves performing reduction processing on the correct image candidate ICC to generate the correct image IC.

[0081] The correct image IC used by the machine learning model M for learning together with the training image IT may be either the correct image IC generated by the configuration of the first embodiment or the correct image IC generated by the configuration of the second embodiment. [Third Embodiment]

[0082] 10 is a block diagram showing a first configuration example of a machine learning device 1 in the third embodiment. In the third embodiment, parts that are the same as those in the first and second embodiments are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate. In the third embodiment, differences from the first and second embodiments will be mainly described.

[0083] The machine learning device 1 includes a machine learning image generation device 10 and a model learning processing unit 20. The machine learning image generation device 10 of this embodiment generates training images IT and supervised images IC from supervised image candidates ICC. The model learning processing unit 20 performs machine learning using the generated training images IT and supervised images IC.

[0084] The machine learning device 1 may be configured such that a processor such as an ASIC (Application Specific Integrated Circuit) including a CPU (Central Processing Unit) or an FPGA (Field Programmable Gate Array) reads and executes a processing program stored in a storage device (or recording medium) such as a memory, thereby fulfilling the functions of each unit. Alternatively, the machine learning device 1 may be configured as a dedicated electronic circuit that fulfills the functions of each unit.

[0085] The machine learning image generating device 10 includes a frequency component adjusting unit 11 and a training image generating unit 12. The frequency component adjusting unit 11 and the training image generating unit 12 receive input of candidate correct image ICCs.

[0086] The training image generation unit 12 further inputs training image imaging system information. The training image imaging system information is imaging system information assumed for the training image IT. Note that in other embodiments described later, imaging system information assumed for the correct image IC (correct image imaging system information) is used. Whether the assumed image is the training image IT or the correct image IC, the imaging system information includes pixel count information, color information, optical characteristic information, noise characteristic information, color filter information, etc.

[0087] The pixel count information is information about the pixel configuration of an image, and corresponds to the pixel configuration of the imaging element of the endoscope to which the machine learning model M is applied. In the above-described numerical example, the pixel count information assumed for the training image IT is information consisting of 720 horizontal and 360 vertical pixels.

[0088] The color information is information for correcting the difference between the color of the input image (in this embodiment, the correct image candidate ICC) and the color of the output image (in this embodiment, the training image IT).

[0089] The optical characteristic information includes, for example, information (corrective PSF) for correcting a point spread function (PSF). The PSF is a point spread function that indicates the spatial distribution of an image formed from a point light source by an imaging optical system.

[0090] The PSF of an input image (in this embodiment, the candidate correct image ICC) is generally different from the PSF assumed for the output image (in this embodiment, the training image IT). Therefore, the optical characteristic information for converting the PSF of the input image into the PSF assumed for the output image is the correction PSF.

[0091] Here, the PSF of an imaging optical system generally depends on the position within the image. Therefore, a correction PSF that depends on the coordinates within the image may be used. Furthermore, using a uniform correction PSF across the entire screen has the advantage of reducing the circuit size and calculation costs.

[0092] Examples of PSFs assumed for the training images IT include the actual PSF of the imaging optical system of the endoscope to which the machine learning model M is applied, or a virtual PSF. The virtual PSF may be blurred more than the actual PSF. By using a virtual PSF that is blurred more than the actual PSF, it is expected that the resolution of the image inferred by the trained model M will be significantly improved. On the other hand, by using a virtual PSF with less blur, it is expected that the image inferred by the trained model M will have reduced resolution but fewer erroneous inferences.

[0093] The noise characteristic information is information that indicates, for example, assuming random noise, how much noise to add to which pixel value at which position. The noise characteristic information preferably includes both information on the noise position and information on the noise amount, but may also include at least one of information on the noise position, the noise amount, and the standard deviation of the noise to be added. When using noise standard deviation information as noise characteristic information, the noise standard deviation information may be uniform regardless of pixel value, or may vary depending on the pixel value.

[0094] The color filter information is information regarding, for example, whether the endoscope to which the machine learning model M is applied is a type equipped with an imaging element having a Bayer array color filter, or a type that sequentially irradiates RGB illumination light onto a monochrome imaging element to sequentially acquire R, G, and B images. Hereinafter, an image acquired by an imaging element having a Bayer array color filter will be referred to as a Bayer image, and an image acquired by a monochrome imaging element with frame sequential irradiation will be referred to as a frame sequential image.

[0095] The training image generation unit 12 generates training images IT from the correct image candidates ICC based on the training image capturing system information as described above.

[0096] Based on the pixel number information, the training image generation unit 12 reduces the correct image candidate ICC from, for example, 1280 pixels wide by 720 pixels high to 640 pixels wide by 360 pixels high, to generate a first intermediate image.

[0097] The training image generation unit 12 performs color correction processing on the first intermediate image based on the color information to generate a second intermediate image.

[0098] Furthermore, the training image generation unit 12 performs PSF correction on the second intermediate image by applying a correction PSF that depends on coordinates within the image, and generates a third intermediate image having blur corresponding to the PSF assumed for the training image IT.

[0099] Furthermore, the training image generating unit 12 generates a fourth intermediate image by adding random noise or the like to the third intermediate image based on noise characteristic information assumed for the training image IT.

[0100] If the correct image candidate ICC is a Bayer image and an endoscope applying machine learning model M generates a Bayer image, or if the correct image candidate ICC is a frame sequential image and an endoscope applying machine learning model M generates a frame sequential image, the training image generation unit 12 outputs a fourth intermediate image as a training image IT.

[0101] On the other hand, if the correct image candidate ICC is a frame sequential image and the endoscope to which the machine learning model M is applied generates a Bayer image, the training image generation unit 12 converts the fourth intermediate image, which is a frame sequential image, into a Bayer image and outputs it as the training image IT.

[0102] Furthermore, when the endoscope generates frame-sequential images, the candidate correct image ICC is preferably a frame-sequential image. However, the training image generator 12 may demosaicing a fourth intermediate image generated from the candidate correct image ICC of the Bayer image to convert it into a frame-sequential image and output it as the training image IT.

[0103] Note that, although the above describes an example of the processing order in which the training image generation unit 12 generates a training image IT from a correct image candidate ICC based on each piece of information contained in the training image capture system information, the processing is not limited to the above order and may be performed in another order.

[0104] The training image generation unit 12 transmits the generated training images IT to the model learning processing unit 20 and transmits information on the Nyquist frequency ftn to the frequency component adjustment unit 11.

[0105] As described above, the frequency component adjustment unit 11 receives information on the Nyquist frequency ftn of the training image IT and the candidate correct image ICC, which has a higher resolution than the training image IT. As in the first embodiment, the frequency component adjustment unit 11 reduces components in the frequency band higher than the Nyquist frequency ftn (or equal to or greater than half the Nyquist frequency ftn) for the candidate correct image ICC, and generates the correct image IC.

[0106] 4, the model learning processing unit 20 causes the machine learning model M to learn through a learning process LP using the training images IT and the correct images IC. The model learning processing unit 20 outputs the learned model M.

[0107] The third embodiment provides substantially the same effects as the first embodiment. Furthermore, the third embodiment allows both the training image IT and the correct image IC to be generated once the correct image candidate ICC is obtained.

[0108] Furthermore, according to the third embodiment, the training image generation unit 12 corrects the input image using training image capture system information, so that it is possible to generate training images IT that are highly consistent with the correct image IC. A trained model M that has undergone machine learning using highly consistent training images IT and the correct image IC is expected to generate post-inference images with more appropriately improved resolution. [Fourth embodiment]

[0109] 11 is a block diagram showing a second configuration example of the machine learning device 1 in the fourth embodiment. In the fourth embodiment, parts that are the same as those in the first to third embodiments are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate. In the fourth embodiment, differences from the first to third embodiments will be mainly described.

[0110] The machine learning device 1 includes a machine learning image generation device 10 and a model learning processing unit 20. The machine learning image generation device 10 of this embodiment generates training images IT and supervised images IC from source images. The model learning processing unit 20 performs machine learning using the generated training images IT and supervised images IC.

[0111] The machine learning image generating device 10 includes a frequency component adjusting unit 11, a training image generating unit 12, and a correct image candidate generating unit 13. The training image generating unit 12 and the correct image candidate generating unit 13 input an original image.

[0112] The training image generation unit 12 further receives the training image capture system information as described above. The training image generation unit 12 generates training images IT from the original images based on the training image capture system information. The training image generation unit 12 transmits the generated training images IT to the model learning processing unit 20 and transmits information on the Nyquist frequency ftn to the frequency component adjustment unit 11.

[0113] The correct image candidate generation unit 13 further inputs correct image imaging system information. As described above, the correct image imaging system information is imaging system information assumed for the correct image IC. The values ​​of each piece of information included in the correct image imaging system information generally differ from the values ​​of each piece of information included in the training image imaging system information. The correct image candidate generation unit 13 generates a correct image candidate ICC from the original image based on the correct image imaging system information in a procedure similar to that of the training image generation unit 12. The correct image candidate generation unit 13 transmits the generated correct image candidate ICC to the frequency component adjustment unit 11.

[0114] As described above, the frequency component adjustment unit 11 generates the correct image IC from the correct image candidate ICC based on the Nyquist frequency ftn.

[0115] As described above, the model learning processing unit 20 causes the machine learning model M to learn using the training images IT and the correct images IC, and outputs the learned model M.

[0116] The fourth embodiment achieves substantially the same effects as the first and third embodiments. Furthermore, by using the correct image capturing system information in addition to the training image capturing system information, both the training image IT and the correct image IC can be generated from a single source image. [Fifth Embodiment]

[0117] 12 is a block diagram showing a third configuration example of the machine learning device 1 in the fifth embodiment. In the fifth embodiment, parts that are the same as those in the first to fourth embodiments are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate. In the fifth embodiment, differences from the first to fourth embodiments will be mainly described.

[0118] The machine-learning image generating device 10 of this embodiment generates training images IT and correct images IC from source images, similar to the fourth embodiment. However, the machine-learning image generating device 10 of this embodiment has a different processing order from the fourth embodiment.

[0119] The machine learning device 1 includes a machine learning image generation device 10 and a model learning processing unit 20. The machine learning image generation device 10 includes a frequency component adjustment unit 11, a training image generation unit 12, and a correct answer image generation unit 14. The frequency component adjustment unit 11 and the training image generation unit 12 input an original image.

[0120] The training image generation unit 12 further receives training image capture system information. The training image generation unit 12 generates training images IT from the original images based on the training image capture system information. The training image generation unit 12 transmits the generated training images IT to the model learning processing unit 20 and transmits information on the Nyquist frequency ftn to the frequency component adjustment unit 11.

[0121] The frequency component adjustment unit 11 receives the original image as a supervised image candidate ICC. Similar to the frequency component adjustment unit 11 described in the first embodiment, the frequency component adjustment unit 11 reduces components in a frequency band higher than the Nyquist frequency ftn (or equal to or greater than half the Nyquist frequency ftn) from the supervised image candidate ICC to generate an adjusted supervised image candidate ICC. The frequency component adjustment unit 11 transmits the adjusted supervised image candidate ICC to the supervised image generation unit 14.

[0122] The correct image generation unit 14 receives the adjusted correct image candidate ICC and further inputs the correct image capturing system information. Based on the correct image capturing system information, the correct image generation unit 14 generates a correct image IC from the adjusted correct image candidate ICC in the same procedure as the correct image candidate generation unit 13. The correct image generation unit 14 transmits the generated correct image IC to the model learning processing unit 20.

[0123] As described above, the model learning processing unit 20 causes the machine learning model M to learn using the training images IT and the correct images IC, and outputs the learned model M.

[0124] In the fourth embodiment, the original image is subjected to image processing based on the correct image capturing system information, and then the frequency component adjustment unit 11 performs the reduction processing. In contrast, in the fifth embodiment, the original image is subjected to the reduction processing by the frequency component adjustment unit 11, and then the original image is subjected to image processing based on the correct image capturing system information to generate the correct image IC. The fifth embodiment provides substantially the same effects as the fourth embodiment. [Sixth embodiment]

[0125] 13 is a block diagram showing a fourth configuration example of the machine learning device 1 in the sixth embodiment. In the sixth embodiment, parts that are the same as those in the first to fifth embodiments are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate. In the sixth embodiment, differences from the first to fifth embodiments will be mainly described.

[0126] The machine learning device 1 of this embodiment has the same components as those of the fourth embodiment, but the information to be transmitted is different.

[0127] The machine learning device 1 includes a machine learning image generation device 10 and a model learning processing unit 20. The machine learning image generation device 10 includes a frequency component adjustment unit 11, a training image generation unit 12, and a correct image candidate generation unit 13.

[0128] As in the fourth embodiment, the training image generation unit 12 generates training images IT from the original images, transmits the training images IT to the model learning processing unit 20, and transmits information on the Nyquist frequency ftn to the frequency component adjustment unit 11. The training image generation unit 12 of this embodiment further transmits additional information added during generation to the correct image candidate generation unit 13.

[0129] The additional information at the time of generation is information indicating the position and amount of noise added to the training image IT by the training image generation unit 12 and / or standard deviation information of the noise included in the training image capture system information.

[0130] The correct image candidate generation unit 13 processes the original image based on the correct image capture system information. In this case, the correct image capture system information of the sixth embodiment does not include noise characteristic information. Therefore, the correct image candidate generation unit 13 adds noise to the processed original image based on the additional information at the time of generation to generate a correct image candidate ICC. Furthermore, if the correct image candidate ICC contains multiple position candidates corresponding to the positions of the noise added to the training image IT, the same noise as added to the training image IT may be added to the pixel values ​​of each of the multiple position candidates. Alternatively, the same noise as added to the training image IT may be added to the pixel values ​​of some of the multiple position candidates, and newly generated random noise based on the standard deviation information of the noise may be added to the pixel values ​​of the other position candidates.

[0131] Thereafter, the correct image candidate generating unit 13 transmits the generated correct image candidate ICC to the frequency component adjusting unit 11 .

[0132] As described above, the frequency component adjustment unit 11 generates the correct image IC from the correct image candidate ICC based on the Nyquist frequency ftn.

[0133] In the fourth embodiment, different noises are added to the training image IT and the reference image IC, whereas in the sixth embodiment, the positions and amounts of noise added to the training image IT and the reference image IC are matched at corresponding locations.

[0134] The sixth embodiment achieves substantially the same effects as the fourth embodiment. Furthermore, according to the sixth embodiment, the position and amount of noise added to the training image IT and the correct image IC are consistent, so the machine learning model M can learn to infer the correct image IC from the training image IT while prioritizing improvement of resolution. [Seventh embodiment]

[0135] 14 is a block diagram showing a fifth configuration example of a machine learning device 1 in the seventh embodiment. In the seventh embodiment, parts that are the same as those in the first to sixth embodiments are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate. In the seventh embodiment, differences from the first to sixth embodiments will be mainly described.

[0136] The machine learning device 1 of this embodiment has the same components as those of the fifth embodiment, but the information to be transmitted is different.

[0137] The machine learning device 1 includes a machine learning image generation device 10 and a model learning processing unit 20. The machine learning image generation device 10 includes a frequency component adjustment unit 11, a training image generation unit 12, and a correct answer image generation unit 14.

[0138] As in the fifth embodiment, the training image generation unit 12 generates training images IT from the original images, transmits the training images IT to the model learning processing unit 20, and transmits information on the Nyquist frequency ftn to the frequency component adjustment unit 11. The training image generation unit 12 of this embodiment further transmits the above-mentioned additional information added at the time of generation to the correct image generation unit 14.

[0139] The correct image generation unit 14 processes the adjusted correct image candidate based on the correct image capturing system information. At this time, the correct image capturing system information of the seventh embodiment does not include noise characteristic information. Therefore, the correct image generation unit 14 adds noise to the processed adjusted correct image candidate based on the additional information at the time of generation to generate the correct image IC.

[0140] In the fifth embodiment described above, individual noise was added to the training image IT and the correct image IC, but in the seventh embodiment, the position and amount of noise added to the training image IT and the correct image IC are consistent at corresponding locations.

[0141] Thereafter, the correct image generating unit 14 transmits the generated correct image IC to the model learning processing unit 20 .

[0142] The seventh embodiment achieves substantially the same effects as the fifth embodiment. Furthermore, according to the seventh embodiment, the position and amount of noise added to the training image IT and the correct image IC are consistent, so the machine learning model M can learn to infer the correct image IC from the training image IT while prioritizing improvement of resolution. [Eighth Embodiment]

[0143] 15 is a block diagram showing an example configuration of a machine learning device 1 that acquires an original image from an endoscope system 31 in the eighth embodiment. In the eighth embodiment, parts that are the same as those in the first to seventh embodiments are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate. In the eighth embodiment, differences from the first to seventh embodiments will be mainly described.

[0144] The machine learning device 1 of this embodiment acquires an original image from an endoscope system 31 .

[0145] The machine learning device 1 includes a machine learning image generation device 10 and a model learning processing unit 20. The machine learning image generation device 10 includes a frequency component adjustment unit 11, a training image generation unit 12, and a correct image candidate generation unit 13.

[0146] The process of generating a trained model M by the machine learning device 1 is basically the same as that described with reference to Fig. 11. However, the machine learning device 1 further acquires original image capturing system information and original image light source information from the endoscope system 31 and performs processing.

[0147] 11 is shown as an example in Fig. 15, but the configurations of the machine learning device 1 shown in Fig. 12 to 14 may also be applied to this embodiment. Furthermore, when an image generated by the endoscope system 31 is used as the correct image candidate ICC, the configuration of the machine learning device 1 shown in Fig. 10 may also be applied to this embodiment.

[0148] The endoscope system 31 includes a light source unit 32, an imaging device 33, and a memory 34. The light source unit 32 irradiates illumination light onto an object 90, which is a subject. A subject light beam, which is return light from the object 90, is incident on the imaging device 33.

[0149] The imaging device 33 is roughly the same as the imaging device 15 shown in Fig. 9 except that the optical low-pass filter 17 is removed. The imaging device 33 forms an image of the subject light beam using the imaging optical system 16, captures an image using the image sensor 18, and outputs an original image. The original image is input to the training image generation unit 12 and the correct image candidate generation unit 13, as described with reference to Fig. 11 .

[0150] The memory 34 is a storage medium that non-volatilely stores information related to the endoscope system 31. The information stored in the memory 34 includes original image capturing system information and original image light source information. The original image capturing system information is imaging system information related to the imaging device 33 that acquires the original image.

[0151] The original image imaging system information includes pixel count information of the imaging element 18, color information of the image acquired by the imaging device 33, optical characteristic information including the PSF of the imaging optical system 16, noise characteristic information related to the imaging element 18 and the readout circuit from the imaging element 18, color filter information indicating whether the imaging element 18 acquires a Bayer image or a frame sequential image, and the like.

[0152] The light source unit 32 is configured to be able to emit illumination light corresponding to, for example, a plurality of types of observation modes. The observation modes include, for example, a white light imaging (WLI) mode, a narrow band imaging (NBI) mode, etc. The original image light source information is information indicating the type of illumination light (WLI illumination light, NBI illumination light, etc.) emitted by the light source unit 32 depending on the observation mode.

[0153] The endoscope system 31 transmits the original image capturing system information and the original image light source information corresponding to the observation mode from the memory 34 to the machine learning device 1.

[0154] The training image generation unit 12 and the correct image candidate generation unit 13 of the machine learning device 1 receive original image capturing system information and original image light source information from the endoscope system 31 .

[0155] The training image generation unit 12 changes, for example, the correction PSF and noise characteristic information included in the training image imaging system information according to the type of illumination light (WLI illumination light, NBI illumination light, etc.) based on the original image light source information. At this time, the training image generation unit 12 may further change the correction PSF and noise characteristic information included in the training image imaging system information based on the optical characteristic information and noise characteristic information included in the original image imaging system information.

[0156] Similarly, the correct image candidate generation unit 13 changes, for example, the correction PSF and noise characteristic information included in the correct image capturing system information according to the type of illumination light (WLI illumination light, NBI illumination light, etc.) based on the original image light source information. At this time, the correct image candidate generation unit 13 may further change the correction PSF and noise characteristic information included in the correct image capturing system information based on the optical characteristic information and noise characteristic information included in the original image capturing system information.

[0157] In addition, when applying the trained model M to the same model (or even the same individual) as the endoscopic system 31 that acquired the original image, it is possible to only add noise without performing PSF correction.

[0158] Furthermore, the training image generating unit 12 and the correct image candidate generating unit 13 may change, for example, the reduction process as needed based on the pixel count information of the image sensor 18 included in the original image imaging system information.

[0159] Furthermore, the training image generation unit 12 and the correct image candidate generation unit 13 may change, for example, the color correction process as needed based on the color information included in the original image capturing system information.

[0160] The training image generation unit 12 and the correct image candidate generation unit 13 may change the conversion process, for example, from a Bayer image to a frame sequential image, or from a frame sequential image to a Bayer image, as necessary based on the color filter information contained in the original image imaging system information.

[0161] According to the eighth embodiment, it is possible to obtain an original image from the endoscope system 31 and achieve substantially the same effects as the third to seventh embodiments. Furthermore, according to the eighth embodiment, it is possible to appropriately change the training image imaging system information and the ground truth image imaging system information based on at least one of the original image imaging system information and the original image light source information, thereby generating an appropriate learning image data set. [Ninth embodiment]

[0162] 16 is a block diagram showing a configuration example in which a trained model M trained by a machine learning device 1 is applied to an endoscope system 41 in the ninth embodiment. In the ninth embodiment, parts that are the same as those in the first to eighth embodiments are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate. In the ninth embodiment, differences from the first to eighth embodiments will be mainly described.

[0163] The machine learning device 1 includes a machine learning image generation device 10 and a model learning processing unit 20. The machine learning image generation device 10 includes a frequency component adjustment unit 11, a training image generation unit 12, and a correct image candidate generation unit 13.

[0164] The process of generating the trained model M by the machine learning device 1 is similar to that described with reference to Fig. 15. That is, in addition to the process described with reference to Fig. 11, original image light source information is further acquired and processed.

[0165] 11 is shown as an example in Fig. 16, but the configurations of the machine learning device 1 shown in Fig. 12 to Fig. 14 may also be applied to this embodiment. Furthermore, when generating training images IT and supervised images IC based on candidate supervised images ICC instead of source images, the configuration of the machine learning device 1 shown in Fig. 10 may also be applied to this embodiment.

[0166] The endoscopic system 41 includes an endoscope 42 and an endoscopic image processing device 44. The imaging device 43 is provided, for example, in the endoscope 42. However, the imaging device 43 may also be provided in a camera head, and the camera head may be attached to an eyepiece of the endoscope 42. The imaging device 43 is connected to the endoscopic image processing device 44.

[0167] The imaging device 43, like the imaging device 33 shown in FIG. 15, has a configuration similar to that of the imaging device 15 shown in FIG. 9 except that the optical low-pass filter 17 is removed.

[0168] The imaging device 43 forms an image of the subject light beam using the imaging optical system 16, captures an image using the imaging element 18, and outputs an endoscopic image. The endoscopic image output from the imaging device 43 becomes an input image I1 (see FIG. 1) to the endoscopic image processing device 44.

[0169] The endoscopic image processing device 44 includes, for example, a processor 44a and a memory 44b. The processor 44a is configured by an ASIC (Application Specific Integrated Circuit) including a CPU (Central Processing Unit) or an FPGA (Field Programmable Gate Array). However, the endoscopic image processing device 44 may be configured as a dedicated electronic circuit that performs the functions of the trained model M.

[0170] The memory 44b is a storage medium that stores (non-volatilely stores) processing programs that realize the functions of each circuit. The processor 44a is connected to wiring extending from the memory 44b. The processor 44a reads and executes the processing programs stored in the memory 44b, thereby realizing the functions of the endoscopic image processing device 44. For example, the endoscopic image processing device 44 realizes an endoscopic image processing method by executing the processing programs.

[0171] The trained model M generated by the machine learning device 1, i.e., the combination of the AI ​​program (algorithm) and the parameters optimized by learning, is stored in the memory 44b of the endoscopic image processing device 44. The memory 44b and the wiring derived from the memory 44b constitute a machine learning model connection unit that can be connected to the trained model M.

[0172] The processor 44a executes the endoscopic image processing method to perform inference on the input image I1 using the trained model M, and outputs the inferred image I2 (see FIG. 1 ). As a result of appropriate inference, the inferred image I2 is an endoscopic image with improved resolution compared to the input image I1.

[0173] According to the ninth embodiment, the trained model M trained using the configurations of the third to seventh embodiments is applied to the endoscopic system 41 to perform appropriate inference with fewer errors, and an output image with improved resolution is obtained.

[0174] Although the present invention has been described above primarily as a machine learning image generation method, a machine learning method, an endoscopic image processing method, a machine learning image generation device, a machine learning device, and an endoscopic image processing device, it is not limited thereto. For example, the present invention may be a computer program for causing a computer to perform processing in accordance with the machine learning image generation method, the machine learning method, and the endoscopic image processing method. Furthermore, the present invention may also be a non-transitory recording medium readable by a computer that records the computer program.

[0175] Examples of recording media for storing a computer program product include portable recording media such as flexible disks, CD-ROMs (Compact Disc Read Only Memory), and DVDs (Digital Versatile Discs), as well as recording media such as HDDs (Hard Disk Drives) and SSDs (Solid State Drives). The recording media may store only a portion of the computer program, not just the entire program. Furthermore, all or part of the computer program may be distributed or provided via a communications network. By installing the computer program from the recording medium onto a computer, or by downloading and installing the computer program via a communications network, a user can read the computer program and execute all or part of the operations, allowing the computer to execute processing in accordance with the machine learning image generation method, the machine learning method, and the endoscopic image processing method.

[0176] Furthermore, the present invention is not limited to the above-described embodiments. In the implementation stage, the components can be modified and embodied without departing from the spirit of the invention. Furthermore, various aspects of the invention can be formed by appropriately combining multiple components disclosed in the above embodiments. For example, some components may be deleted from all the components disclosed in the embodiments. Furthermore, components from different embodiments may be appropriately combined. In this way, it goes without saying that various modifications and applications are possible within the scope of the gist of the invention.

Claims

1. For candidate correct images with higher resolution than the training image paired with the correct image, performing a reduction process on components of at least a part of a frequency band of the training image that is higher than the Nyquist frequency; A method for generating images for machine learning, characterized by generating a correct image for machine learning to improve the resolution of an input image.

2. The method for generating an image for machine learning according to claim 1 , wherein the reduction process reduces components of all frequency bands higher than the Nyquist frequency.

3. The method for generating images for machine learning according to claim 1 , wherein the reduction process reduces all frequency components higher than the Nyquist frequency to zero.

4. The method for generating images for machine learning described in claim 1, characterized in that the reduction process reduces the correct image candidate so that the resolution of the correct image candidate is lower than that of the correct image candidate and higher than that of the training image without reducing the frequency band components of the correct image candidate below the Nyquist frequency of the training image.

5. The method for generating images for machine learning according to claim 4 , wherein the reduction process reduces the correct image candidate so that the resolution thereof is the same as that of the training image.

6. The method for generating images for machine learning according to claim 1, characterized in that the reduction process further reduces components of at least a portion of frequency bands between half the Nyquist frequency and the Nyquist frequency.

7. The method for generating images for machine learning according to claim 1 , wherein the reduction processing is performed by applying a low-pass filter processing to the candidate correct image.

8. The method for generating images for machine learning according to claim 1, characterized in that the reduction process is performed by applying a low-pass filter process to the candidate correct image and reducing the candidate correct image after applying the low-pass filter process.

9. a correct answer image generated by performing a reduction process on components in at least a part of a frequency band higher than the Nyquist frequency of a training image, the correct answer image being a candidate correct answer image having a higher resolution than a training image paired with the correct answer image; Alternatively, the target image is captured by an imaging device equipped with an optical low-pass filter that reduces components of at least a part of a frequency band higher than the Nyquist frequency of the training image; and the training images; A machine learning method characterized by using a machine learning algorithm to perform inference so as to improve the resolution of an input image and train a machine learning model for generating an output image.

10. When the reduction process is performed, The machine learning method according to claim 9 , wherein the reduction process reduces components in all frequency bands higher than the Nyquist frequency.

11. When the reduction process is performed, The machine learning method according to claim 9 , wherein the reduction process reduces all frequency components higher than the Nyquist frequency to zero.

12. When the reduction process is performed, The machine learning method according to claim 9, characterized in that the reduction process reduces the correct image candidate so that the resolution of the correct image candidate is lower than that of the correct image candidate but higher than that of the training image, without reducing the frequency band components of the correct image candidate below the Nyquist frequency of the training image.

13. a correct answer image generated by performing a reduction process on components in at least a part of a frequency band higher than the Nyquist frequency of a training image, the correct answer image being a candidate correct answer image having a higher resolution than a training image paired with the correct answer image; Alternatively, the target image is captured by an imaging device equipped with an optical low-pass filter that reduces components of at least a part of a frequency band higher than the Nyquist frequency of the training image; and the training images; a machine learning model connection unit that can be connected to a machine learning model trained by a machine learning method that trains a machine learning model using a processor; Including, the processor: inputting the received endoscopic image into the machine learning model for performing inference to improve the resolution of the input image and generating an output image; An endoscopic image processing device characterized in that the endoscopic image with improved resolution is output from the machine learning model.

14. 14. The image processing device for endoscopes according to claim 13, wherein the reduction processing reduces components of all frequency bands higher than the Nyquist frequency.

15. 14. The image processing device for endoscopes according to claim 13, wherein the reduction process reduces all frequency band components higher than the Nyquist frequency to zero.

16. The endoscopic image processing device according to claim 13, characterized in that the reduction process reduces the correct image candidate so that the resolution of the correct image candidate is lower than that of the correct image candidate and higher than that of the training image, without reducing the frequency band components of the correct image candidate below the Nyquist frequency of the training image.

17. The machine learning model connection unit a storage medium storing the machine learning model; and Wiring extending from the storage medium; 14. The image processing device for endoscopes according to claim 13, further comprising:

18. A processor comprising: For candidate correct images with higher resolution than the training image paired with the correct image, performing a reduction process on components of at least a part of a frequency band of the training image that is higher than the Nyquist frequency; An endoscopic image processing program that generates the correct image for machine learning to improve the resolution of an input image.