Image generation method for machine learning, machine learning method, and image processing device for endoscopes.
By reducing high-frequency components in ground truth images used for training, the machine learning model generates more accurate output images, addressing the issue of misinferences in super-resolution tasks.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- OLYMPUS MEDICAL SYST CORP
- Filing Date
- 2023-03-31
- Publication Date
- 2026-05-26
AI Technical Summary
Existing machine learning models for super-resolution generate output images with misinferences, particularly in medical images, due to the inclusion of high-frequency components not present in the training images, leading to false patterns such as the appearance of blood vessels where they do not exist.
Reduce components in frequency bands higher than the Nyquist frequency of the training image when generating ground truth images with higher resolution, using either image processing or an optical low-pass filter to create a suitable ground truth image for training, thereby improving the resolution of input images.
Reduces the likelihood of misinferences in output images by ensuring the ground truth images used for training do not contain high-frequency components not originally present in the training images, enhancing the accuracy and resolution of the machine learning model's inference.
Smart Images

Figure 0007866143000001 
Figure 0007866143000002 
Figure 0007866143000003
Abstract
Description
Technical Field
[0001] The present invention relates to an image generation method for machine learning that generates a correct image from a correct image candidate having a higher resolution than a training image, a machine learning method for training a machine learning model using a training image and a correct image, and an endoscopic image processing apparatus using the trained machine learning model.
Background Art
[0002] As technologies for improving the spatial resolution of images, there are super-resolution technologies, edge enhancement technologies, and the like. The super-resolution technology increases the number of pixels constituting the image to enhance the resolution compared to the original image. The edge enhancement technology enlarges specific spatial frequency components in the image to enhance the edges.
[0003] Technologies for performing super-resolution using a machine learning model have been proposed conventionally. The machine learning model makes an inference on an input image, creates an inferred image with enhanced resolution, and outputs the image.
[0004] For example, the machine learning model performs machine learning for super-resolution using learning data including a pair of a training image and a correct image having a higher resolution than the training image. The machine learning model inputs the training image, makes an inference on the training image, and creates an output image. Then, the machine learning model performs learning to adjust the parameters so that the difference from the correct image becomes small by comparing the output image and the correct image. The inference performance of the machine learning model depends on what kind of model is used and the learning data used for machine learning.
[0005] For example, Japanese Patent Laid-Open No. 2022-70035 describes a method for creating a machine learning model with enhanced robustness against noise by creating learning data using, as a correct image, an image obtained by reducing noise in a medical input image, and, as a training image, an image obtained by downsizing the correct image to a low resolution.
[0006] Incidentally, a machine learning model that uses training images and ground truth images with higher resolution than the training images may generate output images that contain high-frequency components not present in the training images. Since these high-frequency components are information that was not originally present in the training images, the inference results cannot be guaranteed. Consequently, false patterns may be generated due to misinference, and for example, in medical images, there is a risk that images may be generated that appear to have blood vessels in locations where there are actually no blood vessels.
[0007] The present invention has been made in view of the above circumstances, and aims to provide a machine learning image generation method, a machine learning method, and an endoscope image processing device that can reduce the possibility that a machine learning model will generate an output image with misinference. [Disclosure of the Invention] [Means for solving the problem]
[0008] A machine learning image generation method according to one aspect of the present invention involves reducing the components of at least a portion of the frequency bands within the frequency band higher than the Nyquist frequency of the training image for a candidate ground truth image with a resolution higher than the training image paired with the ground truth image, thereby generating a ground truth image for machine learning that improves the resolution of the input image.
[0009] A machine learning method according to one aspect of the present invention involves using a ground truth image generated by performing a reduction process on a candidate ground truth image with a resolution higher than the training image pair, applying a reduction process to components in at least some frequency bands within the frequency band higher than the Nyquist frequency of the training image, or using a ground truth image captured by an imaging device equipped with an optical low-pass filter that reduces components in at least some frequency bands within the frequency band higher than the Nyquist frequency of the training image, and the training image. , perform inference to improve the resolution of the input image and generate an output image Train a machine learning model.
[0010] An endoscopic image processing apparatus according to one aspect of the present invention includes a machine learning model connection unit that can be connected to a machine learning model trained by a machine learning method that trains a machine learning model using a ground truth image candidate with a higher resolution than the training image pair, the ground truth image generated by reducing components in at least some frequency bands within the frequency band higher than the Nyquist frequency of the training image, or the ground truth image captured by an imaging device equipped with an optical low-pass filter that reduces components in at least some frequency bands within the frequency band higher than the Nyquist frequency of the training image, and the training image, and a processor, wherein the processor processes the received endoscopic image , perform inference to improve the resolution of the input image and generate an output image The image is input to the machine learning model, and the machine learning model outputs the endoscopic image with improved resolution. [Brief explanation of the drawing]
[0011] [Figure 1] This figure shows the general flow when a machine learning model, according to each embodiment of the present invention, becomes a trained model after learning to improve image resolution and then performs inference. [Figure 2] This figure shows an example in which a machine learning model is constructed using a neural network according to each embodiment. [Figure 3] This is a diagram showing examples of frequency components of training images and ground truth images for each embodiment. [Figure 4] This figure shows an example of how a machine learning model learns using ground truth images generated from candidate ground truth images, according to each embodiment. [Figure 5] These are diagrams showing first to fourth examples of the frequency characteristics of the ground truth images generated from the ground truth image candidates according to each embodiment. [Figure 6] This graph shows a fifth example of the frequency characteristics of the ground truth image generated from the ground truth image candidate according to each embodiment. [Figure 7] This graph shows a sixth example of the frequency characteristics of the ground truth image generated from the ground truth image candidate according to each embodiment. [Figure 8]This block diagram shows an example configuration of a machine learning image generation device that generates a correct image by image processing, according to the first embodiment of the present invention. [Figure 9] This block diagram shows an example configuration of a machine learning image generation device that generates a ground truth image using an optical low-pass filter, according to a second embodiment of the present invention. [Figure 10] This is a block diagram showing a first configuration example of a machine learning device in a third embodiment of the present invention. [Figure 11] This is a block diagram showing a second configuration example of a machine learning device in a fourth embodiment of the present invention. [Figure 12] This is a block diagram showing a third configuration example of a machine learning device in a fifth embodiment of the present invention. [Figure 13] This is a block diagram showing a fourth configuration example of a machine learning device in the sixth embodiment of the present invention. [Figure 14] This is a block diagram showing a fifth configuration example of a machine learning device in a seventh embodiment of the present invention. [Figure 15] This block diagram shows an example configuration of a machine learning device that acquires original images from an endoscope system, according to the eighth embodiment of the present invention. [Figure 16] This block diagram shows an example configuration in which a trained model learned by a machine learning device is applied to an endoscope system in the ninth embodiment of the present invention. [Best Mode for Carrying Out the Invention]
[0012] Embodiments of the present invention will be described below with reference to the drawings. However, the present invention is not limited to the embodiments described below.
[0013] In the drawings, the same or corresponding elements are appropriately labeled with the same reference numerals. Also, the drawings are schematic, and it should be noted that the relationships between the lengths of the respective elements, the ratios of the lengths of the respective elements, the quantities of the respective elements, etc. within one drawing may differ from reality in order to simplify the description. Furthermore, there may be parts where the relationships and ratios of the lengths between the plurality of drawings are different from each other.
[0014] FIGS. 1 to 16 show embodiments of the present invention. FIG. 1 is a diagram showing a general flow when a machine learning model M performs learning to enhance the spatial resolution of an image by a machine learning method and becomes a learned model M, and then performs inference, according to each embodiment. In the specification and drawings, for the sake of simplicity, the "machine learning model" may sometimes be simply referred to as the "model".
[0015] The machine learning model M is a mathematical model such as a neural network (NN: Neural Network), for example. FIG. 2 is a diagram showing an example in which the machine learning model M is configured by a neural network, according to each embodiment.
[0016] The neural network includes an input layer IL, a hidden layer HL, and an output layer OL. Each layer includes a plurality of neurons (units / nodes) ne, and each neuron ne in a certain layer is connected to each neuron ne in the next layer by an edge (line) ed, forming a network structure.
[0017] As the machine learning model M, a deep neural network (DNN) with multiple hidden layers HL and a total depth of 4 or more layers may be used. Also, as the machine learning model M, a convolutional neural network (CNN), an R-CNN (Regions with CNN features) or an FCN (Fully Convolutional Networks) using a CNN, etc. may be used. Note that machine learning is not limited to deep learning, and other known various learning methods may also be used.
[0018] The machine learning model M performs machine learning using learning data LDS in a learning process LP to which a machine learning method is applied. The learning data LDS is a learning image dataset including pairs of training images IT and correct answer images IC. The learning data LDS including the correct answer image IC is also called teacher data. The machine learning model M inputs the training image IT and propagates the data of the input training image IT forward through the input layer IL → hidden layer HL → output layer OL to generate an output image.
[0019] Then, the machine learning model M calculates the difference (loss value) between the output image and the correct answer image IC using a loss function, and in backpropagation, adjusts the parameters to minimize the loss value as much as possible using an optimization algorithm to perform learning.
[0020] The machine learning model M that has been learned through the learning process LP performs inference on the input image I1 in the inference process IP and outputs the post-inference image I2 as the output image. The learned machine learning model M is composed of, for example, a combination of an AI program (algorithm) and parameters optimized by learning.
[0021] FIG. 3 is a chart showing examples of the frequency components of the training image IT and the correct answer image IC according to each embodiment. Note that in this specification, the frequency refers to the spatial frequency.
[0022] Column A of Figure 3 shows an example graph of the frequency components of the training image IT and the ground truth image IC when a ground truth image IC with a higher resolution than the training image IT is used. When the ground truth image IC has a higher resolution than the training image IT, the ground truth image IC generally contains components (high-frequency components) in a frequency band higher than the frequency band of the training image IT (greater than f1 and less than or equal to f2), as shown in the area enclosed by the dashed line in Column A of Figure 3. Here, f1 is the upper limit frequency at which the training image IT has components, and f2 is the upper limit frequency at which the ground truth image IC has components.
[0023] If we use a combination of training image IT and ground truth image IC as shown in column A of Figure 3, the machine learning model M will learn to generate output images with a frequency band between 0 and f2. However, high-frequency components in the frequency band greater than f1 and less than or equal to f2 are information that is not originally present in the training image IT, so the inference results by the trained model cannot be guaranteed. As a result, the trained model may make incorrect inferences and generate false patterns, for example, in medical images, it may generate images that appear to have blood vessels in locations where there are actually no blood vessels.
[0024] Column B of Figure 3 shows an example graph of the frequency components of the training image IT, the candidate ground truth image ICC, and the ground truth image IC when the ground truth image IC is generated from a candidate ground truth image ICC with a higher resolution than the training image IT.
[0025] The ground truth image candidate ICC shown in column B of Figure 3 has the same frequency characteristics as the ground truth image IC shown in column A of Figure 3. The ground truth image IC is generated by reducing the components at frequencies higher than f1 of the ground truth image candidate ICC without reducing the components below f1 of the ground truth image candidate ICC. Here, "reduction" means making the component amount of the ground truth image IC lower than the component amount of the ground truth image candidate ICC at the same frequency, and includes cases where the component amount is set to 0.
[0026] Figure 4 shows an example of how a machine learning model M learns using the ground truth image IC generated from the ground truth image candidate ICC, according to each embodiment.
[0027] In the learning process LP, the machine learning model M uses the training image IT and the ground truth image IC, which is generated by reducing the components at frequencies higher than frequency f1 from the candidate ground truth image ICC, for training. This reduces misinferences in frequency bands that are not originally present in the training image IT when inference is performed by the trained model M. For example, it can suppress misinferences such as generating an image that appears to have blood vessels where there are none.
[0028] Figure 5 is a diagram showing the first to fourth examples of the frequency characteristics of the ground truth image IC generated from the ground truth image candidate ICC for each embodiment.
[0029] In relation to each column in Figure 5, the training image IT consists of multiple pixels arranged with a first pixel pitch, where the first pixel pitch is the sampling distance. The sampling frequency of the training image IT is the reciprocal of the sampling distance. The Nyquist frequency ftn of the training image IT is half the sampling frequency. The training image IT has components in the frequency band below the Nyquist frequency ftn, and no components in the frequency band above the Nyquist frequency ftn (referred to as the high-frequency band as appropriate).
[0030] Furthermore, the candidate ground truth image ICC is composed of multiple pixels arranged with a second pixel pitch smaller than the first pixel pitch, where the second pixel pitch is the sampling distance. Therefore, the candidate ground truth image ICC generally has components in a frequency band higher than the Nyquist frequency ftn of the training image IT.
[0031] Furthermore, in each column of Figure 5, the ground truth image ICs do not have any changes in the amount of components in the frequency band below the Nyquist frequency ftn compared to the ground truth image candidate ICC.
[0032] The ground truth image IC in the first example shown in column A of Figure 5 exhibits a frequency response in which the component amount of the ground truth image candidate ICC is rapidly and smoothly reduced when the frequency exceeds the Nyquist frequency ftn (becomes higher than the Nyquist frequency ftn), resulting in a component amount of zero at a certain frequency higher than the Nyquist frequency ftn.
[0033] The ground truth image IC in the second example shown in column B of Figure 5 has a frequency response in which the component amount of the ground truth image candidate ICC in the high-frequency band is generally reduced. Therefore, although the ground truth image IC has components in the high-frequency band, the component amount of the ground truth image IC in the high-frequency band is smaller than the component amount of the ground truth image candidate ICC.
[0034] In the third example shown in column C of Figure 5, the ground truth image IC discontinuously reduces the component amount of the ground truth image candidate ICC to a value close to zero when the frequency exceeds the Nyquist frequency ftn (becomes higher than the Nyquist frequency ftn). The ground truth image IC has a frequency response with a component amount close to zero in the high-frequency band.
[0035] The correct image IC for the fourth example shown in column D of Figure 5 discontinuously sets the component amount to zero when the frequency exceeds the Nyquist frequency ftn (becomes higher than the Nyquist frequency ftn), making all component amounts in the high-frequency band zero.
[0036] Incidentally, image resolution can be divided into relative resolution, which changes depending on the size of the paper on which the image is displayed, or the size of the display screen (i.e., depending on the number of pixels per unit length), and absolute resolution, which is determined by the number of pixels that make up the image. In this specification, resolution refers to absolute resolution.
[0037] To give one numerical example, suppose the candidate ground truth image (ICC) consists of 1280 pixels wide by 720 pixels high, and the training image (IT) consists of 640 pixels wide by 360 pixels high.
[0038] The ground truth image ICs shown in columns A to C of Figure 5 have components in a frequency band higher than the Nyquist frequency ftn. Therefore, in the numerical example above, the ground truth image IC is composed of 1280 pixels horizontally and 720 pixels vertically, similar to the ground truth image candidate ICC.
[0039] On the other hand, the ground truth image IC shown in column D of Figure 5 does not have components in a frequency band higher than the Nyquist frequency ftn. Therefore, in the above numerical example, the ground truth image IC may be composed of 1280 x 720 pixels, similar to the ground truth image candidate ICC, or it may be composed of 640 x 360 pixels, similar to the training image IT. In the latter case, the ground truth image IC can be generated by reducing the size of the ground truth image candidate ICC.
[0040] Thus, the ground truth image IC does not need to have a higher resolution than the training image IT, unlike the ground truth image candidate ICC; it can have the same resolution as the training image IT.
[0041] Note that the ground truth image IC shown in column A of Figure 5 has components up to a frequency slightly higher than the Nyquist frequency ftn, but its high-frequency bandwidth is narrower than that of the ground truth image candidate ICC. Therefore, in the above numerical example, the ground truth image IC may have a pixel configuration intermediate between 640 x 360 pixels and 1280 x 720 pixels. In this case as well, the ground truth image IC can be generated by reducing the size of the ground truth image candidate ICC.
[0042] The relationship between the resolution of the ground truth image IC and the training image IT is as described above, but the relationship between their spatial resolution is as follows.
[0043] When comparing the component amounts of the ground truth image IC and the training image IT in the frequency band below the Nyquist frequency ftn, the ground truth image IC has a larger component amount than the training image IT. In particular, at higher frequencies in the frequency band below the Nyquist frequency ftn, the component amount of the ground truth image IC is larger than that of the training image IT.
[0044] Therefore, even if the resolution of the ground truth image IC and the training image IT are the same, the spatial resolution of the ground truth image IC is higher than that of the training image IT. Generally, higher image resolution leads to higher image contrast. Consequently, the contrast of the ground truth image IC is higher than that of the training image IT.
[0045] Figure 6 is a graph showing a fifth example of the frequency characteristics of the ground truth image IC generated from the ground truth image candidate ICC according to each embodiment.
[0046] In each example shown in Figure 5, the ground truth image IC showed reduced component levels compared to the ground truth image candidate ICC in frequency bands higher than the Nyquist frequency ftn, but no reduction in component levels in frequency bands below the Nyquist frequency ftn. In contrast, in the fifth example shown in Figure 6, component levels were reduced in some frequency bands even below the Nyquist frequency ftn.
[0047] The training image IT shown in Figure 6 lacks components in some of the higher frequency bands within the frequency range below the Nyquist frequency ftn and above half the Nyquist frequency ftn. In this case, errors may occur in inference by the trained model even in the frequency band below the Nyquist frequency ftn where the training image IT lacks information.
[0048] Therefore, in the fifth example shown in Figure 6, processing is performed to reduce components in at least some of the frequency bands within the frequency range of the ground truth image candidate ICC below the Nyquist frequency fcn (indicated by the white double-headed arrows in Figure 6), at a frequency range of 1 / 2 or more of the Nyquist frequency ftn. In this way, components in the frequency band below the Nyquist frequency ftn may also be reduced or removed as needed.
[0049] Specifically, in the fifth example shown in Figure 6, the training image IT has components in the frequency band below frequency f3, indicated by the dashed line, and no components in the frequency band above frequency f3. Frequency f3 is above half of the Nyquist frequency ftn (ftn / 2) and lower than the Nyquist frequency ftn.
[0050] Therefore, the ground truth image IC is generated by reducing the components of the ground truth image candidate ICC in the frequency band higher than frequency f3 and below the Nyquist frequency fcn of the ground truth image candidate ICC. In the fifth example shown in Figure 6, following the second example shown in column B of Figure 5, the ground truth image IC is an example of a frequency characteristic in which the amount of components of the ground truth image candidate ICC is generally reduced in the frequency band higher than frequency f3.
[0051] Figure 7 is a graph showing a sixth example of the frequency characteristics of the ground truth image IC generated from the ground truth image candidate ICC according to each embodiment.
[0052] The sixth example shown in Figure 7 is an example in which the frequency characteristics of the ground truth image IC are different, while the frequency characteristics of the training image IT and the ground truth image candidate ICC are the same as in the fifth example shown in Figure 6.
[0053] In other words, the correct image IC for the sixth example shown in Figure 7, in accordance with the fourth example shown in column D of Figure 5, discontinuously sets the component amount to 0 when the frequency exceeds f3 (becomes higher than f3), and sets all component amounts in the frequency band higher than f3 to 0.
[0054] As explained with reference to Figures 6 and 7, the ground truth image IC may be generated by reducing the amount of components in the candidate ground truth image ICC in frequency bands higher than the upper limit frequency f3 of the frequency band in which the training image IT has components. The upper limit frequency f3 can be obtained, for example, by frequency analysis of the training image IT.
[0055] Furthermore, for example, when (1 / 2) ≤ α < 1, frequency analysis can be omitted, and the amount of components in the ground truth image candidate ICC in the frequency band greater than or equal to α × ftn can be reduced to generate the ground truth image IC.
[0056] In this case, if α = (1 / 2), the amount of components in the ground truth image candidate ICC in the frequency band of half the Nyquist frequency ftn (ftn / 2) or more will be reduced, and the ground truth image IC will be generated. Setting α = (1 / 2) can prevent misinference in many practical cases. [First Embodiment]
[0057] Figure 8 is a block diagram showing an example configuration of a machine learning image generation device 10 that generates a ground truth image IC by image processing in the first embodiment. 8 The machine learning image generation device 10 shown generates a ground truth image IC from a ground truth image candidate ICC using an image processing-based machine learning image generation method.
[0058] The machine learning image generation device 10 includes a frequency component adjustment unit 11. The frequency component adjustment unit 11 receives information on the Nyquist frequency ftn of the training image IT and a ground truth image candidate ICC with higher resolution than the training image IT. The frequency component adjustment unit 11 performs a reduction process on the ground truth image candidate ICC to reduce components in at least some of the frequency bands within the frequency band higher than the Nyquist frequency ftn, thereby generating a ground truth image IC for use in machine learning to improve the image resolution.
[0059] The reduction process performed by the frequency component adjustment unit 11 may be a process that reduces all frequency band components higher than the Nyquist frequency ftn, as shown in columns B and C of Figure 5.
[0060] Furthermore, the reduction process performed by the frequency component adjustment unit 11 may be a process that reduces all components in frequency bands higher than the Nyquist frequency ftn to zero (reduces them to zero), as shown in column D of Figure 5. In other words, the reduction process includes a removal process that reduces frequency components to zero.
[0061] The frequency component adjustment unit 11 may reduce components in all frequency bands higher than the Nyquist frequency ftn by reducing the ground truth image candidate ICC. Here, the reduction process is a process that changes the pixel configuration of the ground truth image candidate ICC to reduce the number of pixels. The ground truth image IC generated by reducing the ground truth image candidate ICC will have a resolution lower than that of the ground truth image candidate ICC, but higher than that of the training image IT.
[0062] In the numerical example described above, where the candidate ground truth image ICC consists of 1280 pixels horizontally and 720 pixels vertically, and the training image IT consists of 640 pixels horizontally and 360 pixels vertically, the reduction process is the process of changing the ground truth image IC to a pixel configuration of 640 pixels or more but less than 1280 pixels horizontally, and 360 pixels or more but less than 720 pixels vertically.
[0063] In the reduction process, it is preferable not to reduce the components in the frequency band below the Nyquist frequency ftn of the ground truth image candidate ICC. This ensures that the resolution of the ground truth image IC does not decrease below the resolution of the ground truth image candidate ICC in the frequency band below the Nyquist frequency ftn.
[0064] Not only can the resolution of the ground truth image IC not be reduced to the resolution of the ground truth image candidate ICC, but further processing may be performed to improve the resolution of the ground truth image IC to that of the ground truth image candidate ICC in the frequency band below the Nyquist frequency ftn.
[0065] The reduction process performed by the frequency component adjustment unit 11 may be a process that reduces the ground truth image candidate ICC so that it has the same resolution as the training image IT.
[0066] As described above, the reduction process by the frequency component adjustment unit 11 is not limited to frequency bands higher than the Nyquist frequency ftn. That is, the reduction process by the frequency component adjustment unit 11 may further include a process to reduce components in at least some frequency bands within the frequency band that is half or more of the Nyquist frequency ftn and less than or equal to the Nyquist frequency ftn.
[0067] Alternatively, the reduction process performed by the frequency component adjustment unit 11 may be carried out by applying a low-pass filter process as an image processing step to the correct image candidate ICC.
[0068] Alternatively, the reduction process by the frequency component adjustment unit 11 may be performed by both low-pass filtering and reduction processing. For example, the reduction process by the frequency component adjustment unit 11 may be performed by applying low-pass filtering to the ground truth image candidate ICC and then reducing the ground truth image candidate ICC after applying the low-pass filtering.
[0069] According to the first embodiment, the frequency component adjustment unit 11 reduces components in at least a portion of the frequency band higher than the Nyquist frequency ftn for the candidate ground truth image ICC to generate the ground truth image IC. Therefore, the possibility of a machine learning model trained using the training image IT and the ground truth image IC generating an incorrectly inferred output image can be reduced.
[0070] In particular, by setting all frequency bands higher than the Nyquist frequency (ftn) to zero, it is possible to significantly reduce misinferences such as the generation of false patterns.
[0071] Furthermore, when reduction processing is performed using at least one of reduction processing and low-pass filtering processing, the processing load can be significantly reduced and processing speed can be improved compared to, for example, performing Fourier analysis to reduce components in frequency bands higher than a specific frequency (or above a specific frequency) and then performing an inverse Fourier transform to restore them.
[0072] Furthermore, by ensuring that the resolution of the ground truth image IC does not decrease below the resolution of the candidate ground truth image ICC in the frequency band below the Nyquist frequency ftn, the machine learning model can be trained to generate inferred images with higher resolution. [Second Embodiment]
[0073] Figure 9 is a block diagram showing an example configuration of a machine learning image generation device 10 that generates a ground truth image IC using an optical low-pass filter 17 in the second embodiment. In the second embodiment, the same reference numerals are used for parts that are the same as in the first embodiment, and their descriptions are omitted as appropriate. The second embodiment will mainly describe the differences from the first embodiment.
[0074] In the first embodiment, a ground truth image IC was generated from a candidate ground truth image ICC by image processing. In contrast, the machine learning image generation device 10 of the second embodiment shown in Figure 9 generates a ground truth image IC using a machine learning image generation method in which the image sensor 18 captures the subject light beam that has passed through the optical low-pass filter 17.
[0075] The imaging device 15 that generates the correct image IC is, for example, an imaging system that includes an imaging optical system 16, an optical low-pass filter 17, and an image sensor 18.
[0076] The imaging optical system 16 focuses the light beam from the subject and forms an optical image of the subject on the imaging surface of the image sensor 18. The imaging optical system 16 generally includes multiple optical lenses and an optical aperture. However, the imaging optical system 16 may have other configurations.
[0077] The optical low-pass filter 17 is positioned in the path of the subject light beam, cutting out high-frequency components of the optical image and allowing low-frequency components to pass through. Generally, the optical low-pass filter 17 is positioned on the imaging plane of the image sensor 18. However, the optical low-pass filter 17 may be positioned at other locations.
[0078] The optical low-pass filter 17 reduces high-frequency components in at least some of the frequency bands in the optical image that are higher than the Nyquist frequency ftn of the training image IT. Specifically, the optical low-pass filter 17 reduces frequency components in the optical image that are greater than or equal to half the Nyquist frequency ftn.
[0079] The image sensor 18 converts the optical image of the subject, which has been imaged by the imaging optical system 16 via the optical low-pass filter 17, into an optical signal. The image associated with the output signal is a ground truth image IC in which components in at least some frequency bands within the frequency band higher than the Nyquist frequency ftn are reduced.
[0080] According to the second embodiment, by capturing an optical image of the subject formed through the optical low-pass filter 17, it is possible to generate a suitable ground truth image IC for training that can reduce inference errors. In this case, image processing, such as that performed on candidate ground truth image ICs to generate a ground truth image IC, is unnecessary, as in the first embodiment.
[0081] Furthermore, the ground truth image IC that the machine learning model M uses for training along with the training image IT may be either the ground truth image IC generated by the configuration of the first embodiment or the ground truth image IC generated by the configuration of the second embodiment. [Third Embodiment]
[0082] Figure 10 is a block diagram showing a first configuration example of the machine learning device 1 in the third embodiment. In the third embodiment, the same reference numerals are used for parts that are the same as in the first and second embodiments, and their descriptions are omitted as appropriate. In the third embodiment, the differences from the first and second embodiments will be mainly described.
[0083] The machine learning device 1 comprises a machine learning image generation device 10 and a model learning processing unit 20. In this embodiment, the machine learning image generation device 10 generates training images IT and ground truth images IC from ground truth image candidate ICC. The model learning processing unit 20 performs machine learning using the generated training images IT and ground truth images IC.
[0084] The machine learning device 1 may be configured such that a processor, such as an ASIC (Application Specific Integrated Circuit) including a CPU (Central Processing Unit) or an FPGA (Field Programmable Gate Array), reads and executes processing programs stored in a memory or other storage device (or recording medium) to perform the functions of each part. Alternatively, the machine learning device 1 may be configured as a dedicated electronic circuit to perform the functions of each part.
[0085] The machine learning image generation device 10 comprises a frequency component adjustment unit 11 and a training image generation unit 12. The frequency component adjustment unit 11 and the training image generation unit 12 receive candidate ground truth image ICCs as input.
[0086] The training image generation unit 12 further inputs training image imaging system information. The training image imaging system information is the imaging system information assumed for the training image IT. In other embodiments described later, imaging system information assumed for the ground truth image IC (ground truth image imaging system information) is used. Whether the assumed system is the training image IT or the ground truth image IC, the imaging system information includes pixel count information, color information, optical characteristics information, noise characteristics information, color filter information, etc.
[0087] Pixel count information is information about the pixel configuration of an image, and corresponds to the pixel configuration of the image sensor of the endoscope to which the machine learning model M is applied. The pixel count information assumed for the training image IT is, in the numerical example above, horizontal. 64 This information consists of 0 pixels in a vertical 360-pixel array.
[0088] Color information is information used to correct the difference between the color of the input image (in this embodiment, the candidate image ICC) and the color of the output image (in this embodiment, the training image IT).
[0089] Optical characteristic information includes, for example, information for correcting the PSF (Point Spread Function) (corrected PSF). The PSF is a point spreading function that indicates the spatial distribution of the image formed by a point light source using an imaging optical system.
[0090] The PSF of the input image (in this embodiment, the candidate image ICC) is generally different from the PSF expected for the output image (in this embodiment, the training image IT). Therefore, the optical characteristic information used to convert the PSF of the input image to the PSF expected for the output image is called the correction PSF.
[0091] Here, the PSF of the imaging optical system generally depends on the position within the image. Therefore, a PSF that depends on the coordinates within the image may be used as the correction PSF. Furthermore, using a uniform correction PSF across the entire screen has the advantage of reducing circuit size and computational cost.
[0092] Examples of PSFs (Photos per Second) assumed for training images (IT) include the actual PSF or a virtual PSF in the imaging optical system of the endoscope to which the machine learning model M is applied. The virtual PSF may have greater blur than the actual PSF. By using a virtual PSF with greater blur than the actual PSF, it is expected that the resolution of the images inferred by the trained model M will be significantly improved. On the other hand, by using a virtual PSF with less blur, it is expected that the resolution of the images inferred by the trained model M will decrease, but the number of misinferences will be reduced.
[0093] Noise characteristic information is information that indicates, for example, how much noise to add to the pixel values at which locations, assuming random noise. It is preferable that the noise characteristic information includes both noise location information and noise amount information, but it may also include at least one of the following: noise location, noise amount, and the standard deviation of the noise to be added. When using noise standard deviation information as noise characteristic information, the noise standard deviation information may be uniform regardless of the pixel value, or it may be variable depending on the pixel value.
[0094] Color filter information includes, for example, whether the endoscope to which the machine learning model M is applied is equipped with an image sensor having a Bayer array color filter, or whether it is equipped with an image sensor that sequentially acquires R, G, and B images by sequentially irradiating the image with RGB illumination light and imaging with a monochrome image sensor. In the following, images acquired by an image sensor having a Bayer array color filter will be called Bayer images, and images acquired by a monochrome image sensor with sequential illumination will be called sequential images.
[0095] The training image generation unit 12 generates a training image IT from the ground truth image candidate ICC based on the training image acquisition system information described above.
[0096] The training image generation unit 12 generates a first intermediate image by reducing a candidate ground truth image ICC from, for example, 1280 pixels wide x 720 pixels high to 640 pixels wide x 360 pixels high, based on the pixel count information.
[0097] The training image generation unit 12 performs color correction processing on the first intermediate image based on the color information to generate a second intermediate image.
[0098] Furthermore, the training image generation unit 12 applies a coordinate-dependent correction PSF to the second intermediate image to perform PSF correction and generates a third intermediate image having blur corresponding to the assumed PSF for the training image IT.
[0099] Furthermore, the training image generation unit 12 generates a fourth intermediate image by adding random noise and the like to the third intermediate image based on the assumed noise characteristics information for the training image IT.
[0100] If the candidate ground truth image ICC is a Bayer image and the endoscope to which the machine learning model M is applied generates a Bayer image, or if the candidate ground truth image ICC is a plane-sequential image and the endoscope to which the machine learning model M is applied generates a plane-sequential image, the training image generation unit 12 outputs a fourth intermediate image as the training image IT.
[0101] On the other hand, if the ground truth image candidate ICC is a planar sequential image and the endoscope to which the machine learning model M is applied generates a Bayer image, the training image generation unit 12 converts the fourth intermediate image, which is a planar sequential image, into a Bayer image and outputs it as the training image IT.
[0102] Furthermore, when the endoscope generates planar sequential images, it is preferable that the ground truth image candidate ICC is a planar sequential image. However, the training image generation unit 12 may demosaicing the fourth intermediate image generated from the ground truth image candidate ICC of the Bayer image to convert it into a planar sequential image and output it as the training image IT.
[0103] In addition, the above describes an example of the processing order in which the training image generation unit 12 generates a training image IT from a candidate correct image ICC based on the information included in the training image acquisition system information. However, the order is not limited to the above, and processing may be performed in other orders.
[0104] The training image generation unit 12 transmits the generated training image IT to the model learning processing unit 20 and transmits the Nyquist frequency ftn information to the frequency component adjustment unit 11.
[0105] As described above, the frequency component adjustment unit 11 receives information on the Nyquist frequency ftn of the training image IT and a ground truth image candidate ICC with a higher resolution than the training image IT. Similar to the first embodiment, the frequency component adjustment unit 11 reduces components in the frequency band higher than the Nyquist frequency ftn (or more than half the Nyquist frequency ftn) from the ground truth image candidate ICC to generate the ground truth image IC.
[0106] As explained with reference to Figure 4, the model learning processing unit 20 uses the training image IT and the ground truth image IC to train the machine learning model M using the learning process LP. The model learning processing unit 20 then outputs the trained model M.
[0107] The third embodiment achieves substantially the same effects as the first embodiment. Furthermore, according to the third embodiment, if a candidate ground truth image ICC is obtained, both the training image IT and the ground truth image IC can be generated.
[0108] Furthermore, according to the third embodiment, the training image generation unit 12 corrects the input image using training image acquisition system information, thereby enabling the generation of training images IT with high consistency with the ground truth image IC. A trained model M, which has been machine-learned using the highly consistent training images IT and ground truth image IC, is expected to generate inference images with more appropriately improved resolution. [Fourth Embodiment]
[0109] Figure 11 is a block diagram showing a second configuration example of the machine learning device 1 in the fourth embodiment. In the fourth embodiment, the same reference numerals are used for parts that are the same as in the first to third embodiments, and their descriptions are omitted as appropriate. In the fourth embodiment, the differences from the first to third embodiments will be mainly described.
[0110] The machine learning device 1 comprises a machine learning image generation device 10 and a model learning processing unit 20. In this embodiment, the machine learning image generation device 10 generates training images IT and ground truth images IC from the source images. The model learning processing unit 20 performs machine learning using the generated training images IT and ground truth images IC.
[0111] The machine learning image generation device 10 comprises a frequency component adjustment unit 11, a training image generation unit 12, and a ground truth image candidate generation unit 13. The training image generation unit 12 and the ground truth image candidate generation unit 13 receive the original image as input.
[0112] The training image generation unit 12 further inputs the training image acquisition system information as described above. Based on the training image acquisition system information, the training image generation unit 12 generates a training image IT from the original image. The training image generation unit 12 transmits the generated training image IT to the model learning processing unit 20 and transmits the Nyquist frequency ftn information to the frequency component adjustment unit 11.
[0113] The ground truth image candidate generation unit 13 further receives ground truth image imaging system information. As described above, the ground truth image imaging system information is the imaging system information assumed for the ground truth image IC. The values of each piece of information included in the ground truth image imaging system information are generally different from the values of each piece of information included in the training image imaging system information. Based on the ground truth image imaging system information, the ground truth image candidate generation unit 13 generates a ground truth image candidate ICC from the original image using the same procedure as the training image generation unit 12. The ground truth image candidate generation unit 13 transmits the generated ground truth image candidate ICC to the frequency component adjustment unit 11.
[0114] As described above, the frequency component adjustment unit 11 generates a ground truth image IC from the ground truth image candidate ICC based on the Nyquist frequency ftn.
[0115] As described above, the model learning processing unit 20 uses the training image IT and the ground truth image IC to train the machine learning model M and outputs the trained model M.
[0116] The fourth embodiment achieves substantially the same effects as the first and third embodiments. Furthermore, by using ground truth image acquisition system information in addition to training image acquisition system information, both the training image IT and the ground truth image IC can be generated from a single source image. [Fifth Embodiment]
[0117] Figure 12 is a block diagram showing a third configuration example of the machine learning device 1 in the fifth embodiment. In the fifth embodiment, the same reference numerals are used for parts that are the same as in the first to fourth embodiments, and their descriptions are omitted as appropriate. The fifth embodiment will mainly describe the differences from the first to fourth embodiments.
[0118] The machine learning image generation device 10 of this embodiment generates training images IT and ground truth images IC from the source images, similar to the fourth embodiment. However, the processing order of the machine learning image generation device 10 of this embodiment differs from that of the fourth embodiment.
[0119] The machine learning device 1 comprises a machine learning image generation device 10 and a model learning processing unit 20. The machine learning image generation device 10 comprises a frequency component adjustment unit 11, a training image generation unit 12, and a ground truth image generation unit 14. The frequency component adjustment unit 11 and the training image generation unit 12 receive the original image as input.
[0120] The training image generation unit 12 further inputs training image acquisition system information. Based on the training image acquisition system information, the training image generation unit 12 generates a training image IT from the original image. The training image generation unit 12 transmits the generated training image IT to the model learning processing unit 20 and transmits the Nyquist frequency ftn information to the frequency component adjustment unit 11.
[0121] The frequency component adjustment unit 11 receives the original image as a candidate ground truth image ICC. The frequency component adjustment unit 11 reduces components in frequency bands higher than the Nyquist frequency ftn (or more than half of the Nyquist frequency ftn) from the candidate ground truth image ICC, similar to the frequency component adjustment unit 11 described in the first embodiment, and generates an adjusted candidate ground truth image ICC. The frequency component adjustment unit 11 transmits the adjusted candidate ground truth image ICC to the ground truth image generation unit 14.
[0122] The ground truth image generation unit 14 receives the adjusted ground truth image candidate ICC and also receives ground truth image acquisition system information. Based on the ground truth image acquisition system information, the ground truth image generation unit 14 generates a ground truth image IC from the adjusted ground truth image candidate ICC using the same procedure as the ground truth image candidate generation unit 13. The ground truth image generation unit 14 transmits the generated ground truth image IC to the model learning processing unit 20.
[0123] As described above, the model learning processing unit 20 uses the training image IT and the ground truth image IC to train the machine learning model M and outputs the trained model M.
[0124] In the fourth embodiment, the original image was subjected to image processing based on the ground truth image acquisition system information, and then the reduction process was performed by the frequency component adjustment unit 11. In contrast, in the fifth embodiment, the original image was subjected to reduction processing by the frequency component adjustment unit 11, and then the ground truth image IC was generated by image processing based on the ground truth image acquisition system information. The fifth embodiment achieves substantially the same effects as the fourth embodiment. [Sixth Embodiment]
[0125] Figure 13 is a block diagram showing a fourth configuration example of the machine learning device 1 in the sixth embodiment. In the sixth embodiment, the same reference numerals are used for parts that are the same as in the first to fifth embodiments, and their descriptions are omitted as appropriate. The sixth embodiment will mainly describe the differences from the first to fifth embodiments.
[0126] The machine learning device 1 of this embodiment has the same components as the fourth embodiment, but the information transmitted is different.
[0127] The machine learning device 1 comprises a machine learning image generation device 10 and a model learning processing unit 20. The machine learning image generation device 10 comprises a frequency component adjustment unit 11, a training image generation unit 12, and a ground truth image candidate generation unit 13.
[0128] The training image generation unit 12 generates a training image IT from the original image, transmits the training image IT to the model learning processing unit 20, and transmits the Nyquist frequency ftn information to the frequency component adjustment unit 11, similar to the fourth embodiment. In this embodiment, the training image generation unit 12 further transmits generation-additional information to the ground truth image candidate generation unit 13.
[0129] The generation-added information includes information indicating the location and amount of noise added to the training image IT by the training image generation unit 12, and / or standard deviation information of the noise included in the training image acquisition system information.
[0130] The ground truth image candidate generation unit 13 processes the original image based on the ground truth image acquisition system information. In this case, the ground truth image acquisition system information in the sixth embodiment does not include noise characteristic information. Therefore, the ground truth image candidate generation unit 13 adds noise to the processed original image based on the generation-added information to generate a ground truth image candidate ICC. Furthermore, if there are multiple position candidates in the ground truth image candidate ICC that correspond to the positions of the noise added to the training image IT, the same noise that was added to the training image IT may be added to the pixel values of each of the multiple position candidates. Alternatively, the same noise that was added to the training image IT may be added to the pixel values of some of the position candidates, and random noise newly generated from the noise standard deviation information may be added to the pixel values of the other position candidates.
[0131] Subsequently, the ground truth image candidate generation unit 13 transmits the generated ground truth image candidate ICC to the frequency component adjustment unit 11.
[0132] As described above, the frequency component adjustment unit 11 generates a ground truth image IC from the ground truth image candidate ICC based on the Nyquist frequency ftn.
[0133] In the fourth embodiment described above, separate noise was added to the training image IT and the ground truth image IC. In contrast, in the sixth embodiment, the noise added to the training image IT and the ground truth image IC is matched in terms of position and noise amount at corresponding locations.
[0134] According to the sixth embodiment, the same effects as the fourth embodiment are achieved. Furthermore, according to the sixth embodiment, since the location and amount of noise added to the training image IT and the ground truth image IC are matched, the machine learning model M can learn to infer the ground truth image IC from the training image IT prioritizing resolution improvement. [Seventh Embodiment]
[0135] Figure 14 is a block diagram showing a fifth configuration example of the machine learning device 1 in the seventh embodiment. In the seventh embodiment, the same reference numerals are used for parts that are the same as in the first to sixth embodiments, and their descriptions are omitted as appropriate. The seventh embodiment will mainly describe the differences from the first to sixth embodiments.
[0136] The machine learning device 1 of this embodiment has the same components as the fifth embodiment, but the information transmitted is different.
[0137] The machine learning device 1 comprises a machine learning image generation device 10 and a model learning processing unit 20. The machine learning image generation device 10 comprises a frequency component adjustment unit 11, a training image generation unit 12, and a ground truth image generation unit 14.
[0138] The training image generation unit 12 generates a training image IT from the original image, transmits the training image IT to the model learning processing unit 20, and transmits the Nyquist frequency ftn information to the frequency component adjustment unit 11, similar to the fifth embodiment. In this embodiment, the training image generation unit 12 further transmits the above-mentioned generation-added information to the ground truth image generation unit 14.
[0139] The ground truth image generation unit 14 processes the adjusted ground truth image candidates based on the ground truth image acquisition system information. In this case, the ground truth image acquisition system information in the seventh embodiment does not include noise characteristic information. Therefore, the ground truth image generation unit 14 adds noise to the processed adjusted ground truth image candidates based on the generation-added information to generate a ground truth image IC.
[0140] In the fifth embodiment described above, separate noise was added to the training image IT and the ground truth image IC, but in the seventh embodiment, the noise added to the training image IT and the ground truth image IC is matched in terms of position and noise amount at corresponding locations.
[0141] Subsequently, the ground truth image generation unit 14 transmits the generated ground truth image IC to the model learning processing unit 20.
[0142] According to the seventh embodiment, the effects are substantially the same as those of the fifth embodiment. Furthermore, according to the seventh embodiment, since the location and amount of noise added to the training image IT and the ground truth image IC are matched, the machine learning model M can learn to infer the ground truth image IC from the training image IT prioritizing resolution improvement. [Eighth Embodiment]
[0143] Figure 15 is a block diagram showing an example configuration of a machine learning device 1 that acquires the original image from the endoscope system 31 in the eighth embodiment. In the eighth embodiment, the same reference numerals are used for parts that are the same as in the first to seventh embodiments, and their descriptions are omitted as appropriate. The eighth embodiment will mainly describe the differences from the first to seventh embodiments.
[0144] The machine learning device 1 of this embodiment acquires the original image from the endoscope system 31.
[0145] The machine learning device 1 comprises a machine learning image generation device 10 and a model learning processing unit 20. The machine learning image generation device 10 comprises a frequency component adjustment unit 11, a training image generation unit 12, and a ground truth image candidate generation unit 13.
[0146] The process of generating the trained model M using the machine learning device 1 is basically the same as that described with reference to Figure 11. However, the machine learning device 1 further acquires original image acquisition system information and original image light source information from the endoscope system 31 and processes them.
[0147] Figure 15 shows the configuration of the machine learning device 1 in Figure 11 as an example, but the configuration of the machine learning device 1 shown in Figures 12 to 14 may also be applied to this embodiment. Furthermore, if the images generated by the endoscope system 31 are used as the ground truth image candidate ICC, the configuration of the machine learning device 1 shown in Figure 10 may also be applied to this embodiment.
[0148] The endoscope system 31 comprises a light source unit 32, an imaging device 33, and a memory 34. The light source unit 32 illuminates the object 90, which is the subject, with illumination light. The subject light beam, which is the reflected light from the object 90, is incident on the imaging device 33.
[0149] The imaging device 33 is, in general terms, the same configuration as the imaging device 15 shown in Figure 9, but without the optical low-pass filter 17. The imaging device 33 forms an image of the subject light beam with the imaging optical system 16, captures the image with the image sensor 18, and outputs the original image. The original image is input to the training image generation unit 12 and the ground truth image candidate generation unit 13, as explained with reference to Figure 11.
[0150] Memory 34 is a storage medium that non-volatilely stores information related to the endoscope system 31. The information stored in memory 34 includes original image acquisition system information and original image light source information. Original image acquisition system information is imaging system information related to the imaging device 33 that acquires the original image.
[0151] The original image acquisition system information includes pixel count information of the image sensor 18, color information of the image acquired by the imaging device 33, optical characteristic information including PSF of the imaging optical system 16, noise characteristic information related to the image sensor 18 and the readout circuit from the image sensor 18, and color filter information indicating whether the image sensor 18 acquires a Bayer image or a plane sequential image.
[0152] The light source unit 32 is configured to emit illumination light corresponding to multiple types of observation modes, for example. Observation modes include, for example, white light imaging (WLI) mode and narrow band imaging (NBI) mode. The original image light source information is information indicating the type of illumination light (WLI illumination light, NBI illumination light, etc.) emitted by the light source unit 32 according to the observation mode.
[0153] The endoscope system 31 transmits the original image acquisition system information and the original image light source information corresponding to the observation mode from the memory 34 to the machine learning device 1.
[0154] The training image generation unit 12 and the ground truth image candidate generation unit 13 of the machine learning device 1 receive original image acquisition system information and original image light source information from the endoscope system 31.
[0155] The training image generation unit 12 modifies, for example, the correction PSF and noise characteristic information included in the training image acquisition system information, according to the type of illumination light (WLI illumination light, NBI illumination light, etc.) based on the original image light source information. At this time, the training image generation unit 12 may further modify the correction PSF and noise characteristic information included in the training image acquisition system information based on the optical characteristic information and noise characteristic information included in the original image acquisition system information.
[0156] Similarly, the ground truth image candidate generation unit 13 modifies, for example, the correction PSF and noise characteristic information included in the ground truth image acquisition system information, based on the original image light source information and according to the type of illumination light (WLI illumination light, NBI illumination light, etc.). At this time, the ground truth image candidate generation unit 13 may further modify the correction PSF and noise characteristic information included in the ground truth image acquisition system information based on the optical characteristic information and noise characteristic information included in the original image acquisition system information.
[0157] Furthermore, when applying the trained model M to the same model (or even the same individual) as the endoscope system 31 that acquired the original image, it is acceptable to perform only noise addition without performing PSF correction.
[0158] Furthermore, the training image generation unit 12 and the ground truth image candidate generation unit 13 may, if necessary, modify the reduction process based on the pixel count information of the image sensor 18 included in the original image acquisition system information.
[0159] Furthermore, the training image generation unit 12 and the ground truth image candidate generation unit 13 may, if necessary, modify the color correction process based on the color information included in the original image acquisition system information.
[0160] The training image generation unit 12 and the ground truth image candidate generation unit 13 may, if necessary, change the conversion process from, for example, Bayer image to plane sequential image, or from plane sequential image to Bayer image, based on the color filter information included in the original image acquisition system information.
[0161] According to the eighth embodiment, the original image can be acquired from the endoscope system 31, achieving substantially the same effects as in the third to seventh embodiments. Furthermore, according to the eighth embodiment, the training image acquisition system information and the ground truth image acquisition system information can be appropriately modified based on at least one of the original image acquisition system information and the original image light source information to generate an appropriate training image dataset. [Ninth Embodiment]
[0162] Figure 16 is a block diagram showing an example configuration in which a trained model M, trained on the machine learning device 1, is applied to the endoscope system 41 in the ninth embodiment. In the ninth embodiment, the same reference numerals are used for parts that are the same as in the first to eighth embodiments, and their descriptions are omitted as appropriate. In the ninth embodiment, the differences from the first to eighth embodiments will be mainly described.
[0163] The machine learning device 1 comprises a machine learning image generation device 10 and a model learning processing unit 20. The machine learning image generation device 10 comprises a frequency component adjustment unit 11, a training image generation unit 12, and a ground truth image candidate generation unit 13.
[0164] The process of generating the trained model M using machine learning device 1 is similar to that described with reference to Figure 15. In other words, in addition to the process described with reference to Figure 11, the original image light source information is acquired and processed.
[0165] Figure 16 shows the configuration of the machine learning device 1 in Figure 11 as an example, but the configuration of the machine learning device 1 shown in Figures 12 to 14 may also be applied to this embodiment. Furthermore, if the training image IT and the ground truth image IC are generated based on the ground truth image candidate ICC instead of the original image, the configuration of the machine learning device 1 shown in Figure 10 may also be applied to this embodiment.
[0166] The endoscopic system 41 comprises an endoscope 42 and an endoscopic image processing device 44. The imaging device 43 is, for example, mounted on the endoscope 42. However, the imaging device 43 may be mounted on a camera head, and the camera head may be attached to the eyepiece of the endoscope 42. The imaging device 43 is connected to the endoscopic image processing device 44.
[0167] The imaging device 43, like the imaging device 33 shown in Figure 15, is basically the same as the imaging device 15 shown in Figure 9, but without the optical low-pass filter 17.
[0168] The imaging device 43 forms an image of the subject light beam with the imaging optical system 16, captures the image with the image sensor 18, and outputs an endoscopic image. The endoscopic image output from the imaging device 43 becomes the input image I1 (see Figure 1) to the endoscopic image processing device 44.
[0169] The endoscopic image processing device 44 includes, for example, a processor 44a and a memory 44b. The processor 44a is composed of an ASIC (Application Specific Integrated Circuit) including a CPU (Central Processing Unit), an FPGA (Field Programmable Gate Array), etc. However, the endoscopic image processing device 44 may also be configured as a dedicated electronic circuit that performs the functions of a trained model M.
[0170] Memory 44b is a storage medium that stores (non-volatilely stores) processing programs that realize the functions of each circuit. Processor 44a is connected to the wiring derived from memory 44b. The functions of the endoscopic image processing device 44 are realized when processor 44a reads and executes the processing programs stored in memory 44b. For example, the endoscopic image processing device 44 realizes an endoscopic image processing method by executing a processing program.
[0171] The trained model M generated by the machine learning device 1, that is, the combination of the AI program (algorithm) and the parameters optimized through training, is stored in the memory 44b of the endoscopic image processing device 44. The memory 44b and the wiring derived from the memory 44b constitute a machine learning model connection section that can be connected to the trained model M.
[0172] The processor 44a executes an endoscopic image processing method, performing inference on the input image I1 using the trained model M, and outputting the inferred image I2 (see Figure 1). As a result of appropriate inference, the inferred image I2 becomes an endoscopic image with improved resolution compared to the input image I1.
[0173] According to the ninth embodiment, the trained model M learned in the configuration of the third to seventh embodiments is applied to the endoscope system 41 to perform appropriate inference with fewer errors and obtain output images with improved resolution.
[0174] Although the above description primarily focuses on the present invention as a machine learning image generation method, a machine learning method, an endoscopic image processing method, a machine learning image generation apparatus, a machine learning apparatus, and an endoscopic image processing apparatus, the present invention is not limited to these. For example, the present invention may be a computer program for causing a computer to perform processing according to the machine learning image generation method, the machine learning method, and the endoscopic image processing method. Furthermore, the present invention may also be a non-temporary recording medium readable by a computer that stores the computer program, etc.
[0175] Some examples of recording media for storing computer program products include portable recording media such as flexible disks, CD-ROMs (Compact Disc Read-only memory), and DVDs (Digital Versatile Discs), or recording media such as HDDs (Hard Disk Drives) and SSDs (Solid State Drives). The recording media may store not only the entire computer program, but also only a portion of it. Furthermore, the entire computer program, or a portion thereof, may be distributed or provided via a communication network. By installing the computer program from the recording media to a computer, or by downloading and installing the computer program via a communication network, the computer can read the computer, execute all or part of its operations, and perform the processing described above in accordance with the machine learning image generation method, machine learning method, and endoscopic image processing method.
[0176] Furthermore, the present invention is not limited to the embodiments described above. The present invention can be implemented by modifying its components during the implementation stage, without departing from the spirit of the invention. In addition, various forms of the invention can be formed by appropriately combining the multiple components disclosed in the above embodiments. For example, some components may be deleted from all the components disclosed in the embodiments. Furthermore, components from different embodiments may be appropriately combined. Thus, it goes without saying that various modifications and applications are possible without departing from the spirit of the invention.
Claims
1. For candidate ground truth images with higher resolution than the training images that are paired with the correct image, A reduction process is performed on components in at least some frequency bands within the frequency band higher than the Nyquist frequency of the aforementioned training image. A method for generating images for machine learning, characterized by generating the aforementioned ground truth images for machine learning to improve the resolution of input images.
2. The image generation method for machine learning according to claim 1, characterized in that the reduction process reduces components in all frequency bands higher than the Nyquist frequency.
3. The reduction process is characterized by setting all frequency band components higher than the Nyquist frequency to zero, as described in claim 1, for machine learning image generation.
4. The reduction process is characterized in that it reduces the candidate ground truth image so that its resolution is lower than that of the candidate ground truth image and higher than that of the training image, without reducing the components of the candidate ground truth image in the frequency band below the Nyquist frequency of the training image. This is the machine learning image generation method according to claim 1.
5. The reduction process is characterized by reducing the candidate ground truth image so that it has the same resolution as the training image, as described in claim 4 for machine learning image generation method.
6. The image generation method for machine learning according to claim 1, further characterized in that the reduction process reduces components in at least a portion of the frequency bands within a frequency band that is half or more of the Nyquist frequency and less than or equal to the Nyquist frequency.
7. The image generation method for machine learning according to claim 1, characterized in that the reduction process is performed by applying a low-pass filter to the candidate ground truth image.
8. The reduction process is performed by applying a low-pass filter to the candidate ground truth image, and then reducing the size of the candidate ground truth image after applying the low-pass filter, as described in claim 1.
9. For candidate ground truth images with a higher resolution than the training images that are paired with the ground truth images, a reduction process is performed on components in at least some frequency bands within the frequency band higher than the Nyquist frequency of the training images to generate the ground truth image. Alternatively, the ground truth image captured by an imaging device equipped with an optical low-pass filter that reduces components in at least some frequency bands within the frequency band higher than the Nyquist frequency of the training image, The aforementioned training image and, A machine learning method characterized by using to perform inference to improve the resolution of an input image and training a machine learning model to generate an output image.
10. When performing the reduction process, The machine learning method according to claim 9, characterized in that the reduction process reduces components in all frequency bands higher than the Nyquist frequency.
11. When performing the reduction process, The machine learning method according to claim 9, characterized in that the reduction process reduces all frequency band components higher than the Nyquist frequency to zero.
12. When performing the reduction process, The machine learning method according to claim 9, characterized in that the reduction process reduces the candidate ground truth image so that its resolution is lower than the resolution of the candidate ground truth image and higher than the resolution of the training image, without reducing the components of the candidate ground truth image in the frequency band below the Nyquist frequency of the training image.
13. For candidate ground truth images with a higher resolution than the training images that are paired with the ground truth images, a reduction process is performed on components in at least some frequency bands within the frequency band higher than the Nyquist frequency of the training images to generate the ground truth image. Alternatively, the ground truth image captured by an imaging device equipped with an optical low-pass filter that reduces components in at least some frequency bands within the frequency band higher than the Nyquist frequency of the training image, The aforementioned training image and, A machine learning model connection unit that can connect to the machine learning model trained by a machine learning method that trains a machine learning model using, Processor and Includes, The aforementioned processor, The received endoscopic image is input into the machine learning model that performs inference to improve the resolution of the input image and generates an output image. An endoscope image processing device characterized by outputting the endoscopic image with improved resolution from the machine learning model.
14. The endoscope image processing apparatus according to claim 13, characterized in that the reduction process reduces components in all frequency bands higher than the Nyquist frequency.
15. The endoscopic image processing apparatus according to claim 13, characterized in that the reduction process reduces all frequency band components higher than the Nyquist frequency to zero.
16. The reduction process is characterized in that it reduces the candidate ground truth image so that its resolution is lower than that of the candidate ground truth image and higher than that of the training image, without reducing the components of the candidate ground truth image in the frequency band below the Nyquist frequency of the training image. This is the image processing apparatus for endoscopes according to claim 13.
17. The aforementioned machine learning model connection unit is A storage medium on which the aforementioned machine learning model is stored, Wiring derived from the aforementioned storage medium, The endoscopic image processing apparatus according to claim 13, characterized by including the following:
18. The processor includes: For candidate ground truth images with higher resolution than the training images that are paired with the correct image, The training image is subjected to a reduction process on components in at least some of the frequency bands higher than the Nyquist frequency. An endoscope image processing program that generates the aforementioned ground truth image for machine learning to improve the resolution of the input image.