Trained model and information processing device

A trained model enhances image contrast in low-resolution images by correcting for pixel and contrast differences between imaging systems, addressing the limitations of existing models in converting low-resolution images to high-quality images for endoscopes.

JP2026068588APending Publication Date: 2026-04-22OLYMPUS MEDICAL SYST CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
OLYMPUS MEDICAL SYST CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Trained models using high-resolution images cannot effectively convert low-resolution images into images with higher contrast than those captured by high-resolution cameras, leading to issues in image quality for small-diameter endoscopes and mass-produced imagers.

Method used

A trained model is developed using a ground truth image and training image generation process that corrects for the difference in pixel count and contrast between different imaging systems, involving reduction processes and blur addition to enhance contrast in low-resolution images.

Benefits of technology

The trained model can generate images with higher contrast than high-resolution images, improving image quality in endoscopic applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068588000001_ABST
    Figure 2026068588000001_ABST
Patent Text Reader

Abstract

This provides a pre-trained model that can generate images with higher contrast than high-resolution images acquired by high-resolution cameras. [Solution] The trained model is machine-learned using a ground truth image obtained by applying a first reduction process to an original image of a predetermined subject captured by a first imaging system, and a training image obtained by applying a second reduction process to the original image and then applying a blurring process to correct the difference in contrast due to the difference in the number of pixels between the original image and the processing target image captured by a second imaging system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learned model and an information processing device.

Background Art

[0002] There is an increasing demand for the development of endoscopes with a small diameter that cause less pain to patients, single-use type endoscopes that prevent the spread of infectious diseases, and the like. On the other hand, excessive thinning of the endoscope and deterioration of the image quality associated with the use of mass-produced imagers are major causes such as overlooking lesions and stress on doctors. Therefore, there is a strong demand for a technology that improves the image quality while using a small-diameter endoscope or a mass-produced imager.

[0003] As one method for realizing image quality improvement, deep learning is known. In this deep learning, learning data including a pair of a correct image, which is a high-resolution image captured by a high-resolution camera (high-end machine), and a training image, which is a low-resolution image obtained by degrading the correct image by adding blur or the like, is learned by AI to generate a learned model. By performing an inference to convert a low-resolution image, which is a processing target image, into a high-resolution image using the learned model thus generated, image generation with extremely high image quality is realized.

[0004] For example, Patent Document 1 describes a method for creating a machine learning model with enhanced robustness against noise by creating learning data using, as a correct image, an image obtained by reducing noise in a medical input image and, as a training image, an image obtained by reducing the resolution of the correct image.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, since the high-resolution images used for training are acquired by a high-resolution camera, a trained model trained using these high-resolution images cannot convert low-resolution images acquired by a low-resolution camera (low-end camera) into images with higher contrast than high-resolution images acquired by a high-resolution camera.

[0007] The present invention has been made in view of the above circumstances, and aims to provide a trained model that can generate images with higher contrast than high-resolution images acquired by a high-resolution camera, and an information processing device that generates the trained model. [Means for solving the problem]

[0008] A trained model according to one aspect of the present invention is machine-learned using a ground truth image obtained by applying a first reduction process to an original image of a predetermined subject captured by a first imaging system, and a training image obtained by applying a second reduction process to the original image and then applying a blurring process to correct the difference in contrast due to the difference in the number of pixels between the original image and the processing target image captured by a second imaging system.

[0009] Furthermore, a trained model in another aspect of the present invention includes: a first ground truth image obtained by applying a first reduction process to a first source image taken of a predetermined subject with a first imaging system; a first training image obtained by applying a second reduction process to the first source image or a first intermediate image obtained by applying an imaging system simulation process to the first source image, and applying a blur addition process to correct the difference in contrast due to the difference in the number of pixels with the processing target image taken with a third imaging system; a first training set consisting of the first ground truth image and the first training image as a pair; and the predetermined subject with a second imaging system. Machine learning is performed using a second ground truth image obtained by applying a third reduction process to a second source image captured by the third imaging system, a second training image obtained by applying a fourth reduction process to the second source image or a second intermediate image obtained by applying the imaging system simulation process to the second source image, and applying a blurring process to correct the difference in contrast due to the difference in the number of pixels between it and the target image captured by the third imaging system, a second training set consisting of the second ground truth image and the second training image paired together, and multiple training sets including the first training set and the second training set.

[0010] Furthermore, a trained model in another aspect of the present invention is machine-learned using: a ground truth image obtained by applying a first reduction process to the original image captured by the first imaging system, and then applying a second reduction process with a different aspect ratio to the image that has undergone the first reduction process; and a training image obtained by applying a third reduction process to the original image, and then applying a blurring process to correct the difference in contrast due to the difference in the number of pixels between the original image and the image to be processed captured by the second imaging system, and then applying a fourth reduction process with a different aspect ratio to the image that has undergone the blurring process.

[0011] Furthermore, an information processing device in another aspect of the present invention comprises: a ground truth image generation unit that performs a first reduction process to reduce the size of an image on a source image of a predetermined subject captured by a first imaging system; and a training image generation unit that performs a second reduction process to reduce the size of the image on the source image and applies a blurring process to correct the difference in contrast due to the difference in the number of pixels between the source image and a processing target image captured by a second imaging system. [Effects of the Invention]

[0012] The present invention provides a trained model that can generate images with higher contrast than high-resolution images acquired by a high-resolution camera, and an information processing device that generates the trained model. [Brief explanation of the drawing]

[0013] [Figure 1] This figure shows an example of the processing flow for model creation and trained model inference according to the first embodiment. [Figure 2] This figure shows an example of the relationship between the original image and the ground truth image in the processing flow of the model creation process of the first embodiment. [Figure 3] This figure shows an example of the relationship between the source image and the training image in the processing flow of the model creation process of the first embodiment. [Figure 4] This figure shows an example of the processing flow for model creation and trained model inference according to the second embodiment. [Figure 5] This figure shows an example of the relationship between the source image and the ground truth image in the processing flow of the model creation process of the second embodiment. [Figure 6] This figure shows an example of the relationship between the source image and the training image in the processing flow of the model creation process of the second embodiment. [Figure 7] This figure shows an example of the processing flow for model creation and trained model inference according to the third embodiment. [Figure 8] This figure shows an example of the relationship between the source image and the ground truth image in the processing flow of the model creation process of the third embodiment. [Figure 9] This figure shows an example of the relationship between the source image and the training image in the processing flow of the model creation process of the third embodiment. [Figure 10] This figure shows an example of the processing flow for model creation and trained model inference according to the fourth embodiment. [Figure 11]It is a diagram showing an example of the relationship between the original image and the correct answer image in the processing flow of the model creation process according to the fourth embodiment. [Figure 12] It is a diagram showing an example of the relationship between the original image and the training image in the processing flow of the model creation process according to the fourth embodiment. [Figure 13] It is a diagram showing an example of the processing flow of the model creation process and the learned model inference process according to the fifth embodiment. [Figure 14] It is a block diagram of the machine learning device 1 which is an information processing device according to the sixth embodiment. [Figure 15] It is a block diagram of a configuration in which a learned model learned by a machine learning device which is an information processing device according to the seventh embodiment is applied to an endoscope system.

Embodiments for Carrying Out the Invention

[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings. (First Embodiment) FIG. 1 is a diagram showing an example of the processing flow of the model creation process and the learned model inference process according to the first embodiment. FIG. 2 is a diagram showing an example of the relationship between the original image and the correct answer image in the processing flow of the model creation process according to the first embodiment. FIG. 3 is a diagram showing an example of the relationship between the original image and the training image in the processing flow of the model creation process according to the first embodiment.

[0015] As shown in FIG. 1, the model creation process S100 generates a correct answer image 11 and a training image 14 from the input original image 10. Then, the model creation process S100 executes a learning process S150 using the correct answer image 11 and the training image 14, and generates a learned model M. The original image 10 is an image captured by the high-resolution first imaging system. The detailed processing flow of the model creation process S100 will be described later.

[0016] The trained model inference process S200 performs inference on the input image 31 using the trained model M generated by the model creation process S100, and outputs an output image 32. The image 31 is an image captured by a second imaging system with a lower resolution than the first imaging system.

[0017] The hardware constituting the trained model inference process S200 is, for example, a general-purpose processor such as a CPU. In this case, the program describing the inference algorithm and the parameters used in that inference algorithm are stored as the trained model M, for example in a memory unit. Alternatively, the trained model inference process S200 may be a dedicated processor in which the inference algorithm is implemented in hardware. A dedicated processor is, for example, an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). In this case, the parameters used in the inference algorithm are stored as the trained model M, for example in a memory unit.

[0018] An inference algorithm can utilize a neural network. The parameters are the weight coefficients of the inter-node connections in the neural network. A neural network includes an input layer where image data is input, an intermediate layer that performs calculations on the data input through the input layer, and an output layer that outputs image data based on the calculation results from the intermediate layer. A Convolutional Neural Network (CNN) is preferred as the neural network used for inference. However, various AI (Artificial Intelligence) technologies can be employed, not just CNNs.

[0019] To realize these technologies, conventional general-purpose computing circuits such as CPUs and FPGAs can be used, but since much of the processing in neural networks involves matrix multiplication, specialized devices such as GPUs and TPUs (Tensor Processing Units) are sometimes used. In recent years, these artificial intelligence (AI) dedicated hardware, called "Neural Network Processing Units (NPUs)," have been designed to be integrated and embedded together with CPUs and other circuits, and are sometimes part of the processing circuit.

[0020] The model creation process S100 receives the source image 10 captured by the first imaging system as input. The source image 10 is a higher resolution image than the processing target image 31 captured by the second imaging system. Furthermore, the source image 10 is the image that will be used as the basis for the ground truth image 11 and training image 14 used in the learning process S150.

[0021] The ground truth image reduction process S110 reduces the original image 10 to generate the ground truth image 11. The process of reducing the original image 10 may involve interpolating pixels or downsampling pixels. By reducing the image, for example, a structure such as blood vessels that was represented by 10 pixels can be represented by 5 pixels while maintaining almost the same contrast. In other words, the number of pixels required to represent a structure is reduced while maintaining almost the same contrast.

[0022] In other words, the ground truth image reduction process S110 converts the spatial frequency characteristics per pixel to the higher frequency side, generating a ground truth image 11 with higher contrast than the original image 10. That is, as shown in Figure 2, the ground truth image reduction process S110 generates a ground truth image 11 with fewer pixels and higher contrast than the original image 10.

[0023] Note that the number of pixels in the correct image 11 is the same as the number of pixels in the training image 14, which will be described later, but it is not limited to this. The number of pixels in the correct image 11 is only required to be less than the number of pixels in the original image 10.

[0024] Image sensor information 21 and optical system information 22 are input to the imaging system simulation process S120. The image sensor information 21 includes information on the image sensor of the first imaging system and the image sensor of the second imaging system. The optical system information 22 includes information on the optical system of the first imaging system and the optical system of the second imaging system.

[0025] The imaging system simulation process S120 uses the image sensor information 21 and the optical system information 22 to reduce the characteristics of the original image 10 captured by the first imaging system and degrade the frequency characteristics so that they are similar to the characteristics of the processing target image 31 captured by the second imaging system, thereby generating a first intermediate image 12. As a result, as shown in Figure 3, a first intermediate image 12 is generated that has fewer pixels and lower contrast than the original image 10. The characteristics of this first intermediate image 12 are approximately the same as those of the processing target image 31 captured by the second imaging system.

[0026] In other words, the imaging system simulation process S120 performs a process to convert the original image 10, which was captured by the first imaging system, into an image that appears as if it were captured by the second imaging system. Specifically, based on the optical characteristics information and noise characteristics information of the first imaging system and the optical characteristics information and noise characteristics information of the second imaging system, a correction filter corresponding to the difference in optical system characteristics between the first and second imaging systems and correction noise corresponding to the difference in noise characteristics are calculated and added to the original image 10. In addition, the original image is reduced in size so that it has a number of pixels similar to that of an image captured by the second imaging system.

[0027] Noise characteristic information refers to information such as random noise, indicating where and how much noise is added. It is preferable that the noise characteristic information includes information on noise location and noise amount, but it may include only one of them. When using standard deviation information of noise as noise characteristic information, the noise amount may be uniform or variable relative to the pixel value.

[0028] If the imaging methods of the first and second imaging systems are different, a process is performed to simulate the imaging methods. Two known imaging methods are sequential and simultaneous. In the sequential imaging method, multiple colors of illumination light are irradiated sequentially, and images are captured by the monochrome image sensor at the timing when each color of light is irradiated. In the simultaneous imaging method, white light is irradiated, and the image sensor has multiple pixels with different light-receiving colors.

[0029] For example, if the first imaging system uses a sequential imaging method and the second imaging system uses a simultaneous imaging method, the imaging method of the second imaging system is simulated by generating the signal for each pixel of the simultaneous imaging sensor based on the signal from at least one of the multiple color light signals acquired by the pixels of the monochrome image sensor of the first imaging system.

[0030] By simulating the imaging method, the characteristics of the first intermediate image 12 are made to be similar to those of the image to be processed 31.

[0031] The training image reduction process S130 reduces the first intermediate image 12 to generate a second intermediate image 13 (see Figure 3). The process of reducing the first intermediate image 12 is, for example, a process of interpolating or downsampling pixels. The training image reduction process S130 converts the spatial frequency per pixel to the higher frequency side. As a result, as shown in Figure 3, a second intermediate image 13 is generated that has fewer pixels than the first intermediate image 12 and has higher contrast than the first intermediate image 12.

[0032] The pixel count difference correction process S140 corrects the second intermediate image 13, which has a higher contrast than the first intermediate image 12 due to the training image reduction process S130, so that it has a contrast similar to the image captured by the second imaging system, thereby generating the training image 14.

[0033] Specifically, the pixel count difference correction process S140 performs a blur addition process to correct the difference in contrast caused by the difference in the number of pixels between the training image 14 and the image to be processed 31. In other words, the pixel count difference correction process S140 applies a correction filter and degrades the frequency characteristics that were increased by the training image reduction process S130 to be equivalent to the frequency characteristics of the image to be processed 31 (or equivalent to the frequency characteristics of the first intermediate image 12). As a result, as shown in Figure 3, a training image 14 is generated that has fewer pixels than the image to be processed 31 (first intermediate image 12) and has a similar level of contrast to the image to be processed 31 (first intermediate image 12).

[0034] The correction filter can be, for example, a Gaussian filter, but is not limited to this; any filter capable of adding blur can be used. Alternatively, the blur addition process may be performed using PSF (Point Spread Function) data obtained from the optical system data.

[0035] Noise may be added again to compensate for the amount of noise reduced by the pixel difference correction process.

[0036] It is desirable to perform the imaging system simulation process, the reduction process, the pixel count difference correction process, and the noise re-addition process individually in this order, but the order in which the processes are performed can be changed, or multiple processes can be performed simultaneously.

[0037] The learning process S150 performs machine learning using the high-contrast ground truth image 11 generated by the ground truth image reduction process S110 and the low-contrast training image 14 generated by the imaging system simulation process S120, training image reduction process S130, and pixel count difference correction process S140, and outputs a trained model M. The pair of ground truth image 11 and training image 14 is a training dataset used in machine learning, and is also called training data.

[0038] In this embodiment, a ground truth image reduction process S110 is applied to the original image 10 captured by the first imaging system (high resolution) to generate a ground truth image 11 with higher contrast than the original image 10. Furthermore, an imaging system simulation process S120 is used to create an image equivalent to the processing target image 31 (first intermediate image 12) from the original image 10. Subsequently, a training image reduction process S130 and a pixel count difference correction process S140 are performed on the first intermediate image 12 to generate a training image 14. During training, a training process S150 is performed using the ground truth image 11 and the training image 14 as a pair to generate a trained model M.

[0039] During inference, the image captured by the second imaging system (low resolution) is input as the image to be processed 31 to the trained model inference process S200. Inference is performed using the trained model M, and an output image (inference image) 32 is produced that has higher contrast and resolution than the original image 10 captured by the first imaging system (high resolution).

[0040] As a result, the trained model M of this embodiment can output output images with higher contrast than high-resolution images captured by a high-resolution camera.

[0041] (Second embodiment) Next, a second embodiment will be described. Generally, AI cannot be trained without training images with higher contrast than images captured by a high-resolution camera. Therefore, when images with higher contrast than those captured by a high-resolution camera are required, it is necessary to create ground truth images with higher contrast than those captured by a high-resolution camera and train the AI ​​with them. In the second embodiment, an example of performing machine learning using self-images captured by a camera such as a high-resolution camera as the source images for the ground truth and training images will be described.

[0042] Figure 4 shows an example of the processing flow for the model creation process and the trained model inference process according to the second embodiment. Figure 5 shows an example of the relationship between the original image and the ground truth image in the processing flow of the model creation process of the second embodiment. Figure 6 shows an example of the relationship between the original image and the training image in the processing flow of the model creation process of the second embodiment. Note that in Figures 4, 5, and 6, components similar to those in Figures 1, 2, and 3 are denoted by the same reference numerals and their descriptions are omitted.

[0043] The model creation process S300 receives the original image 10 captured by the first imaging system as input. The trained model inference process S400 receives the image to be processed 33 captured by the first imaging system as input. In other words, in this embodiment, the original image 10 and the image to be processed 33 are images captured by the same imaging system (first imaging system).

[0044] The ground truth image reduction process S110 reduces the original image 10 to generate the ground truth image 11. The ground truth image reduction process S110 converts the spatial frequency per pixel to the higher frequency side, resulting in the generation of a ground truth image 11 with higher contrast than the original image 10. In other words, as shown in Figure 5, the ground truth image reduction process S110 generates a ground truth image 11 with fewer pixels and higher contrast than the original image 10. Note that the number of pixels in the ground truth image 11 only needs to be less than the number of pixels in the original image 10.

[0045] The training image reduction process S330 reduces the original image 10 to generate an intermediate image 15 (see Figure 6). The training image reduction process S330 converts the spatial frequency per pixel to the higher frequency side, generating an intermediate image 15 that has fewer pixels than the original image 10 and higher contrast than the original image 10.

[0046] As described above, the original image 10 and the image to be processed 33 are images captured by the same first imaging system and have similar characteristics. Therefore, the intermediate image 15 generated by the training image reduction process S330 has fewer pixels than the image to be processed 33 and is an image with higher contrast than the image to be processed 33.

[0047] The pixel count difference correction process S340 corrects the intermediate image 15, which has higher contrast than the target image 33 due to the training image reduction process S330, so that it has a similar contrast to the target image 33, and generates the training image 14.

[0048] Specifically, the pixel count difference correction process S340 performs a blur addition process to correct the contrast difference caused by the difference in the number of pixels between the training image 14 and the image to be processed 33. In other words, the pixel count difference correction process S340 applies a correction filter and degrades the frequency characteristics that were increased by the training image reduction process S330 to be equivalent to the frequency characteristics of the image to be processed 33. As a result, as shown in Figure 6, a training image 14 is generated that has fewer pixels than the image to be processed 33 and has a similar level of contrast to the image to be processed 33.

[0049] Noise may be added again to compensate for the amount of noise reduced by the pixel difference correction process.

[0050] It is preferable to perform the reduction process, the pixel difference correction process, and the noise re-addition process individually in this order, but the order in which the processes are performed can be changed, or multiple processes can be performed simultaneously.

[0051] The learning process S150 performs machine learning using the high-contrast ground truth image 11 generated by the ground truth image reduction process S110 and the low-contrast training image 14 generated by the training image reduction process S330 and the pixel count difference correction process S340, and outputs the trained model M.

[0052] During inference, the image to be processed 33 captured by the first imaging system's camera is input to the trained model inference process S400, and inference is performed using the trained model M, outputting an output image (inference image) 34 with higher contrast than the image to be processed 33 captured by the first imaging system's camera.

[0053] In this way, by generating a trained model M using a self-image captured with the same first imaging system as the image to be processed 33 as the source image 10 of the training dataset (ground truth image 11 and training image 14), it is possible to generate output images with higher contrast than the self-image captured by the camera, even when the performance of the high-resolution camera is improved.

[0054] (Third embodiment) Next, a third embodiment will be described. Recent advancements in AI technology have made it possible to generate images equivalent to high-resolution images taken with high-resolution cameras from low-resolution images taken with low-resolution cameras. In other words, the performance gap between high-resolution and low-resolution cameras is narrowing, and a decrease in the availability and performance degradation of high-resolution cameras are expected. As a result, there is a possibility that a sufficient supply of training images taken with high-resolution cameras may not be available.

[0055] Therefore, it is necessary to create training images (training datasets) with higher contrast than images captured by a high-resolution camera from images captured by a low-resolution camera. In the third embodiment, an example of performing machine learning using low-resolution images captured by a low-resolution camera as the source images for the ground truth and training images will be described.

[0056] Figure 7 shows an example of the processing flow for the model creation process and the trained model inference process according to the third embodiment. Figure 8 shows an example of the relationship between the original image and the ground truth image in the processing flow of the model creation process of the third embodiment. Figure 9 shows an example of the relationship between the original image and the training image in the processing flow of the model creation process of the third embodiment. Note that in Figures 7, 8, and 9, components similar to those in Figures 1, 2, and 3 are denoted by the same reference numerals and their descriptions are omitted.

[0057] The model creation process S500 receives the original image 16 captured by the first imaging system as input. The trained model inference process S600 receives the image 35 to be processed captured by the second imaging system as input.

[0058] The first imaging system is a low-resolution imaging system, and the second imaging system is a higher-resolution imaging system than the first imaging system. In other words, in this embodiment, the original image 16 is a lower-resolution image than the processing target image 35 captured by the high-resolution second imaging system.

[0059] The ground truth image reduction process S110 reduces the original image 16 to generate the ground truth image 11. The ground truth image reduction process S110 converts the spatial frequency per pixel to the higher frequency side, generating a ground truth image 11 with higher contrast than the original image 16. In other words, as shown in Figure 8, the ground truth image reduction process S110 generates a ground truth image 11 with fewer pixels and higher contrast than the original image 10. Note that the number of pixels in the ground truth image 11 only needs to be less than the number of pixels that results in a contrast equal to or greater than that of the image 35 being processed.

[0060] The training image reduction process S530 reduces the original image 16 to generate an intermediate image 17 (see Figure 9). The training image reduction process S530 converts the spatial frequency per pixel to the higher frequency side, generating an intermediate image 17 that has fewer pixels than the processing target image 35 and has higher contrast than the processing target image 35.

[0061] The pixel count difference correction process S540 corrects the intermediate image 17, which has higher contrast than the target image 35 due to the training image reduction process S530, so that it has a similar contrast to the target image 35, and generates the training image 14.

[0062] Specifically, the pixel count difference correction process S540 performs a blur addition process to correct the difference in contrast caused by the difference in the number of pixels between the training image 14 and the image to be processed 35. In other words, the pixel count difference correction process S540 applies a correction filter and degrades the frequency characteristics that were increased by the training image reduction process S530 to be equivalent to the frequency characteristics of the image to be processed 35. As a result, as shown in Figure 9, a training image 14 is generated that has fewer pixels than the image to be processed 35 and has a similar level of contrast to the image to be processed 35.

[0063] Noise may be added again to compensate for the amount of noise reduced by the pixel difference correction process.

[0064] It is desirable to perform the imaging system simulation process, the reduction process, the pixel count difference correction process, and the noise re-addition process individually in this order, but the order in which the processes are performed can be changed, or multiple processes can be performed simultaneously.

[0065] The learning process S150 performs machine learning using the high-contrast ground truth image 11 generated by the ground truth image reduction process S110 and the low-contrast training image 14 generated by the training image reduction process S550 and the pixel count difference correction process S540, and outputs the trained model M.

[0066] During inference, the image 35 to be processed, captured by the second imaging system's camera, is input to the trained model inference process S600. Inference is performed using the trained model M, and an output image (inference image) 36 with higher contrast than the image 35 to be processed, captured by the second imaging system's camera, is output.

[0067] Thus, in this embodiment, a trained model M is generated using low-resolution images captured by the low-resolution camera of the first imaging system as the source images 16 for the ground truth image 11 and the training image 14. As a result, even if high-resolution images that serve as the source images for the ground truth image 11 and the training image 14 are unavailable due to a decrease in the availability or performance degradation of high-resolution cameras, it is possible to generate an output image 36 with higher contrast than the processing target image 35 captured by the high-resolution camera of the second imaging system.

[0068] (Fourth embodiment) Next, a fourth embodiment will be described. To create an AI specifically for medical applications, it is necessary to acquire images for inference using an imaging device and to train the AI ​​with training images that are appropriate for the purpose. A vast number of training images are required for pre-training, but as in the first to third embodiments, it is difficult to acquire a sufficient number of training images when using images captured by a single imaging system as the source images.

[0069] Therefore, in the fourth embodiment, we will describe a model creation process that creates a trained model using images captured by multiple cameras, or in other words, multiple imaging systems, as source images.

[0070] Figure 10 shows an example of the processing flow for the model creation process and the trained model inference process according to the fourth embodiment. Figure 11 shows an example of the relationship between the original image and the ground truth image in the processing flow of the model creation process of the fourth embodiment. Figure 12 shows an example of the relationship between the original image and the training image in the processing flow of the model creation process of the fourth embodiment. Note that in Figures 10, 11, and 12, components similar to those in Figures 1, 2, and 3 are denoted by the same reference numerals and their descriptions are omitted.

[0071] The model creation process S700 includes the first training image set creation process S710_1, the second training image set creation process S710_2, ..., the Nth training image set creation process S710_N, and the training process S150.

[0072] The first training image set creation process S710_1, the second training image set creation process S710_2, ..., and the Nth training image set creation process S710_N are input to the first source image 18_1, the second source image 18_2, ..., and the Nth source image 18_N, each captured by a different camera. The first source image 18_1 is an image captured by the first imaging system, the second source image 18_2 is an image captured by the second imaging system, and the Nth source image 18_N is an image captured by the Nth imaging system. In addition, the trained model inference process S800 is input to the processing target image 37, for example, captured by the third imaging system.

[0073] For example, the first imaging system has a higher resolution than the third imaging system, and the second imaging system has a lower resolution than the third imaging system. Therefore, as shown in Figure 11, the first source image 18_1 is a high-resolution image with more pixels and higher contrast than the image to be processed 37, while the second source image 18_2 is a low-resolution image with fewer pixels and lower contrast than the image to be processed 37.

[0074] In the first training image set creation process S710_1, a first ground truth image reduction process S110_1 is performed on the first source image 18_1 to reduce the number of pixels. As a result of this first ground truth image reduction process S110_1, a first ground truth image 11_1 is generated that has fewer pixels and higher contrast than the processing target image 37, as shown in Figure 11.

[0075] Furthermore, in the first training image set creation process S710_1, a first training image reduction process S730_1 is performed to reduce the number of pixels in the first source image 18_1. This first training image reduction process S730_1 generates an intermediate image 19_1 with fewer pixels and higher contrast than the processing target image 37, as shown in Figure 12.

[0076] The first pixel difference correction process S740_1 corrects the intermediate image 19_1, which has higher contrast than the target image 37 due to the first training image reduction process S730_1, so that it has a contrast similar to the target image 37, thereby generating the first training image 14_1.

[0077] Specifically, the first pixel count difference correction process S740_1 performs a blur addition process to correct the contrast difference caused by the difference in the number of pixels between the first training image 14_1 and the image to be processed 37. In other words, the first pixel count difference correction process S740_1 applies a correction filter and degrades the frequency characteristics that were increased by the first training image reduction process S730_1 to be equivalent to the frequency characteristics of the image to be processed 37. As a result, as shown in Figure 12, the first training image 14_1 is generated which has fewer pixels than the image to be processed 37 and has a similar level of contrast to the image to be processed 37.

[0078] Additionally, noise may be added again to compensate for the amount of noise reduced by the pixel count difference correction process.

[0079] It is desirable to perform the imaging system simulation process, the reduction process, the pixel count difference correction process, and the noise re-addition process individually in this order, but the order in which the processes are performed can be changed, or multiple processes can be performed simultaneously.

[0080] The pairs of the first correct image 11_1 and the first training image 14_1 generated in this way are input to the learning process S150 as the first learning image set TS_1.

[0081] Similarly, the second training image set creation process S710_2, ... and the Nth training image set creation process S710_N generate the second training image set TS_2, ... and the Nth training image set TS_N, which are then input to the training process S150.

[0082] In the learning process S150, machine learning is performed using multiple training image sets, including the first training image set TS_1, the second training image set TS_2, ..., and the Nth training image set TS_N, to generate a trained model M.

[0083] During inference, the image 37 to be processed, captured by the third imaging system's camera, is input to the trained model inference process S800. Inference is performed using the trained model M, and an output image (inference image) 38 with higher contrast than the image 37 to be processed, captured by the third imaging system's camera, is output.

[0084] As described above, by generating training images, even if it is not possible to prepare a sufficient number of training images that are suitable for the purpose of the inference target camera, images captured by multiple imaging systems can be reused as source images for training, thus ensuring that a sufficient quantity of training images are prepared.

[0085] (Fifth embodiment) Next, a fifth embodiment will be described. Various types of image sensors can be used to capture the subject, such as monochrome, Bayer, and complementary color sensors. For example, in images captured with a complementary color sensor, the number of pixels changes during the image processing process, and the effect of AI processing changes depending on where it is applied. Therefore, even if the ratio of vertical to horizontal dimensions is not the expected value, it is necessary to apply AI processing using a model that has been trained to match that ratio.

[0086] Therefore, in the fifth embodiment, we will describe an example of creating training images that take into account different ratios in the vertical and horizontal directions and performing machine learning.

[0087] Figure 13 shows an example of the processing flow for the model creation process and the trained model inference process according to the fifth embodiment. In Figure 13, components similar to those in Figure 1 are denoted by the same reference numerals and their explanation is omitted.

[0088] The model creation process S900 receives the source image 10 as input. The source image 10 can be an image captured with various image sensors such as monochrome, Bayer type, or complementary color type. By applying the ground truth image reduction process S110 to the source image 10, the ground truth intermediate image 41 is generated.

[0089] The vertical reduction process S910 for the ground truth image generates the ground truth image 42 by performing a reduction process with a different aspect ratio on the ground truth intermediate image 41. Specifically, it reduces the number of vertical pixels in the ground truth intermediate image 41 to half the number of horizontal pixels by interpolating or downsampling pixels in the vertical (vertical) direction of the ground truth intermediate image 41.

[0090] Furthermore, by applying the training image reduction process S130 and the pixel count difference correction process S140 to the original image 10, an intermediate training image 43 is generated.

[0091] Additionally, noise may be added again to compensate for the amount of noise reduced by the pixel count difference correction process.

[0092] It is preferable to perform the reduction process, the pixel difference correction process, and the noise re-addition process individually in this order, but the order in which the processes are performed can be changed, or multiple processes can be performed simultaneously.

[0093] The vertical reduction processing of the training image S920 generates the training image 44 by performing a reduction process with a different aspect ratio on the training intermediate image 43. Specifically, the vertical pixels of the training intermediate image 43 are interpolated or decimated to reduce the number of vertical pixels in the training intermediate image 43 to half the number of horizontal pixels.

[0094] The learning process S150 performs machine learning using the ground truth image 42 generated by the ground truth image vertical reduction process S910 and the training image 44 generated by the training image vertical reduction process S920, and outputs the trained model M.

[0095] During inference, the odd or even fields of the captured image 51 taken by the complementary color camera are input to the trained model inference process S1000 as the image to be processed 52. Inference is performed using the trained model M, and an output image (inference image) 53 is output. This makes it possible to output an output image 53 with uniform characteristics in both the horizontal and vertical directions.

[0096] (Sixth embodiment) Figure 14 is a block diagram of the machine learning device 1, which is an information processing device that acquires the original image from the endoscope system 81. The machine learning device 1 comprises a machine learning image generation device 60 and a model learning processing unit 70. The machine learning image generation device 60 comprises a training image generation unit 62 and a ground truth image generation unit 63.

[0097] The endoscope system 81 comprises a memory 84 and a first imaging system 85 (a light source unit 82 and an imaging device 83). The light source unit 82 irradiates illumination light onto an object 100, which is a predetermined subject. The subject light beam, which is the reflected light from the object 100, is incident on the imaging device 83.

[0098] The imaging device 83 forms an image of the subject's light beam, captures the image with the image sensor 18, and outputs the original image. The original image is input to the training image generation unit 62 and the ground truth image generation unit 63.

[0099] Memory 84 is a storage medium that non-volatilely stores information related to the endoscope system 81. The information stored in memory 84 includes first imaging system information. First imaging system information is information related to the imaging device 83 that acquires the original image. Second imaging system information is stored in the memory (not shown) of the machine learning device 1, etc.

[0100] The first imaging system information includes pixel count information of the image sensor, color information of the image acquired by the imaging device 83, optical characteristic information including PSF of the imaging optical system, noise characteristic information related to the image sensor and the readout circuit from the image sensor, and color filter information indicating whether the image sensor acquires a Bayer image, a plane sequential image, or a complementary color image.

[0101] The light source unit 82 is configured to emit illumination light corresponding to multiple types of observation modes, for example. Observation modes include, for example, white light imaging (WLI) mode and narrowband imaging (NBI) mode. The original image light source information is information indicating the type of illumination light (WLI illumination light, NBI illumination light, etc.) emitted by the light source unit 82 according to the observation mode.

[0102] The endoscope system 81 transmits the first imaging system information from the memory 84 to the machine learning device 1.

[0103] The training image generation unit 62 of the machine learning device 1 receives first imaging system information from the endoscope system 81 and second imaging system information from the machine learning image generation device 60. The ground truth image generation unit 63 receives first imaging system information from the endoscope system 81.

[0104] The training image generation unit 62 changes the pixel reduction ratio and correction filter according to the type of illumination light (WLI illumination light, NBI illumination light, etc.) based on the original image light source information. At this time, the training image generation unit 62 may further change the pixel reduction ratio and correction filter based on the optical characteristics information and noise characteristics information included in the first imaging system information and the second imaging system information.

[0105] Similarly, the ground truth image generation unit 63 changes the pixel reduction ratio, for example, according to the type of illumination light (WLI illumination light, NBI illumination light, etc.) based on the original image light source information. At this time, the ground truth image generation unit 63 may further change the pixel reduction ratio based on the optical characteristic information included in the first imaging system information and the second imaging system information.

[0106] Furthermore, the training image generation unit 62 and the ground truth image generation unit 63 may, if necessary, change the color correction process based on the color information included in the first imaging system information.

[0107] The training image generation unit 62 and the ground truth image generation unit 63 may, as necessary, change the conversion process from, for example, a Bayer image to a planar sequential image, or from a planar sequential image to a Bayer image, based on the color filter information included in the first imaging system information.

[0108] (Seventh Embodiment) Figure 15 is a block diagram showing an example configuration in which a trained model M, learned by a machine learning device 1 which is an information processing device, is applied to an endoscope system 91.

[0109] The endoscopic system 91 comprises an endoscope 92 and an endoscopic image processing device 94. The imaging device 93 is a second imaging system that forms an image of the subject light beam, captures the image with an image sensor, and outputs an endoscopic image. The endoscopic image output from the imaging device 93 becomes the input image to the endoscopic image processing device 94.

[0110] The endoscopic image processing device 94 includes, for example, a processor 94a and a memory 94b. The processor 94a is composed of an ASIC (Application Specific Integrated Circuit) including a CPU (Central Processing Unit), an FPGA (Field Programmable Gate Array), etc. However, the endoscopic image processing device 94 may also be configured as a dedicated electronic circuit that performs the functions of a trained model M.

[0111] Memory 94b is a storage medium that stores (non-volatilely stores) processing programs that realize the functions of each circuit. Processor 94a is connected to the wiring derived from memory 94b. The functions of the endoscopic image processing device 94 are realized when processor 94a reads and executes the processing programs stored in memory 94b. For example, the endoscopic image processing device 94 realizes an endoscopic image processing method by executing a processing program.

[0112] The trained model M generated by the machine learning device 1, that is, the combination of the AI ​​program (algorithm) and the parameters optimized through learning, is stored in the memory 94b of the endoscopic image processing device 94. The memory 94b and the wiring derived from the memory 94b constitute a machine learning model connection section that can be connected to the trained model M.

[0113] The processor 94a executes an endoscopic image processing method, performing inference on the input image (the image to be processed, or the captured image) using a pre-trained model M, and outputs the inferred image. As a result of appropriate inference, the inferred image becomes an endoscopic image with improved resolution compared to the input image.

[0114] The present invention is not limited to the embodiments described above, and various modifications and alterations are possible without changing the essence of the invention. For example, the present invention can be applied to AI using images to detect objects with poor visibility (for example, pedestrians with poor visibility using an in-vehicle camera). [Explanation of Symbols]

[0115] 10...Original image, 11...Correct image, 12...First intermediate image, 13...Second intermediate image, 14...Training image, 18...Image sensor, 21...Image sensor information, 22...Optical system information, 31...Image to be processed, 32...Output image, 60...Machine learning image generation device, 62...Training image generation unit, 63...Correct image generation unit, 70...Model learning processing unit, 81...Endoscope system, 82...Light source unit, 83...Imaging device, 84 Memory, 85...First imaging system, 91...Endoscope system, 92...Endoscope, 93...Imaging device, 94...Endoscope image processing device, 94a...Processor, 94b...Memory, 100...Object, S100...Model creation process, S110...Ground truth image reduction process, S120...Imaging system simulation process, S130...Training image reduction process, S140...Pixel count difference correction process, S150...Learning process, S200...Trained model inference process, M...Trained model

Claims

1. A ground truth image is obtained by applying a first reduction process to the original image, which is an image of a predetermined subject captured by the first imaging system, and The original image is subjected to a second reduction process to reduce the size of the image, and a training image is obtained by applying a blurring process to correct the difference in contrast due to the difference in the number of pixels between the original image and the processing target image captured by the second imaging system. A pre-trained model characterized by being machine-learned using [a specific method / tool].

2. The second reduction process for reducing the aforementioned image is performed on the first intermediate image, which is the original image to which an imaging system simulation process has been further applied. The trained model according to claim 1, characterized in that the first imaging system that captures the original image has a higher resolution than the second imaging system that captures the image to be processed.

3. The trained model according to claim 1, characterized in that the first imaging system for capturing the original image and the second imaging system for capturing the image to be processed are the same imaging system.

4. The trained model according to claim 1, characterized in that the first imaging system that captures the original image has a lower resolution than the second imaging system that captures the image to be processed.

5. A first ground truth image is obtained by applying a first reduction process to a first original image, which is obtained by capturing a predetermined subject with a first imaging system, and A first training image is obtained by applying a second reduction process to the first source image or a first intermediate image obtained by applying an imaging system simulation process to the first source image, and then applying a blur addition process to correct the difference in contrast due to the difference in the number of pixels with the processing target image captured by the third imaging system. A first training set consisting of the first correct image and the first training image paired together, A second ground truth image is obtained by applying a third reduction process to a second original image, which is obtained by capturing the predetermined subject with the second imaging system, and a third reduction process is applied to reduce the size of the image. A second training image is obtained by applying a fourth reduction process to the second source image or a second intermediate image obtained by applying the imaging system simulation process to the second source image, and by applying a blur addition process to correct the difference in contrast due to the difference in the number of pixels with the processing target image captured by the third imaging system, A second training set consisting of the second correct image and the second training image paired together, A trained model characterized by being machine-learned using a plurality of training sets, including the first training set and the second training set.

6. A ground truth image is obtained by applying a first reduction process to the original image of a predetermined subject captured by a first imaging system, and then applying a second reduction process with a different aspect ratio to the image that has undergone the first reduction process. A third reduction process is applied to the original image to reduce its size, a blurring process is applied to correct the difference in contrast due to the difference in pixel count with the processing target image captured by the second imaging system, and a training image is obtained by applying a fourth reduction process with a different aspect ratio to the image that has undergone the blurring process. A pre-trained model characterized by being machine-learned using [a specific method / tool].

7. A ground truth image generation unit performs a first reduction process on the original image, which is an image of a predetermined subject captured by a first imaging system, to reduce the size of the image. An information processing apparatus comprising: a training image generation unit that applies a second reduction process to the original image to reduce its size, and a blurring process to correct the difference in contrast due to the difference in the number of pixels between the original image and the processing target image captured by a second imaging system; and a training image generation unit.

Citation Information

Patent Citations

  • Medical image processing device, medical image processing method, and program

    JP2022070035A