Image processing device, its operating method, and endoscope system

The image processing device enhances super-resolution by degrading source images to create learning inputs, using a specialized learning model with multiple layers to output high-resolution images, addressing the challenge of generating high-quality images without high-quality source images.

JP7821701B2Active Publication Date: 2026-02-27FUJIFILM CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022134466
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-25
Publication Date
2026-02-27
Estimated Expiration
2042-08-25

AI Technical Summary

Technical Problem

Existing deep learning methods struggle to generate high-resolution images when high-quality source images are not available, leading to inadequate super-resolution results that fail to reflect the characteristics of degraded input images.

Method used

An image processing device that includes a processor performing degradation processing on source images to generate learning input images, using a learning model with specific layers to output high-resolution images, and updating the model based on calculated losses to enhance the super-resolution capability.

Benefits of technology

Accurately generates images with resolutions higher than those used for machine learning training, ensuring the super-resolution images reflect the characteristics of the input images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007821701000001
    Figure 0007821701000001
  • Figure 0007821701000002
    Figure 0007821701000002
  • Figure 0007821701000003
    Figure 0007821701000003
Patent Text Reader

Abstract

To provide an image processing device, an actuation method thereof, and an endoscope system capable of accurately performing super-resolution to generate an image with a resolution equal to or higher than that of an image used for learning in machine learning.SOLUTION: A processor generates a trained model by updating a learning model that takes as input an input image for learning obtained by performing degradation processing on a source image and outputs a first output image for learning and a second output image for learning that have a larger number of pixels than the input image for leaning. A second intermediate layer of the learning model outputs a feature map to be input to a first output layer that outputs the first output image for learning based on feature maps from a folded layer and a first intermediate layer. A third intermediate layer outputs a feature map to be input to a second output layer that outputs the second output image for learning based on the feature map from the folded layer.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device that performs super-resolution on an image, an operating method thereof, and an endoscope system. [Background technology]

[0002] Deep learning is sometimes used to accurately perform super-resolution to increase the resolution of unknown images. For example, Patent Document 1 describes a neural network that receives a low-resolution image and a classification score as input and outputs a high-resolution image. Patent Document 1 also describes that the low-resolution image used as learning data to be input to the neural network may be generated by downsampling a high-resolution image. It also describes that an image captured by a camera with a small number of pixels may be used as the low-resolution image, and an image captured of the same subject by a camera with a large number of pixels may be used as the high-resolution image. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2020-24612 Summary of the Invention [Problem to be solved by the invention]

[0004] In order to train deep learning to generate high-resolution images, it is desirable to prepare high-resolution images as source images for learning. However, depending on the specifications of the device that captures or generates the source images, it may not be possible to prepare high-quality source images for learning. Obtaining a super-resolution image with a higher resolution than the source image using deep learning is particularly difficult when such a source image is not available. Furthermore, when a high-resolution source image is not available, the super-resolution image generated by deep learning may not adequately reflect the characteristics of the degraded image input to deep learning before super-resolution processing.

[0005] The present invention aims to provide an image processing device, an operating method thereof, and an endoscopic system that can accurately perform super-resolution to generate images with a resolution higher than that of images used for machine learning training. [Means for solving the problem]

[0006] The image processing device of the present invention includes a processor, which acquires a source image, performs degradation processing on the source image to generate a learning input image, inputs the learning input image to a learning model to output a first learning output image and a second learning output image having a larger number of pixels than the learning input image, calculates a first loss based on the first learning output image and the source image, calculates a second loss based on the second learning output image and the source image, and updates the learning model based on the first loss and the second loss to generate a trained model, the learning model having an input layer, a first hidden layer, a folding layer, a second hidden layer, a first output layer, a third hidden layer, and a second output layer, and the input layer receives the learning input image and calculates a feature to be input to the first hidden layer. The first intermediate layer receives the feature map output by the input layer and outputs a feature map to be input to the folding layer and the second intermediate layer. The folding layer outputs a feature map to be input to the second and third intermediate layers based on the feature map input from the first intermediate layer. The second intermediate layer outputs a feature map to be input to the first output layer based on the feature map input from the folding layer and the feature map input from the first intermediate layer. The first output layer outputs a first learning output image based on the feature map input from the second intermediate layer. The third intermediate layer outputs a feature map to be input to the second output layer based on the feature map input from the folding layer. The second output layer outputs a second learning output image based on the feature map input from the third intermediate layer.

[0007] The amount of information in the training input image, the first training output image, the second training output image, and the feature map is determined by the number of pixels, the number of channels, and the number of bits according to the data type, and it is preferable that the folding layer reduces the amount of information in the feature map output by the folding layer to be smaller than the amount of information in the training input image by an information amount reduction process that changes the number of channels or the number of bits.

[0008] The information amount reduction process is preferably a process of reducing the number of channels, and is preferably a process of reducing the number of bits.

[0009] The first intermediate layer preferably performs processing to reduce the number of pixels in the feature map input from the input layer.

[0010] The second and third hidden layers preferably perform processing to increase the number of pixels in the feature map input from the folding layer.

[0011] The degradation process preferably includes a process of reducing the number of pixels of the source image. The degradation process preferably includes a filtering process and / or a noise addition process. The processor preferably further inputs the training input image to the second intermediate layer.

[0012] The processor inputs an inference input image having a first number of pixels into the trained model, and outputs a super-resolution image having a second number of pixels greater than the first number of pixels, and it is preferable that the ratio of the second number of pixels to the first number of pixels is equal to the ratio of the number of pixels of the first training output image and the second training output image to the number of pixels of the training input image.

[0013] A method for operating an image processing device of the present invention includes the steps of acquiring a source image, performing degradation processing on the source image to generate a learning input image, inputting the learning input image into a learning model to output a first learning output image and a second learning output image having a larger number of pixels than the learning input image, calculating a first loss based on the first learning output image and the source image, calculating a second loss based on the second learning output image and the source image, and updating the learning model based on the first loss and the second loss to generate a trained model, wherein the learning model has an input layer, a first hidden layer, a folding layer, a second hidden layer, a first output layer, a third hidden layer, and a second output layer, and the input layer is configured to receive the learning input image as input. The first intermediate layer receives the feature map output by the input layer and outputs a feature map to be input to the folding layer and the second intermediate layer. The folding layer outputs a feature map to be input to the second and third intermediate layers based on the feature map input from the first intermediate layer. The second intermediate layer outputs a feature map to be input to the first output layer based on the feature map input from the folding layer and the feature map input from the first intermediate layer. The first output layer outputs a first learning output image based on the feature map input from the second intermediate layer. The third intermediate layer outputs a feature map to be input to the second output layer based on the feature map input from the folding layer. The second output layer outputs a second learning output image based on the feature map input from the third intermediate layer.

[0014] The image processing device of the present invention includes a processor, which acquires a source image and performs degradation processing on the source image to generate a training input image, inputs the training input image to a generator to output a first training output image and a second training output image having a larger number of pixels than the training input image, inputs the first training output image and the second training output image to a classifier to output a first classification result based on the first training output image and a second classification result based on the second training output image, calculates a first classifier loss based on the first classification result and a second classifier loss based on the second classification result, updates the classifier based on the first classifier loss and the second classifier loss, calculates a first generator loss based on the first classification result and the source image and a second generator loss based on the second classification result and the source image, and updates the generator based on the first generator loss and the second generator loss to generate a trained generator, the generator includes an input layer, a first The neural network has an intermediate layer, a folded layer, a second intermediate layer, a first output layer, a third intermediate layer, and a second output layer. The input layer receives a learning input image as input and outputs a feature map to be input to the first intermediate layer. The first intermediate layer receives a feature map output by the input layer as input and outputs a feature map to be input to the folded layer and the second intermediate layer. The folded layer outputs a feature map to be input to the second and third intermediate layers based on the feature map input from the first intermediate layer. The second intermediate layer outputs a feature map to be input to the first output layer based on the feature map input from the folded layer and the feature map input from the first intermediate layer. The first output layer outputs a first learning output image based on the feature map input from the second intermediate layer. The third intermediate layer outputs a feature map to be input to the second output layer based on the feature map input from the folded layer. The second output layer outputs a second learning output image based on the feature map input from the third intermediate layer.

[0015] The image processing device of the present invention is an image processing device including a processor, An endoscopic image having a first number of pixels is acquired, and by inputting the endoscopic image into a trained model, a super-resolution image having a second number of pixels greater than the first number of pixels is output. The trained model is generated by using a training input image having a fourth number of pixels smaller than the third number of pixels generated by degradation processing a source endoscopic image having a third number of pixels equal to or less than the first number of pixels, and updating the training model to output a first training output image and a second training output image having a fifth number of pixels greater than the fourth number of pixels based on a first loss based on the first training output image and the source endoscopic image and a second loss based on the second training output image and the source endoscopic image.

[0016] The processor preferably controls the display of the super-resolution image and information indicating that the endoscopic image has been subjected to high-resolution processing.

[0017] Preferably, the ratio of the second number of pixels to the first number of pixels is equal to the ratio of the fifth number of pixels to the fourth number of pixels. Preferably, the third number of pixels is equal to the first number of pixels.

[0018] The endoscopic system of the present invention comprises the above-mentioned image processing device, an endoscope that generates an endoscopic image by photographing a subject, and a display, and the processor controls the display to display a super-resolution image. [Effects of the Invention]

[0019] According to the present invention, it is possible to perform super-resolution with high accuracy to generate images having a resolution equal to or higher than that of images used for machine learning training. [Brief explanation of the drawings]

[0020] [Figure 1] FIG. 2 is a block diagram showing functions of the image processing device. [Figure 2] FIG. 2 is a block diagram showing the functions of a learning unit. [Figure 3] FIG. 2 is a block diagram showing the functions of an inference unit. [Figure 4]FIG. 1 is a block diagram showing the functions of a learning model. [Figure 5] FIG. 2 is an explanatory diagram showing the functions of the first intermediate layer, the second intermediate layer, and the third intermediate layer. [Figure 6] FIG. 10 is an explanatory diagram showing an information amount reduction process. [Figure 7] FIG. 10 is a block diagram showing the functions of a learning unit when GAN is applied. [Figure 8] FIG. 10 is an explanatory diagram showing the function of the inference unit when the input image for inference is an endoscopic image. [Figure 9] FIG. 10 is an explanatory diagram showing the function of the learning unit when the input image for inference is an endoscopic image. [Figure 10] FIG. 1 is a block diagram showing the functions of a trained model. [Figure 11] 10A and 10B are image diagrams showing examples of when a super-resolution image and a notification display are displayed. [Figure 12] 1 is a flowchart showing the flow of functions according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0021] 1, the image processing device 10 includes a learning unit 11 and an inference unit 12. The learning unit 11 uses a source image for learning input to the image processing device 10 from a database 20 to optimize parameters of a learning model 100 to which machine learning is applied, thereby generating a trained model 200.

[0022] Machine learning applied to the learning model 100 includes decision trees, support vector machines, random forests, regression analysis, deep learning, reinforcement learning, deep reinforcement learning, neural networks, convolutional neural networks, generative adversarial networks, etc. The specific configuration of the learning model 100 will be described later.

[0023] The inference unit 12 generates a super-resolution image that has the characteristics of the unknown image and has a higher resolution than the unknown image by inputting an input image for inference, which is an unknown image different from the source image and transmitted from the database 20, to the trained model 200 generated by the learning unit 11. As will be described in detail later, the number of pixels of the input image for inference input to the trained model 200 is equal to or greater than the number of pixels of the source image used for training.

[0024] The database 20 stores source images used for training the learning model 100 and inference input images that are used as original images for generating super-resolution images through inference by the trained model. The database 20 is a storage, file server, cloud storage, or the like that stores images. The database 20 may be part of a system that directly or indirectly cooperates with the image processing device 10, such as a hospital information system (so-called HIS (Hospital Information Systems)) or a PACS (Picture Archiving and Communication Systems).

[0025] The source images or inference input images stored in the database 20 are transmitted from the modality 30. The image processing device 10 of this embodiment is suitable when the modality 30 is a medical image generating device that generates medical images, such as an endoscope, a radiographic imaging device, or an ultrasound imaging device. The medical images are endoscopic images, radiographic images, ultrasound images, etc. The image processing device 10 of this embodiment is particularly suitable when the modality 30 is an endoscope, and the source images and inference input images are endoscopic images. An example in which the source images and inference input images are endoscopic images will be described in detail later.

[0026] The image processing device 10, the database 20, and the modality 30 are connected to each other via wired or wireless communication. Wireless connection includes connection via a network, such as the Internet or a LAN (Local Area Network).

[0027] As shown in Fig. 2, the learning unit 11 has a degradation processing unit 40, a learning model 100, an evaluation unit 50, and an update unit 60. Also, as shown in Fig. 3, the inference unit 12 has a trained model 200 generated by training the learning model 100, and a display control unit 70. In the image processing device 10, programs related to various processes are incorporated into a program memory (not shown). A control unit (not shown) configured by a processor executes the programs in the program memory, thereby realizing the functions of the degradation processing unit 40, the learning model 100, the evaluation unit 50, and the update unit 60 of the learning unit 11, and the trained model 200 and the display control unit 70 of the inference unit 12.

[0028] The image processing device 10 may be configured so that the learning unit 11 and the inference unit 12 are separated and provided in different devices, and they communicate with each other. Also, the image processing device 10 may be configured so that the components of the learning unit 11 and the inference unit 12 are separated and provided in different devices, and they communicate with each other. In this case, each device is provided with a control unit constituted by a processor.

[0029] As shown in FIG. 2, the database 20 inputs a source image 21 to the degradation processing unit 40 of the learning unit 11. The degradation processing unit 40 performs degradation processing on the source image 21 to generate a learning input image 41 with a smaller number of pixels than the source image 21. The degradation processing is a resolution reduction process on the source image 21 that reduces the number of pixels in the source image 21. The number of pixels is the number of pixels contained in an image, represented by the width and height of the image, and is also called the number of pixels. In this specification, the term "number of pixels" is used to mean "resolution." The greater the number of pixels an image has, the greater the image resolution, allowing complex structures to be represented in detail. On the other hand, the smaller the number of pixels an image has, the lower the resolution, resulting in a rough image with blurred contours.

[0030] The degradation process includes filtering and / or noise addition to the source image 21. The filtering process is a process in which a filter such as a Gaussian filter, an averaging filter, a median filter, or a bilateral filter is applied to the source image to degrade the source image 21 by blurring or shrinking. The noise addition process is a process in which noise is added to degrade the source image 21 by randomly setting the pixel values ​​of the pixels in the source image 21 to their maximum or minimum values, changing the brightness of the pixels in the source image 21 using random numbers, or the like.

[0031] The training input image 41 obtained by the degradation process is input to the training model 100. The training model 100 performs feature extraction and high-resolution processing on the training input image 41, and outputs a first training output image 101 and a second training output image 102. Because the training model 100 performs high-resolution processing on the training input image 41, the first training output image 101 and the second training output image 102 have a larger number of pixels than the training input image 41. The training model 100 has one input layer, two intermediate layers, and two output layers. Details of the two intermediate layers and two output layers of the training model 100 in this embodiment will be described later.

[0032] The evaluation unit 50 applies the first learning output image 101 and the source image 21 to a loss function, which is a model for evaluation, to calculate a first loss 51. It also applies the second learning output image 102 and the source image 21 to the loss function to calculate a second loss 52. It is preferable to use mean squared error (MSE) to calculate the first loss 51 and the second loss 52. The smaller the first loss 51, the smaller the difference between the first learning output image 101 and the source image 21. The smaller the second loss 52, and the smaller the difference between the second learning output image 102 and the source image 21. The closer the loss is to "0," the higher the output accuracy of the learning model 100. Hereinafter, the term "loss" will be used to refer to either or both of the first loss 51 and the second loss 52.

[0033] The update unit 60 sets and updates the parameters of the learning model 100 so that the loss approaches "0" (minimizes). The calculation of the loss by the evaluation unit 50 and the update of the parameters by the update unit 60 are repeated until the first loss 51 and the second loss 52 reach preset values. The value for instructing the end of the loss calculation and parameter update may be a value within a certain range, or may be greater than or less than a certain threshold. Learning the learning model 100, i.e., generating the trained model 200, refers to a parameter optimization process for minimizing the loss. The optimized parameters are used as parameters of the trained model 200 in the inference unit 12. The image processing device 10 may be provided with a parameter storage memory (not shown) for storing the parameters.

[0034] As shown in Fig. 3, a trained model 200 generated by training the training model 100 performs feature extraction and high-resolution processing on an input image for inference 201 sent from a database 20, and outputs a super-resolution image 202. Note that super-resolution refers to high-resolution processing that generates a high-resolution image from a low-resolution image, which is an input signal.

[0035] The super-resolution image 202 output by the trained model 200 is transmitted to the display control unit 70. The display control unit 70 performs signal processing for displaying the super-resolution image 202 on the display 80, and controls the display of the super-resolution image 202 on the display 80.

[0036] The configuration of the learning model 100 will be described below with reference to Fig. 4. It is preferable that the learning model 100 is a convolutional neural network (CNN).

[0037] As shown in FIG. 4, the learning model 100 includes an input layer 110, a first hidden layer 120, a folding layer 130, a second hidden layer 140, a first output layer 150, a third hidden layer 160, and a second output layer 170.

[0038] The network configured with the input layer 110, the first hidden layer 120, and the folding layer 130 is a network that extracts features from the training input image 41, and is a network that corresponds to the encoder of the model having an encoder-decoder structure.

[0039] The input layer 110 receives training input images 41 from the degradation processor 40 and outputs a feature map 111 to be input to the first hidden layer 120. The training input images 41 are preferably integer data types, and the feature map 111 is preferably floating-point data types. The data type of the training input images 41 may be converted from integer to floating-point data by the degradation processor 40, or may be converted at a stage from the degradation processor 40 to the input to the input layer 110, or may be converted at a stage from the input layer 110 to the input to a convolutional layer, which will be described later.

[0040] The first hidden layer 120 receives the feature map 111 output by the input layer 110, and outputs a feature map 120a to be input to the folding layer 130 and the second hidden layer 140. The first hidden layer 120 performs convolution and / or pooling on the feature map 111, and outputs the feature map 120a obtained by extracting the features of the training input image 41.

[0041] Convolution is a process of applying a filter to input image data and extracting (outputting) a feature map that indicates the position of the filter's pattern in the input image data. The filter is also called a convolution kernel or simply a kernel. The number of pixels in the feature map extracted by convolution can be made the same as the number of pixels in the input feature map or can be made smaller by setting padding that interpolates pixel values ​​around the image data and the interval (stride) at which the filter is applied to the image data. In this specification, the term "feature map" is also used to mean "feature amount."

[0042] Pooling is a process that summarizes the values ​​of local regions belonging to each feature map and reduces the number of pixels in the feature map, which is image data. A local region is a region consisting of multiple pixels centered around one pixel in the feature map. Pooling methods include max pooling and average pooling. Max pooling is a process that selects the maximum pixel value among the pixel values ​​of the pixels included in the local region and sets it as the pixel value of the pixel in the output feature map. Average pooling is a process that selects the average pixel value of the pixels included in the local region and sets it as the pixel value of the pixel in the output feature map. The process of reducing the number of pixels in a feature map by convolution or pooling is also called downsampling. The first hidden layer 120 preferably reduces the number of pixels in the feature map input from the input layer 110 by downsampling.

[0043] The folding layer 130 outputs a feature map 131 to be input to the second hidden layer 140 and the third hidden layer 160 based on the feature map 120a input from the first hidden layer 120. Similar to the first hidden layer 120, the folding layer 130 performs convolution and / or pooling on the feature map 120a.

[0044] The number of pixels in the feature map 131 output from the folding layer 130 may be the same as that of the training input image 41, but it is preferable that the number of pixels in the feature map 131 be smaller than that of the training input image 41 in order to speed up the output processing in training and inference.

[0045] The network configured with the second hidden layer 140 and the first output layer 150 performs a high-resolution process on the feature map 131 having the features of the training input image 41, which is output from the folding layer 130, by further using the feature map 120a output from the first hidden layer 120, and outputs a first training output image 101 having a larger number of pixels than the training input image 41. The network configured with the second hidden layer 140 and the first output layer 150 corresponds to the decoder in the model of the encoder-decoder structure.

[0046] As shown in Fig. 4, the learning model 100 of this embodiment has two decoders. The network consisting of the second hidden layer 140 and the first output layer 150 is the first decoder. The network consisting of the third hidden layer 160 and the second output layer 170 (described later) is the second decoder.

[0047] The second hidden layer 140 outputs a feature map 140a to be input to the first output layer 150 based on the feature map 131 input from the folding layer 130 and the feature map 120a input from the first hidden layer 120. High-resolution processing includes upsampling, which arranges pixel values ​​of pixels constituting the feature map at intervals of several pixels and interpolates the values ​​of the pixels between them, and upconvolution, which combines upsampling without interpolating pixel values ​​with convolution. Upsampling is also called unpooling, and upconvolution is also called transposed convolution or deconvolution. The second hidden layer 140 and the third hidden layer 160 (described later) perform processing to increase the number of pixels in the feature map 131 input from the folding layer 130 by upsampling or upconvolution.

[0048] In the second hidden layer 140, the feature map 120a output from the first hidden layer 120 is skip-connected. By the skip connection, the second hidden layer 140 outputs a feature map obtained by performing high-resolution processing on the feature map 131 and the feature map 120a as the feature map 140a.

[0049] That is, when focusing on an encoder composed of an input layer 110, a first hidden layer 120, and a folding layer 130, and a first decoder composed of a second hidden layer 140 and a first output layer 150, a network called a U-net is formed in which the first hidden layer 120 and the second hidden layer 140 have a bilaterally symmetrical shape. It is generally known that a U-net can share features output from the hidden layer of the encoder with the decoder by connecting the encoder layer with the corresponding decoder layer, thereby enabling the decoder to output with extremely high accuracy.

[0050] In the learning model 100 of this embodiment, the feature map 120a output from the first hidden layer 120 is connected to the second hidden layer 140, thereby enabling efficient and highly accurate resolution enhancement of the feature map 131 having the features of the learning input image 41. The folding layer 130 is the last layer in the encoder and corresponds to the bottom of the U-shape of the "folding" in U-net.

[0051] The first output layer 150 outputs a first learning output image 101 based on the feature map 140a input from the second hidden layer 140. The first output layer 150 outputs the first learning output image 101 by applying an activation function such as a ReLU (Rectified Linear Unit) function to the feature map 140a. The activation function is also applied to the feature maps convolved in the first hidden layer 120, the second hidden layer 140, and the third hidden layer 160. The first learning output image 101 output from the first output layer 150 is transmitted to the evaluation unit 50.

[0052] Preferably, the data type of the feature map 140a is floating point, and the data type of the first training output image 101 is integer. The conversion of the data type of the first training output image 101 from floating point to integer may be performed in the first output layer 150, or may be performed at a stage from the first output layer 150 to the stage before input to the evaluation unit 50.

[0053] The network, which is a second decoder and is composed of a third hidden layer 160 and a second output layer 170, performs high-resolution processing on the feature map 131 output from the folding layer 130, and outputs a second training output image 102 having a larger number of pixels than the training input image 41.

[0054] In the second decoder, the third hidden layer 160 outputs a feature map 160a to be input to the second output layer 170 based on the feature map 131. Unlike the second hidden layer 140, the third hidden layer 160 does not perform skip connection of the feature map 120a output from the first hidden layer 120.

[0055] The second output layer 170 outputs the second learning output image 102 based on the feature map 160a input from the third hidden layer 160. The second output layer 170 applies an activation function to the feature map 160a to output the second learning output image 102. The second learning output image 102 is transmitted to the evaluation unit 50.

[0056] Preferably, the data type of the feature map 160a is floating point, and the data type of the second learning output image 102 is integer. The conversion of the data type of the second learning output image 102 from floating point to integer may be performed in the second output layer 170, or may be performed at a stage from the second output layer 170 to the input to the evaluation unit 50.

[0057] The evaluation unit 50 calculates a first loss 51 by applying a loss function to the first learning output image 101 output from the first output layer 150 and the source image 21 transmitted from the database 20 and comparing them. The evaluation unit 50 also calculates a second loss 52 by applying a loss function to the second learning output image 102 output from the second output layer 170 and the source image 21 transmitted from the database 20 and comparing them. The loss is used in the parameter optimization process by the update unit 60, as described above.

[0058] As described above, by configuring and training the learning model 100 to include not only the first decoder to which the feature map 120a from the encoder is connected, but also the second decoder to which the feature map 120a from the encoder is not connected, it is possible to update the encoder parameters so that the encoder can extract more important features of the training input images 41. When the encoder parameters are updated in this way, the parameters of the first decoder and the second decoder are also updated so as to output images that have been subjected to high-resolution processing and that more strongly reflect the features of the training input images 41.

[0059] The configuration and functions of the learning model 100 will be described in more detail with reference to FIG. 5. As shown in FIG. 5, the first hidden layer 120 has multiple convolution layers 121, 122, 123, and 124 that perform convolution. The second hidden layer 140 has multiple upsampling layers 141, 142, 143, and 144 that perform high-resolution processing. The third hidden layer 160 also has multiple upsampling layers 161, 162, 163, and 164. Although not shown in FIG. 5, the second hidden layer 140 and the third hidden layer 160 preferably have convolution layers downstream of each upsampling layer. Furthermore, the first hidden layer 120 may have pooling layers downstream of each convolution layer.

[0060] 5 illustrates a network in which the first hidden layer 120, the second hidden layer 140, and the third hidden layer 160 each have four convolutional layers or upsampling layers, but the number of convolutional layers or upsampling layers included in each of the first hidden layer 120, the second hidden layer 140, and the third hidden layer 160 is not limited to this. It is preferable that the number of upsampling layers in the third hidden layer 160 be the same as the number of upsampling layers in the second hidden layer 140.

[0061] The amount of information of each image data input and output in the learning model 100 will be described below. The amount of information of image data is the number of elements or memory capacity. The number of elements is determined by the number of pixels of each image included in the image data and the number of channels of the image data. As a specific example, the number of elements is (number of pixels) x (number of channels). The amount of memory is determined by the number of pixels of each image included in the image data, the number of channels of the image data, and the number of bits according to the data type of each image. As a specific example, the amount of memory is (number of pixels) x (number of channels) x (number of bits).

[0062] Integer types and floating-point types are data types used to process image data. Integer types include byte, short, int, and long, while floating-point types include float and double. Integer types include unsigned integers, which represent positive integers, and signed integers, which can represent both positive and negative integers. The number of bits corresponding to a data type varies depending on the programming language; for example, an "unsigned int8" data type is 8 bits, and a "float32" data type is 32 bits.

[0063] In FIG. 5, a specific example of the number of elements is attached to the right of each image data. For example, the number of elements of the source image 21 sent from the database 20 to the degradation processing unit 40 and the evaluation unit 50 is "1024 x 1024 x 3." This indicates that the number of pixels (width x height) is "1024 x 1024" and the number of channels is "3." In FIG. 5, an example is shown in which the source image 21 is a color image, and therefore the number of channels is "3," indicating that there are three types: R channel, G channel, and B channel. Note that if the source image 21 is a monochrome image, the number of channels is "1."

[0064] The number of elements in the training input image 41 generated by performing degradation processing in the degradation processing unit 40 is "512 x 512 x 3," which is a smaller number of pixels than the source image 21. The number of elements in the feature map 111 input from the input layer 110 to the convolutional layer 121 of the first hidden layer 120 is "512 x 512 x 3." Here, if the data type of the feature map 111 is unsigned int8, the memory size of the feature map 111 is 6,291,456 bits (512 x 512 x 3 x 8). If the data type of the feature map 111 is converted from unsigned int8 to float32, the memory size of the feature map 111 becomes 25,165,824 bits (512 x 512 x 3 x 32).

[0065] Convolutional layers 121, 122, 123, and 124 and folding layer 130 included in the encoder perform convolution of feature maps in stages to obtain feature maps 121a, 122a, 123a, 124a, and 131. In the example shown in Fig. 5, the later the convolutional layer, the smaller the number of extracted pixels in the output feature map.

[0066] Specifically, the feature map 121a has 512 x 512 x 64 elements and a memory capacity of 536,870,912 bits (512 x 512 x 64 x 32). The feature map 122a has 256 x 256 x 128 elements and a memory capacity of 268,435,456 bits (256 x 256 x 128 x 32). The feature map 123a has 128 x 128 x 256 elements and a memory capacity of 134,217,728 bits (128 x 128 x 256 x 32). The feature map 124a has 64 x 64 x 512 elements and a memory capacity of 67,108,864 bits (64 x 64 x 512 x 32). The number of elements in the feature map 131 is "32 x 32 x 1024", and the memory capacity is 33,554,432 bits (32 x 32 x 1024 x 32).

[0067] In general, in feature extraction in a convolutional neural network, the number of channels in the feature map output from each convolutional layer is gradually increased to maintain the amount of information, and the number of channels in the output feature map corresponds to the number of filters used in each convolutional layer.

[0068] 5, the convolutional layer 124 applies 512 filters to the feature map 123a, resulting in an output feature map 124a with 512 channels. The folding layer 130, which performs feature extraction at the final stage of the encoder, applies 1024 filters to the feature map 124a, resulting in an output feature map 131 with 1024 channels.

[0069] In the encoder and first decoder, which are U-nets, feature maps 121a, 122a, 123a, and 124a output from the convolutional layers 121, 122, 123, and 124 of the first hidden layer are input to upsampling layers 141, 142, 143, and 144 of the second hidden layer 140, respectively, via skip connections.

[0070] In the upsampling layers 141, 142, 143, and 144 included in the first decoder, the resolution of the feature map is increased stepwise to obtain feature maps 141a, 142a, 143a, and 144a. In the example shown in Fig. 5, the later the upsampling layer, the higher the resolution of the output feature map, with a larger number of pixels.

[0071] The upsampling layer 141 receives the feature map 131 from the folding layer 130 and the feature map 124a output from the corresponding convolutional layer 124, and outputs the feature map 141a. The number of elements in the feature map 141a is 128×128×256, and the memory size is 134,217,728 bits (128×128×256×32).

[0072] Upsampling layer 142 receives feature map 141a and feature map 123a output from corresponding convolution layer 123, and outputs feature map 142a. Feature map 142a has 256×256×128 elements and a memory size of 268,435,456 bits (256×256×128×32).

[0073] The upsampling layer 143 receives the feature map 142a and the feature map 122a output from the corresponding convolutional layer 122, and outputs the feature map 143a. The feature map 143a has 512 x 512 x 64 elements and a memory capacity of 536,870,912 bits (512 x 512 x 64 x 32).

[0074] Upsampling layer 144 receives feature map 143a and feature map 121a output from the corresponding convolutional layer 121, and outputs feature map 144a. The number of elements in feature map 144a is 1024×1024×32, and the memory size is 1,073,741,824 bits (1024×1024×32×32).

[0075] Note that the training input images 41 may be further input to the upsampling layer 144, which is the upsampling layer immediately preceding the first output layer 150. By using the training input images 41 in the resolution enhancement process, it is possible to increase the accuracy of the output of the first decoder. Note that when the training input images 41 are input to the upsampling layer 144, the data type of the training input images 41 needs to be converted to the same data type as the feature map 143a before being input.

[0076] The first output layer 150 applies an activation function to the feature map 144a and outputs a first training output image 101. The number of elements in the first training output image 101 is 1024 × 1024 × 3. If the data type of the first training output image 101 is float32, the memory size of the first training output image 101 is 100,663,296 bits (1024 × 1024 × 3 × 32). If the data type of the first training output image 101 is converted from float32 to unsigned int8, the memory size of the first training output image 101 becomes 25,165,824 bits (1024 × 1024 × 3 × 8).

[0077] Although FIG. 5 shows an example in which the feature maps are connected to corresponding layers between the encoder and the first decoder, the connections from the encoder are not limited to the corresponding layers.

[0078] The second decoder does not have a skip connection from the encoder, and performs a high-resolution process on feature map 131 from folding layer 130 in stages using upsampling layers 161, 162, 163, and 164 included in the second decoder, thereby obtaining feature maps 161a, 162a, 163a, and 164a. In the example shown in Figure 5, the later the upsampling layer, the higher the resolution of the output feature map, resulting in a larger number of pixels.

[0079] Specifically, the feature map 161a has 128 x 128 x 256 elements and a memory capacity of 134,217,728 bits (128 x 128 x 256 x 32). The feature map 162a has 256 x 256 x 128 elements and a memory capacity of 268,435,456 bits (256 x 256 x 128 x 32). The feature map 163a has 512 x 512 x 64 elements and a memory capacity of 536,870,912 bits (512 x 512 x 64 x 32). The feature map 164a has 1024 x 1024 x 32 elements and a memory capacity of 1,073,741,824 bits (1024 x 1024 x 32 x 32).

[0080] The second output layer 170 applies an activation function to the feature map 164a and outputs the second training output image 102. The number of elements in the second training output image 102 is 1024 × 1024 × 3. If the data type of the second training output image 102 is float32, the memory size of the second training output image 102 is 100,663,296 bits (1024 × 1024 × 3 × 32). If the data type of the second training output image 102 is converted from float32 to unsigned int8, the memory size of the second training output image 102 becomes 25,165,824 bits (1024 × 1024 × 3 × 8). The number of pixels in the second training output image 102 output from the second decoder is preferably the same as the number of pixels in the first training output image 101 output from the first decoder.

[0081] The evaluation unit 50 calculates a first loss 51 and a second loss 52 using a source image 21, a first learning output image 101, and a second learning output image 102, each having the number of elements of "1024×1024×3".

[0082] As described above, by inputting a training input image 41 having a pixel count of "512 x 512" into the training model 100, a first training output image 101 and a second training output image 102 having a pixel count of "1024 x 1024" are output. That is, the training model 100 in the example shown in FIG. 5 is a training model that outputs an image with a resolution increased by four times the pixel count of the input image. Therefore, the trained model 200 generated by training the training model 100 shown in FIG. 5 outputs a super-resolution image having four times the pixel count of the input unknown image (input image for inference). For example, by inputting an input image for inference having a pixel count of "1024 x 1024" as an unknown image into the trained model 200 generated by training the training model 100 shown in FIG. 5, a super-resolution image having a pixel count of "2048 x 2048" can be obtained.

[0083] The number of pixels of the source image 21, the learning input image 41, the first learning output image 101, and the second learning output image 102 are not limited to the above examples. For example, the learning model 100 may be designed to output an image with a high resolution by increasing the number of pixels of the input image by 16 times, 64 times, 256 times, etc.

[0084] That is, in the inference unit 12, an inference input image 201 having a first number of pixels is input to a trained model 200, and a super-resolution image 202 having a second number of pixels greater than the first number of pixels is output (see FIG. 3). This trained model 200 is generated by using a source image 21 having a third number of pixels and a training input image 41 having a fourth number of pixels smaller than the third number of pixels, which is generated by performing degradation processing on the source image 21, and updating parameters of a training model 100 that outputs a first training output image 101 and a second training output image 102 having a fifth number of pixels greater than the fourth number of pixels.

[0085] Here, the third number of pixels in the source image 21 is equal to or less than the first number of pixels in the inference input image 201. The ratio of the second number of pixels in the super-resolution image 202 to the first number of pixels in the inference input image 201 is equal to the ratio of the fifth number of pixels in the first learning output image 101 and the second learning output image 102 to the fourth number of pixels in the learning input image 41. The ratio can be, for example, 4 times, 16 times, 64 times, 256 times, etc. 2n times (n is a natural number greater than or equal to 1).

[0086] As described above, in addition to the U-net structure, the learning model 100 is provided with a second decoder that branches off from the folding layer 130 and generates an image that has been subjected to high-resolution processing without receiving features from the encoder. This allows the learning model 100 to be trained to generate an output image that reflects particularly important features of a degraded image while enjoying the benefits of the U-net, which can be trained with high accuracy and efficiency.

[0087] In image processing using CNNs, feature extraction is often performed by gradually reducing the resolution of the input image through downsampling. In this case, the number of pixels in the feature map output from the later convolutional layers becomes smaller. For this reason, it is common to increase the number of channels in the feature map in order to maintain the amount of information in the feature map as a whole in later stages.

[0088] Even in super-resolution using U-net, the encoder maintains redundant information by gradually increasing the number of channels instead of gradually decreasing the resolution of the output feature map. High-precision super-resolution can be achieved by connecting the redundantly maintained feature map close to the original image to the decoder. A learning model consisting solely of U-net has the advantage of being able to achieve super-resolution that accurately restores the original image corresponding to the source image 21.

[0089] However, in "restoring the original image," it is difficult to achieve super-resolution beyond the number of pixels of the source image 21. In the case of a learning model consisting only of a U-net, the output first learning output image 101 is an image in which the typical features of the learning input image 41 are made high-resolution. Therefore, by further providing a second decoder that is not connected to the intermediate features of the encoder as in the U-net, it is possible to train an encoder and decoder that achieve super-resolution to generate a super-resolution image having a number of pixels greater than the number of pixels of the source image 21.

[0090] The second training output image 102 is an image output by using only the feature map 131 from the folding layer 130, from which the features of the training input image 41 are most strongly extracted, for the resolution enhancement process, without using image data from an intermediate stage where feature extraction is performed for the resolution enhancement process. By configuring the training model 100 as described above, the second loss 52 becomes large in the early stages of training. However, by proceeding with training to reduce the second loss 52, the accuracy of feature extraction in the encoder can be improved compared to the accuracy of feature extraction in an encoder of a training model consisting only of a U-net. This allows the training model 100 to output a high-resolution image that strongly reflects the features of the original image. As a result, during inference, it becomes possible to output a super-resolution image 202 having a number of pixels greater than the number of pixels of the source image 21 used during training.

[0091] As described above, by training the learning model 100 having the first decoder to which the feature map from the encoder is connected and the second decoder to which the feature map from the encoder is not connected, it becomes possible to obtain a high-resolution image that strongly reflects the features of the original image. Here, by performing an information amount reduction process that reduces the amount of information in the feature map 131 output from the folding layer 130, the accuracy of feature extraction by the encoder can be further improved.

[0092] The information amount reduction process is a process of changing the number of channels or bits of the information amount (number of elements or memory amount) determined by the number of pixels, the number of channels, and the number of bits according to the data type. It is preferable that the information amount of the feature map 131 output by the folding layer 130 is made smaller than the information amount of the training input image 41 by the information amount reduction process.

[0093] The information amount reduction process of changing the number of channels involves reducing the number of channels of the feature map 131 output by the folding layer 130. Specifically, the folding layer 130 reduces the number of channels of the input / output feature map before or after performing convolution or pooling.

[0094] For example, as shown in FIG. 6, the number of elements (1,048,576) of the feature map 131, which was "32×32×1024" in FIG. 5, is reduced to "32×32×128" (number of elements: 131,072). In the example shown in FIG. 6, if the data type of the feature map 131 is float32, the memory amount is reduced from 33,554,432 bits to 4,194,304 bits. In FIG. 6, the number of elements of the training input image 41 is "512×512×3" (number of elements: 786,432, memory amount: 6,291,456 bits (512×512×3×8)). Therefore, the number of elements of the feature map 131 after the information amount reduction process is smaller than that of the training input image 41.

[0095] The information amount reduction process of changing the number of bits involves reducing the number of bits according to the data type of the feature map 131 output by the folding layer 130. Specifically, the folding layer 130 converts the data type of the feature map before or after performing convolution or pooling, ultimately reducing the memory size of the feature map 131 output by the folding layer 130.

[0096] 6, if the data type of the feature map 131 output by the folding layer 130 is converted to float16, the memory size of the feature map 131 will be 2,097,152 bits (32×32×128×16). If the data type of the feature map 131 is converted to int8, the memory size of the feature map 131 will be 1,048,576 bits (32×32×128×8).

[0097] 5 or 6, if the data type of the feature map 131 is converted to a data type with a bit count of 2, the memory size of the feature map 131 becomes 2,097,152 bits (32×32×1024×8) or 1,048,576 bits (32×32×128×8). If the data type of the training input images 41 (and the feature map 111) is unsigned int8, the memory size of the training input images 41 is 6,291,456 bits (512×512×3×8). Therefore, the memory size of the feature map 131 that has undergone a combination of channel number reduction and data type conversion, or information amount reduction processing through data type conversion, is smaller than that of the training input images 41.

[0098] As described above, by reducing the amount of information in the feature map 131, which is the result of feature extraction of the training input image in the encoder, the second decoder generates the second training output image 102 from a feature map with less information. This allows the training model 100 to be updated to further improve the accuracy of feature extraction by the encoder, and also allows the training model 100 to be updated to make the first training output image 101 and the second training output image 102 output from the two decoders into images with higher resolution that more strongly reflect the features of the training input image 41.

[0099] The configuration for performing the information amount reduction process described above is particularly suitable when the number of pixels of an unknown image input to the trained model 200 exceeds the number of pixels of the source image 21. For example, the number of pixels of the unknown image, an inference input image 201, is 1280 × 960, and the number of pixels of the source image 21 is 512 × 512. In this case, the training model 100 is designed to output an image having 64 times the number of pixels of the training input image 41, and outputs a first training output image 101 and a second training output image 102 having a pixel count of 2048 × 2048 from the training input image 41 whose pixel count has been reduced to 256 × 256 by degradation processing. In this case, when the inference input image 201 having a pixel count of 1280 × 960 is input to the generated trained model 200, a super-resolution image 202 having a pixel count of 10240 × 7680 and a resolution of 8K or higher can be generated.

[0100] A Generative Adversarial Network (GAN) may be applied to optimize the parameters of the learning model 100. A GAN is generally configured by linking two learning models, a generator and a discriminator, and the parameters of the generator and discriminator are updated, respectively, to train the entire network. Known original learning data or data output by the generator is input to the discriminator. The discriminator outputs a discrimination result as a result of discriminating whether the input data is original learning data (true) or data output by the generator (false).

[0101] In training the generator, parameters are optimized so that the data output by the generator is not classified as "false" by the classifier, in other words, so that the data output by the generator is classified as "true" by the classifier. In training the classifier, parameters are optimized to improve the accuracy of true / false determination. The parameters are updated using losses calculated by applying a loss function for the classifier and a loss function for the generator. The loss for the classifier is calculated by applying the classification result output by the classifier to the loss function for the classifier. The loss for the generator is calculated by applying the loss for the classifier to the loss function for the generator.

[0102] In this embodiment, a learning model 100 that outputs a first learning output image 101 and a second learning output image 102 is used as a generator, and as shown in Fig. 7, a classification learning model 300 that is different from the learning model 100 that is the generator is provided as a classifier in the learning unit 11. The classification learning model 300 receives the first learning output image 101 and the second learning output image 102 output by the learning model 100 as input, and outputs a first classification result 301 and a second classification result 302.

[0103] The classification learning model 300 is a learning model that outputs the source image 21 as "true" and any image that is not the source image 21 as "false," and performs a classification of the authenticity of the first training output image 101 and the second training output image 102 output from the training model 100 as a generator. The first classification result 301 is the result of the classification of the authenticity of the first training output image 101 performed by the classification learning model 300. The second classification result 302 is the result of the classification of the authenticity of the second training output image 102 performed by the classification learning model 300.

[0104] First classification result 301 and second classification result 302 are input to evaluation unit 50. Evaluation unit 50 applies first classification result 301 to a loss function for the classifier to calculate a first classifier loss as the loss of classification learning model 300, and also applies second classification result 302 to the loss function for the classifier to calculate a second classifier loss. Update unit 60 updates classification learning model 300 by optimizing the parameters of classification learning model 300 based on the first classifier loss and the second classifier loss.

[0105] Furthermore, the evaluation unit 50 applies the first classification result and the source image 21 to a loss function for the generator to calculate a first loss 51 (first generator loss) as the loss of the learning model 100 as a generator. That is, the first generator loss is a loss calculated using the first classification result based on the first training output image 101 and the source image 21. Similarly, the evaluation unit 50 applies the second classification result and the source image 21 to a loss function for the generator to calculate a second loss 52 (second generator loss) as the loss of the learning model 100. That is, the second generator loss is a loss calculated using the first classification result based on the second training output image 102 and the source image 21.

[0106] The update unit 60 updates the learning model 100 by optimizing the parameters of the learning model 100 as a generator based on the first loss 51 (first generator loss) and the second loss 52 (second generator loss). In this case, the trained model 200 (trained generator) is the learning model 100 as a trained generator. As described above, by employing a GAN and configuring a network so that the learning model 100 serves as a generator, highly accurate super-resolution can be achieved. The above configuration is particularly suitable even when there are a small number of source images 21.

[0107] The trained model 200 generated by updating the learning model 100 in this embodiment is suitable for a case where an unknown image, that is, an input image for inference 201, is an endoscopic image. The endoscopic image is an image generated by an endoscope using an endoscope as the modality 30 to capture an image of a subject. In this case, as shown in FIG. 8, the inference unit 12 inputs an endoscopic image 203 having a first number of pixels stored in the database 20 into the trained model 200, and outputs a super-resolution image 204 having a second number of pixels that is greater than the first number of pixels.

[0108] The trained model 200 that outputs a super-resolution image 204 having a second number of pixels is generated by using a training input image 41 having a fourth number of pixels, which is smaller than the third number of pixels, generated by performing degradation processing on a source endoscopic image 221 having a third number of pixels, as shown in Figure 9, and outputs a first training output image 101 and a second training output image 102 having a fifth number of pixels, which is greater than the fourth number of pixels, by an update unit 60 updating the training model 100 using a first loss 51 calculated by applying the first training output image 101 and the source endoscopic image 221 to a loss function calculated by an evaluation unit 50, and a second loss 52 calculated by applying the second training output image 102 and the source endoscopic image 221 to a loss function.

[0109] As in the case where the source image 21 and the inference input image 201 are images other than endoscopic images, the ratio of the second number of pixels in the super-resolution image 204 to the first number of pixels in the endoscopic image 203 is equal to the ratio of the fifth number of pixels in the first learning output image 101 and the second learning output image 102 to the fourth number of pixels in the learning input image 41.

[0110] For example, when the learning model 100 is designed to output a learning output image having four times the number of pixels of the learning input image 41, the trained model 200 receives an endoscopic image 203 having a pixel count of "512 x 512" and outputs a super-resolution image 204 having a pixel count of "1024 x 1024." In this case, the pixel count of the source endoscopic image 221 is set to "512 x 512," and degradation processing is performed on the source endoscopic image 221 to obtain a learning input image 41 having a pixel count of "256 x 256," which is input to the learning model 100. The learning model 100 outputs a first learning output image 101 and a second learning output image 102 having a pixel count of "512 x 512," which is four times the pixel count of the learning input image 41. In this case, the evaluation unit 50 calculates the first loss 51 and the second loss 52 by applying the first learning output image 101 and the second learning output image 102, each having a pixel count of "512 x 512", and the source endoscopic image 221, each having a pixel count of "512 x 512", to a loss function.

[0111] As in the case where the source image 21 and the input image for inference 201 are images other than endoscopic images, the third number of pixels of the source endoscopic image 221 is equal to or less than the first number of pixels of the endoscopic image 203, which is an unknown image input to the trained model 200. In particular, as in the above example, it is preferable that the third number of pixels, which is the number of pixels of the source endoscopic image 221, and the first number of pixels, which is the number of pixels of the endoscopic image 203, are equal to each other.

[0112] In addition, the learning model 100 may be designed to output a learning output image having 16 times the number of pixels of the learning input image 41, and if the number of pixels of the source image 21 is "512 x 512", an endoscopic image 203 having a pixel count of "1280 x 960" may be input to generate a super-resolution image 204 having a pixel count of "5120 x 3840" and a resolution of 4K or higher.

[0113] As shown in FIG. 10 , the trained model 200 is preferably configured with an input layer 110, a first hidden layer 120, a folding layer 130, a second hidden layer 140 that receives a feature map from the first hidden layer 120, and a first output layer 150. For training the training model 100, a third hidden layer 160, which is a second decoder, and a second output layer 170 are required to obtain a training output image with higher accuracy, but one decoder is sufficient to output a super-resolution image from an unknown image. Furthermore, by omitting the third hidden layer 160 and the second output layer 170 from the trained model 200, the memory constituting the processor can be reduced, thereby improving processing speed.

[0114] Furthermore, when displaying the super-resolution image 202 received from the trained model 200 on the display 80, the display control unit 70 preferably displays the super-resolution image 202 and information indicating that the image is an endoscopic image that has been subjected to high-resolution processing. For example, as in the example of the super-resolution image 204 shown in Fig. 11, a notification display 210 indicating that four times super-resolution (Super Resolution) has been performed on the endoscopic image is displayed on the display 80, displaying "x4 SR."

[0115] The super-resolution image 204 generated by the trained model 200, which is a machine learning model, is an artificially generated image and therefore cannot be used for diagnosis by a doctor. However, it is useful when, for example, an endoscopic image is to be displayed on a large display 80, such as when a doctor shows an endoscopic image to a patient to explain, or when multiple people observe an endoscopic image, and when an area of ​​interest, such as a lesion or treatment target, is to be enlarged or observed. Therefore, by displaying the notification display 210 on the super-resolution image 204, a person observing the super-resolution image 204 can observe the super-resolution image with high definition while recognizing that the image displayed on the display 80 is an artificially generated image.

[0116] Because there are recommended standards for storing endoscopic images in the database 20 or for communicating between multiple databases 20, there is a practical limit to the number of pixels in the image that can be obtained. Furthermore, it may be difficult to capture an image with a large number of pixels depending on the machine specifications of the endoscope. In such a situation, the image processing device 10 of this embodiment and the endoscopic system including the image processing device 10, an endoscope, and a display 80 can generate a super-resolution image with a number of pixels that exceeds the limit of the number of pixels in an endoscopic image that can actually be obtained.

[0117] A series of steps in the operation method of the image processing device 10 of this embodiment will be described using the flowchart of FIG. 12. First, the learning unit 11 acquires a source image from the database 20 (step ST101). Next, the degradation processing unit 40 performs degradation processing on the source image 21 to generate a learning input image 41 (step ST102). The learning unit 11 inputs the learning input image 41 to the learning model 100 and outputs a first learning output image 101 and a second learning output image 102 (collectively referred to as "learning output images" in FIG. 12) that have a larger number of pixels than the learning input image 41 (step ST103). Next, the evaluation unit 50 calculates a first loss based on the first learning output image 101 and the source image 21, and further calculates a second loss based on the second learning output image 102 and the source image 21 (calculate loss) (step ST104). Finally, the learning model 100 is updated based on the first loss and the second loss (step ST105), thereby generating the learned model 200 (step ST106).

[0118] In the above embodiment, the hardware structure of the processing units that execute various processes, such as the degradation processing unit 40, the learning model 100, the trained model 200, the evaluation unit 50, the update unit 60, and the display control unit 70, is the following various processors: The various processors include a CPU (Central Processing Unit), which is a general-purpose processor that executes software (programs) and functions as various processing units, a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), whose circuit configuration can be changed after manufacture, and a dedicated electrical circuit, which is a processor with a circuit configuration designed specifically for executing various processes.

[0119] A single processing unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, multiple FPGAs, or a combination of a CPU and an FPGA). Also, multiple processing units may be configured with a single processor. Examples of multiple processing units configured with a single processor include, first, a configuration in which one processor is configured with a combination of one or more CPUs and software, as typified by client or server computers, and this processor functions as multiple processing units. Second, a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a System on Chip (SoC). In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure.

[0120] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit formed by combining circuit elements such as semiconductor elements, and the hardware structure of the memory unit is a storage device such as a hard disk drive (HDD) or a solid state drive (SSD). [Explanation of symbols]

[0121] 10 Image processing device 11 Learning Department 12 Reasoning part 20 databases 21 Source Image 30 Modalities 40 Deterioration processing section 41 training input images 50 Evaluation Department 51 1st loss 52 Second loss 60 Update section 70 Display control unit 80 Display 100 Learning Models 101 First training output image 102 Second training output image 110 Input Layer 111, 120a, 121a, 122a, 123a, 124a, 131, 140a, 141a, 142a, 143a, 144a, 160a, 161a, 162a, 163a, 164a feature maps 120 First Middle Class 121, 122, 123, 124 Convolutional Layers 130 Folded Layer 140 Second Middle Class 141, 142, 143, 144, 161, 162, 163, 164 upsampling layers 150 First output layer 160 Third Middle Class 170 Second output layer 200 pre-trained models 201 Input image for inference 202, 204 Super-resolution images 203 Endoscopic Images 210 Notification display 221 Source Endoscopic Images 300 Discriminative Learning Model 301 First Identification Result 302 Second Identification Result

Claims

1. a processor; The processor: Get the source image, generating a learning input image by performing degradation processing on the source image; inputting the learning input image into a learning model to output a first learning output image and a second learning output image having a larger number of pixels than the learning input image; Calculating a first loss based on the first training output image and the source image; Calculating a second loss based on the second training output image and the source image; generating a trained model by updating the training model based on the first loss and the second loss; the learning model has an input layer, a first hidden layer, a folding layer, a second hidden layer, a first output layer, a third hidden layer, and a second output layer; the input layer receives the learning input image as input and outputs a feature map to be input to the first intermediate layer; the first hidden layer receives the feature map output by the input layer and outputs the feature map to be input to the folding layer and the second hidden layer; the folding layer outputs the feature map to be input to the second hidden layer and the third hidden layer based on the feature map input from the first hidden layer; the second hidden layer outputs the feature map to be input to the first output layer based on the feature map input from the folding layer and the feature map input from the first hidden layer; the first output layer outputs the first learning output image based on the feature map input from the second hidden layer; the third hidden layer outputs the feature map to be input to the second output layer based on the feature map input from the folding layer; The second output layer outputs the second learning output image based on the feature map input from the third intermediate layer.

2. an amount of information of the training input image, the first training output image, the second training output image, and the feature map is determined by a number of bits according to the number of pixels, the number of channels, and a data type; 2. The image processing device according to claim 1, wherein the folding layer reduces the amount of information of the feature map output by the folding layer to be smaller than the amount of information of the training input image by performing an information amount reduction process that changes the number of channels or the number of bits.

3. The image processing device according to claim 2 , wherein the information amount reduction process is a process of reducing the number of channels.

4. The image processing device according to claim 2 , wherein the information amount reduction process is a process of reducing the number of bits.

5. The image processing device according to claim 3 , wherein the first intermediate layer performs processing to reduce the number of pixels in the feature map input from the input layer.

6. The image processing device according to claim 5 , wherein the second hidden layer and the third hidden layer perform processing to increase the number of pixels in the feature map input from the folding layer.

7. The image processing device according to claim 1 , wherein the degradation process includes a process of reducing the number of pixels of the source image.

8. The image processing device according to claim 7 , wherein the degradation processing includes filtering and / or noise addition processing.

9. The processor: The image processing device according to claim 1 , wherein the learning input image is further input to the second intermediate layer.

10. The processor: inputting an inference input image having a first number of pixels into the trained model, and outputting a super-resolution image having a second number of pixels greater than the first number of pixels; 2. The image processing device according to claim 1, wherein the ratio of the second number of pixels to the first number of pixels is equal to the ratio of the number of pixels of the first learning output image and the second learning output image to the number of pixels of the learning input image.

11. obtaining a source image; generating a learning input image by performing degradation processing on the source image; a step of inputting the learning input image into a learning model to output a first learning output image and a second learning output image having a larger number of pixels than the learning input image; calculating a first loss based on the first training output image and the source image; calculating a second loss based on the second training output image and the source image; and generating a trained model by updating the training model based on the first loss and the second loss; the learning model has an input layer, a first hidden layer, a folding layer, a second hidden layer, a first output layer, a third hidden layer, and a second output layer; the input layer receives the learning input image as input and outputs a feature map to be input to the first intermediate layer; the first hidden layer receives the feature map output by the input layer and outputs the feature map to be input to the folding layer and the second hidden layer; the folding layer outputs the feature map to be input to the second hidden layer and the third hidden layer based on the feature map input from the first hidden layer; the second hidden layer outputs the feature map to be input to the first output layer based on the feature map input from the folding layer and the feature map input from the first hidden layer; the first output layer outputs the first learning output image based on the feature map input from the second hidden layer; the third hidden layer outputs the feature map to be input to the second output layer based on the feature map input from the folding layer; A method for operating an image processing device, wherein the second output layer outputs the second learning output image based on the feature map input from the third intermediate layer.

12. a processor; The processor: Get the source image, generating a learning input image by performing degradation processing on the source image; inputting the learning input image to a generator to output a first learning output image and a second learning output image having a larger number of pixels than the learning input image; The first learning output image and the second learning output image are input to a classifier, outputting a first classification result based on the first learning output image and a second classification result based on the second learning output image; calculating a first classifier loss based on the first classification result and a second classifier loss based on the second classification result; updating the classifier based on the first classifier loss and the second classifier loss; calculating a first generator loss based on the first classification result and the source image and a second generator loss based on the second classification result and the source image; generating a trained generator by updating the generator based on the first generator loss and the second generator loss; the generator has an input layer, a first hidden layer, a folding layer, a second hidden layer, a first output layer, a third hidden layer, and a second output layer; the input layer receives the learning input image as input and outputs a feature map to be input to the first intermediate layer; the first hidden layer receives the feature map output by the input layer and outputs the feature map to be input to the folding layer and the second hidden layer; the folding layer outputs the feature map to be input to the second hidden layer and the third hidden layer based on the feature map input from the first hidden layer; the second hidden layer outputs the feature map to be input to the first output layer based on the feature map input from the folding layer and the feature map input from the first hidden layer; the first output layer outputs the first learning output image based on the feature map input from the second hidden layer; the third hidden layer outputs the feature map to be input to the second output layer based on the feature map input from the folding layer; The second output layer outputs the second learning output image based on the feature map input from the third intermediate layer.

13. An image processing device including a processor, The processor: acquiring an endoscopic image having a first number of pixels; inputting the endoscopic image into a trained model to output a super-resolution image having a second number of pixels greater than the first number of pixels; The trained model is an image processing device generated by using a training input image having a fourth number of pixels smaller than the third number of pixels generated by degradation processing a source endoscopic image having a third number of pixels equal to or less than the first number of pixels, and updating the training model to output a first training output image and a second training output image having a fifth number of pixels larger than the fourth number of pixels based on a first loss based on the first training output image and the source endoscopic image and a second loss based on the second training output image and the source endoscopic image.

14. The processor: The image processing device according to claim 13 , wherein the image processing device performs control to display the super-resolution image and information indicating that the endoscopic image has been subjected to high-resolution processing.

15. 15. The image processing device according to claim 13, wherein a ratio of the second number of pixels to the first number of pixels is equal to a ratio of the fifth number of pixels to the fourth number of pixels.

16. The image processing device according to claim 15 , wherein the third number of pixels is equal to the first number of pixels.

17. The image processing device according to claim 13; an endoscope that captures an image of a subject to generate the endoscopic image; a display; The processor: An endoscope system that controls the display of the super-resolution image on the display.

Citation Information

Patent Citations

  • Image processing device, image processing method, processing device, processing method and program

    JP2020024612A

  • Training method of image processing model, image processing method, apparatus, and device

    US20220261965A1

  • Learning device, super-resolving device, learning method, super-resolving method, and program

    WO2018235168A1

  • Learning method, learning system, learned model, program, and super-resolution image generation device

    WO2020175446A1

  • Super resolution using convolutional neural network

    WO2021163844A1