Medical image processing device, method of operating the medical image processing device, and program

JP7927957B2Active Publication Date: 2026-10-01CANON KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025152990
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-03-29
Filing Date
2025-09-16
Publication Date
2026-10-01
Estimated Expiration
2039-10-03

AI Technical Summary

Benefits of technology

【0012】 本発明の一つによれば、従来よりも画像診断に適した画像を生成することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927957000001
    Figure 0007927957000001
  • Figure 0007927957000002
    Figure 0007927957000002
  • Figure 0007927957000003
    Figure 0007927957000003
Patent Text Reader

Abstract

To provide a medical image processing device that can generate images more suitable for image diagnosis than conventional ones.SOLUTION: A medical image processing device comprises: an acquisition unit for acquiring a first image which is a medical image of a prescribed portion of a subject; a high-quality image enhancement unit for generating, from the first image, a second image whose image quality is enhanced compared to the first image, using a high-quality image enhancement engine including a machine learning engine; and a display control unit for causing a display unit to display a composite image obtained by combining the first image and the second image based on a ratio obtained by using information regarding at least some regions in the first image.SELECTED DRAWING: Figure 34
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a medical image processing apparatus, medical image processing Device operation method and program. [Background Art]

[0002] In the medical field, in order to identify a subject's disease and observe the degree of the disease, images are acquired by various imaging apparatuses, and image diagnosis is performed by medical staff. Types of imaging apparatuses include, for example in the field of radiology, an X-ray imaging apparatus, an X-ray computed tomography (CT) apparatus, a magnetic resonance imaging (MRI) apparatus, a positron emission tomography (PET) apparatus, and a single photon emission computed tomography (SPECT) apparatus, among others. Also, for example in the field of ophthalmology, there are a fundus camera, a scanning laser ophthalmoscope (SLO), an optical coherence tomography (OCT) apparatus, and an OCT angiography (OCTA) apparatus.

[0003] In order to perform accurate image diagnosis and complete it in a short time, high image quality such as low noise, high resolution and spatial resolution, and appropriate gradation of an image acquired by an imaging apparatus is important. Also, an image in which a region to be observed or a lesion is emphasized can also be useful.

[0004] However, for many imaging apparatuses, some kind of trade-off is required in order to acquire an image suitable for image diagnosis, such as an image with high image quality. For example, there is a method of purchasing a high-performance imaging apparatus in order to acquire a high-image-quality image, but this often requires more investment than acquiring a low-performance one.

[0005] Furthermore, with CT scans, for example, it may be necessary to increase the patient's radiation dose to obtain images with less noise. Similarly, with MRI scans, contrast agents with potential side effects may be used to enhance images of the area to be observed. Also, with OCT scans, for example, the scanning time may be longer if the area to be scanned is large or high spatial resolution is required. Additionally, some imaging devices require multiple image acquisitions to obtain high-quality images, which increases the scanning time.

[0006] Patent Document 1 discloses a technology that uses an artificial intelligence engine to convert previously acquired images into higher-resolution images in order to cope with the rapid advancements in medical technology and the need for simple imaging in emergencies. With such a technology, for example, images acquired through simple imaging with minimal cost can be converted into higher-resolution images. [Prior art documents] [Patent Documents]

[0007] [Patent Document 1] Japanese Patent Publication No. 2018-5841 [Overview of the Initiative] [Problems that the invention aims to solve]

[0008] However, even high-resolution images may not be suitable for diagnostic imaging. For example, even high-resolution images may not adequately identify the object to be observed if they are noisy or have low contrast.

[0009] In contrast, one of the objectives of the present invention is to provide a medical image processing apparatus, a medical image processing method, and a program that can generate images more suitable for medical imaging than conventional methods. [Means for solving the problem]

[0010] A medical image processing apparatus according to one embodiment of the present invention is: An acquisition unit acquires a first image, which is a medical image of a predetermined part of the subject, The system comprises a display control unit that controls the display unit to display either a second image obtained by inputting the first image as an input image to the trained model, or the first image. The first image and the The A first display screen in which either of the two images is displayed, and the first image and the The In a case where, among the two images, a second display screen shows either image 2 or image 2, and the display screen is changed from one display screen to the other in response to instructions from the examiner, If the first image is displayed on one of the display screens, the modification is made so that the first image is displayed on the other display screen, and If the second image is displayed on one of the display screens, the modification is made so that the second image is displayed on the other display screen.

[0011] Furthermore, medical image processing according to another embodiment of the present invention Device operation The method is, The aforementioned medical image processing device is The first image is a medical image of a designated area of ​​the subject, The aforementioned medical image processing device is This includes controlling the display unit to display either a second image obtained by inputting the first image as an input image to the trained model, or the first image, The first image and the The A first display screen in which either of the two images is displayed, and the first image and the The In a case where, among the two images, a second display screen shows either image 2 or image 2, and the display screen is changed from one display screen to the other in response to instructions from the examiner, If the first image is displayed on one of the display screens, the modification is made so that the first image is displayed on the other display screen, and When the second image is displayed on said one display screen, said change is performed such that the second image is displayed on the other display screen. Effects of the Invention

[0012] According to one aspect of the present invention, an image more suitable for image diagnosis than conventional techniques can be generated. Brief Description of Drawings

[0013] [Figure 1] An example of the configuration of a neural network related to image quality enhancement processing is shown. [Figure 2] An example of the configuration of a neural network related to imaging region estimation processing is shown. [Figure 3] An example of the configuration of a neural network related to image authenticity evaluation processing is shown. [Figure 4] An example of the schematic configuration of the image processing apparatus according to the first embodiment is shown. [Figure 5] It is a flow diagram showing an example of the flow of image processing according to the first embodiment. [Figure 6] It is a flow diagram showing another example of the flow of image processing according to the first embodiment. [Figure 7] It is a flow diagram showing an example of the flow of image processing according to the second embodiment. [Figure 8] It is a diagram for explaining image processing according to the fourth embodiment. [Figure 9] It is a flow diagram showing an example of the flow of image quality enhancement processing according to the fourth embodiment. [Figure 10] It is a diagram for explaining image processing according to the fifth embodiment. [Figure 11] It is a flow diagram showing an example of the flow of image quality enhancement processing according to the fifth embodiment. [Figure 12] It is a diagram for explaining image processing according to the sixth embodiment. [Figure 13] It is a flow diagram showing an example of the flow of image quality enhancement processing according to the sixth embodiment. [Figure 14]This is a diagram illustrating the image processing according to the sixth embodiment. [Figure 15] An example of a schematic configuration of the image processing apparatus according to the seventh embodiment is shown. [Figure 16] This is a flowchart showing an example of the image processing flow according to the seventh embodiment. [Figure 17] An example of a user interface according to the seventh embodiment is shown. [Figure 18] An example of a schematic configuration of the image processing apparatus according to the ninth embodiment is shown. [Figure 19] This is a flowchart showing an example of the image processing flow according to the ninth embodiment. [Figure 20] An example of a schematic configuration of the image processing apparatus according to the twelfth embodiment is shown. [Figure 21A] This is a flowchart showing an example of the image quality enhancement process according to the 13th embodiment. [Figure 21B] This flowchart shows another example of the image quality enhancement process according to the 13th embodiment. [Figure 22] An example of a schematic configuration of the image processing apparatus according to the 17th embodiment is shown. [Figure 23] This is a flowchart showing an example of the image processing flow according to the 17th embodiment. [Figure 24] This shows an example of a neural network configuration for image quality enhancement processing. [Figure 25] An example of a schematic configuration of the image processing apparatus according to the 19th embodiment is shown. [Figure 26] This is a flowchart showing an example of the image processing flow according to the 19th embodiment. [Figure 27] This is a flowchart showing an example of the image processing flow according to the 21st embodiment. [Figure 28] An example of a training image for image quality enhancement processing is shown. [Figure 29] An example of an input image for image quality enhancement processing is shown. [Figure 30] An example of a schematic configuration of the image processing apparatus according to the 22nd embodiment is shown. [Figure 31] This is a flowchart showing an example of the image processing flow according to the 22nd embodiment. [Figure 32] This is a diagram illustrating a wide-angle image according to the 22nd embodiment. [Figure 33] This is a diagram illustrating the image enhancement process according to the 23rd embodiment. [Figure 34] An example of a user interface according to the 24th embodiment is shown. [Figure 35] An example of a schematic configuration of the image processing apparatus according to the 25th embodiment is shown. [Figure 36] An example of the configuration of a neural network used as a machine learning engine in the modified example 6 is shown. [Figure 37] An example of the configuration of a neural network used as a machine learning engine in the modified example 6 is shown. [Figure 38] An example of a user interface according to the 24th embodiment is shown. [Modes for carrying out the invention]

[0014] Hereinafter, exemplary embodiments for carrying out the present invention will be described in detail with reference to the drawings. However, the dimensions, materials, shapes, and relative positions of components described in the following embodiments are arbitrary and can be changed according to the configuration of the device to which the present invention is applied or various conditions. In addition, the same reference numerals are used between drawings to indicate elements that are identical or functionally similar.

[0015] <Explanation of Terms> First, we will explain the terms used in this specification.

[0016] In the network described herein, each device may be connected by a wired or wireless line. Here, the lines connecting each device in the network include, for example, dedicated lines, local area network (hereinafter referred to as LAN) lines, wireless LAN lines, internet lines, Wi-Fi (registered trademark), and Bluetooth (registered trademark).

[0017] A medical image processing system may consist of two or more devices capable of communicating with each other, or it may consist of a single device. Furthermore, each component of the medical image processing system may consist of software modules executed by a processor such as a CPU (Central Processing Unit) or MPU (Micro Processing Unit). Each of these components may also consist of circuits performing specific functions, such as ASICs. It may also be composed of any combination of other hardware and any software.

[0018] Furthermore, the medical images processed by the medical image processing apparatus or medical image processing method according to the embodiments described below include images acquired using any modality (imaging apparatus, imaging method). The medical images to be processed may include medical images acquired by any imaging device, etc., or images created by a medical image processing device or medical image processing method according to the embodiments described below.

[0019] Furthermore, the medical images to be processed are images of a predetermined part of the subject (test subject), and the image of the predetermined part includes at least a portion of that predetermined part of the subject. The medical images may also include other parts of the subject. The medical images may be still or moving images, and may be black and white or color images. Furthermore, the medical images may represent the structure (morphology) of the predetermined part, or they may represent its function. Images representing function include, for example, OCTA images, Doppler OCT images, fMRI images, and ultrasound Doppler images that represent hemodynamics (blood flow rate, blood flow velocity, etc.). The predetermined part of the subject may be determined according to the subject being photographed, and may include the human eye (test eye), brain, lungs, intestines, heart, pancreas, kidneys, liver and other organs, head, chest, legs, and arms.

[0020] Furthermore, the medical image may be a tomographic image of the subject or a frontal image. Frontal images include, for example, frontal images of the fundus, frontal images of the anterior segment, fluorescently scanned fundus images, and En-Face images generated using data from OCT (3D OCT data) in at least a portion of the depth range of the subject being scanned. Note that the En-Face image may also be an OCTA En-Face image (motion contrast frontal image) generated using data from 3D OCTA data (3D motion contrast data) in at least a portion of the depth range of the subject being scanned. Also, 3D OCT data and 3D motion contrast data are examples of 3D medical image data.

[0021] Furthermore, an imaging device is a device for capturing images used for diagnosis. Imaging devices include, for example, devices that obtain images of a predetermined area of ​​a subject by irradiating that area with light, X-rays or other radiation, electromagnetic waves, or ultrasound, and devices that obtain images of a predetermined area by detecting radiation emitted from a subject. More specifically, imaging devices according to the following embodiments include at least an X-ray imaging device, a CT scanner, an MRI scanner, a PET scanner, a SPECT scanner, an SLO scanner, an OCT scanner, an OCTA scanner, a fundus camera, and an endoscope.

[0022] Furthermore, the OCT device may include time-domain OCT (TD-OCT) devices and Fourier-domain OCT (FD-OCT) devices. The Fourier-domain OCT device may also include spectral-domain OCT (SD-OCT) devices and wavelength-swept OCT (SS-OCT) devices. In addition, the SLO device and OCT device may include wavefront-compensated SLO (AO-SLO) devices and wavefront-compensated OCT (AO-OCT) devices using wavefront compensating optical systems. In addition, the SLO device and OCT device may include polarization SLO (PS-SLO) devices and polarization OCT (PS-OCT) devices for visualizing information related to polarization phase difference and depolarization.

[0023] An image management system is a device and system that receives and stores images captured by an imaging device and images that have undergone image processing. Furthermore, the image management system can transmit images in response to requests from connected devices, perform image processing on stored images, and request image processing from other devices. An image management system can include, for example, a picture archiving and communication system (PACS). In particular, the image management system according to the following embodiment includes a database capable of storing various information, such as patient information and shooting time, associated with the received images. The image management system is also connected to a network and can send and receive images, convert images, and send and receive various information associated with stored images in response to requests from other devices.

[0024] Image acquisition conditions refer to various pieces of information obtained when an image is captured by an imaging device. These conditions may include, for example, information about the imaging device, information about the facility where the imaging was performed, information about the examination related to the imaging, information about the person taking the image, and information about the subject. They may also include, for example, the date and time of imaging, the name of the area being imaged, the imaging area, the imaging angle, the imaging method, the image resolution and grayscale, the image size, the applied image filter, the image data format, and information about the radiation dose. The imaging area may include surrounding areas shifted from a specific imaging area, or areas encompassing multiple imaging areas.

[0025] Shooting conditions can be stored within the data structure that makes up the image, as separate shooting condition data from the image, or in a database or image management system related to the imaging device. Therefore, shooting conditions can be obtained by procedures corresponding to the means by which the imaging device stores its shooting conditions. Specifically, shooting conditions can be obtained, for example, by analyzing the data structure of the image output by the imaging device, obtaining shooting condition data corresponding to the image, or accessing an interface for obtaining shooting conditions from a database related to the imaging device.

[0026] It should be noted that some imaging conditions cannot be acquired for reasons such as not being saved, depending on the imaging device. For example, the imaging device may not have a function to acquire or save specific imaging conditions, or such a function may be disabled. In addition, some imaging conditions may not be saved if they are deemed unrelated to the imaging device or imaging itself. Furthermore, for example, imaging conditions may be concealed, encrypted, or only accessible with the necessary rights. However, even if imaging conditions are not saved, they may still be obtainable. For example, by performing image analysis, it may be possible to identify the name of the imaging body part or the imaging area.

[0027] A machine learning model is a model that has been trained (learned) in advance using appropriate training data (training data) for any machine learning algorithm. Training data consists of one or more pairs of input data and output data (correct data). The format and combination of input and output data in the training data pairs can be any format suitable for the desired configuration, such as one being an image and the other a number, one being a group of images and the other a string, or both being images.

[0028] Specifically, one example is training data (hereinafter referred to as "first training data") consisting of pairs of images acquired by OCT and corresponding imaging area labels. The imaging area labels are unique numerical values ​​or strings that represent the body part. Another example of training data is training data (hereinafter referred to as "second training data") consisting of pairs of noisy, low-resolution images acquired by normal OCT imaging and high-resolution images obtained by taking multiple OCT scans and processing them to enhance image quality.

[0029] When input data is given to a machine learning model, output data is output according to the design of the machine learning model. For example, the machine learning model outputs output data that is likely to correspond to the input data, according to the trends trained using the training data. The machine learning model can also output the probability of each type of output data corresponding to the input data as a numerical value, according to the trends trained using the training data. Specifically, for example, if an image acquired by OCT is input to a machine learning model trained with first training data, the machine learning model will output the imaging area labels of the imaging areas captured in the image, and output the probability for each imaging area label. Also, for example, if a noisy, low-resolution image acquired by normal OCT imaging is input to a machine learning model trained with second training data, the machine learning model will output a high-resolution image equivalent to an image that has been taken multiple times by OCT and processed for high-resolution enhancement. Note that, from the perspective of maintaining quality, the machine learning model can be configured not to use its own output data as training data.

[0030] Furthermore, machine learning algorithms include deep learning techniques such as convolutional neural networks (CNNs). In deep learning techniques, different parameter settings for the layers and nodes that make up the neural network may result in different degrees of accuracy in reproducing the trained trends in the output data. For example, in a deep learning machine learning model using the first set of training data, more appropriate parameter settings may increase the probability of outputting the correct image location label. Also, for example, in a deep learning machine learning model using the second set of training data, more appropriate parameter settings may result in the output of higher-quality images.

[0031] Specifically, parameters in a CNN can include, for example, the kernel size of the filters, the number of filters, the stride value, and the dilation value set for the convolutional layers, as well as the number of output nodes in the fully connected layers. The parameter set and the number of training epochs can be set to values ​​favorable to the intended use of the machine learning model, based on the training data. For example, based on the training data, the parameter set and number of epochs can be set to output the correct imaging site label with a higher probability or to output higher quality images.

[0032] Here is an example of how to determine the parameter set and the number of epochs. First, 70% of the pairs constituting the training data are randomly set for training, and the remaining 30% for evaluation. Next, the machine learning model is trained using the training pairs, and at the end of each training epoch, the training evaluation value is calculated using the evaluation pairs. The training evaluation value is, for example, the average of the values ​​obtained by evaluating the output when the input data constituting each pair is input to the machine learning model during training, and the output data corresponding to the input data, using a loss function. Finally, the parameter set and the number of epochs at which the training evaluation value is smallest are determined as the parameter set and the number of epochs for the machine learning model. Note that by dividing the pairs constituting the training data into training and evaluation sets and determining the number of epochs in this way, it is possible to prevent the machine learning model from overfitting to the training pairs.

[0033] A high-resolution engine (a trained model for high-resolution enhancement) is a module that outputs a high-resolution image by enhancing an input low-resolution image. Here, high-resolution enhancement in this specification means converting an input image into an image with a resolution suitable for medical imaging, and a high-resolution image is an image that has been converted into an image with a resolution suitable for medical imaging. Low-resolution images are, for example, two-dimensional or three-dimensional images acquired by X-ray, CT, MRI, OCT, PET, or SPECT, or three-dimensional moving images from continuously scanned CT scans, etc., that were taken without any settings to achieve particularly high resolution. Specifically, low-resolution images include, for example, images acquired by low-dose imaging with X-ray equipment or CT, imaging with MRI without contrast agents, short-time imaging with OCT, and OCTA images acquired with a small number of scans.

[0034] Furthermore, the characteristics of image quality suitable for medical imaging depend on what needs to be diagnosed in each type of medical imaging. Therefore, it is difficult to generalize, but for example, image quality suitable for medical imaging includes low noise, high contrast, colors and gradations that make the subject easy to observe, large image size, and high resolution. It can also include image quality in which non-existent objects or gradients that were drawn during the image generation process have been removed from the image.

[0035] Furthermore, using high-resolution images with low noise and high contrast for image analysis, such as vascular analysis using OCTA or region segmentation using CT or OCT images, often results in more accurate analysis than using low-resolution images. Therefore, high-resolution images output by image enhancement engines can be useful not only for diagnostic imaging but also for image analysis.

[0036] The image processing method constituting the image enhancement method in the embodiment described below performs processing using various machine learning algorithms such as deep learning. In addition to processing using machine learning algorithms, the image processing method may also perform various image filtering processes, matching processes using a database of high-resolution images corresponding to similar images, and knowledge-based image processing.

[0037] In particular, Figure 1 shows an example of a CNN configuration for improving the image quality of two-dimensional images. This CNN configuration includes a group of multiple convolutional processing blocks 100. Each convolutional processing block 100 includes a convolutional layer 101, a batch normalization layer 102, and an activation layer 103 using a RectifierLinear Unit. This CNN configuration also includes a merger layer 104 and a final convolutional layer 105. The merger layer 104 combines the output values ​​of the convolutional processing blocks 100 with the pixel values ​​that make up the image by concatenating or adding them. The final convolutional layer 105 outputs the pixel values ​​that make up the high-quality image Im120, which were combined in the merger layer 104. In this configuration, the pixel value group that makes up the input image Im110 is output via the convolutional processing block 100, and this output group is combined with the original pixel value group that makes up the input image Im110 in the synthesis layer 104. Subsequently, the combined pixel value group is formed into a high-resolution image Im120 in the final convolutional layer 105.

[0038] For example, by setting the number of convolutional processing blocks 100 to 16 and the parameters of the convolutional layer group 101 to a filter kernel size of 3 pixels wide, 3 pixels high, and 64 filters, a certain level of image quality improvement can be obtained. However, as described in the explanation of the machine learning model above, in practice, a better set of parameters can be set using training data appropriate to the usage of the machine learning model. Furthermore, if it is necessary to process three-dimensional or four-dimensional images, the filter kernel size may be extended to three or four dimensions.

[0039] Furthermore, when using certain image processing techniques, such as those employing CNNs, it is necessary to pay attention to image size. Specifically, to address issues such as insufficient enhancement of the peripheral areas of high-resolution images, it should be noted that different image sizes may be required for the input low-resolution image and the output high-resolution image.

[0040] For the sake of clarity, although not explicitly stated in the embodiments described later, if a high-resolution enhancement engine is used that requires different image sizes for the input image and the output image, the image size is adjusted as appropriate. Specifically, the image size is adjusted for input images, such as images used as training data for training a machine learning model or images input to the high-resolution enhancement engine, by padding or by merging the surrounding shooting area of ​​the input image. The padding area is filled with a fixed pixel value, filled with neighboring pixel values, or mirror padding is performed in accordance with the characteristics of the high-resolution enhancement method in order to effectively enhance image quality.

[0041] Furthermore, image enhancement techniques may be implemented using only one image processing method, or they may be implemented using a combination of two or more image processing methods. Alternatively, multiple image enhancement methods may be implemented in parallel to generate multiple high-resolution image sets, and the highest-resolution image may be selected as the final high-resolution image. The selection of the highest-resolution image may be performed automatically using an image quality evaluation index, or it may be performed according to the examiner's (user's) instructions by displaying multiple high-resolution image sets on a user interface provided on an arbitrary display unit.

[0042] Note that in some cases, input images that have not been enhanced may be more suitable for image diagnosis, so input images may be included in the final image selection. Furthermore, parameters may be input to the image enhancement engine along with the low-resolution images. For example, parameters specifying the degree of enhancement or the image filter size used in the image processing technique may be input to the image enhancement engine along with the input images.

[0043] A location estimation engine is a module that estimates the location or region of an image taken from an input image. The location estimation engine can output the location of the location or region depicted in the input image, or the probability that it is that location or region for each location or region label at the required level of detail.

[0044] Depending on the imaging device, the imaging area and region may not be saved as imaging conditions, or the device may not be able to acquire and save them. Furthermore, even if the imaging area and region are saved, the necessary level of detail may not be saved. For example, the imaging area may only be saved as "posterior segment," but it may not be clear whether it is the "macula," "optic nerve head," "macula and optic nerve head," or "other." In another example, the imaging area may only be saved as "breast," but it may not be clear whether it is the "right breast," "left breast," or "both." Therefore, by using an imaging area estimation engine, it is possible to estimate the imaging area and region of the input image in these cases.

[0045] The image and data processing methods that constitute the estimation method of the imaging location estimation engine perform processing using various machine learning algorithms such as deep learning. In addition to or instead of processing using machine learning algorithms, the image and data processing methods may also perform any existing estimation processing such as natural language processing, matching processing using databases of similar images and similar data, and knowledge-based processing. The training data for training the machine learning model constructed using the machine learning algorithm can be images labeled with imaging parts or regions. In this case, the training data images are used as input data, and the labels of imaging parts or regions are used as output data.

[0046] In particular, Figure 2 shows an example of a CNN configuration for estimating the location of a two-dimensional image. This CNN configuration includes a group of 200 convolutional processing blocks, each consisting of a convolutional layer 201, a batch normalization layer 202, and an activation layer 203 using a normalized linear function. Furthermore, the CNN configuration includes a final convolutional layer 204, a fully connected layer 205, and an output layer 206. The fully connected layer 205 fully connects the output values ​​of the convolutional processing block 200. The output layer 206 uses the Softmax function to output the estimated probability for each expected imaging site label for the input image Im210 as the Result 207. In such a configuration, for example, if the input image Im210 is an image of the "macula," the highest probability for the imaging site label corresponding to the "macula" will be output.

[0047] For example, by setting the number of convolutional processing blocks 200 to 16, and the parameters of the convolutional layer group 201 to a filter kernel size of 3 pixels wide, 3 pixels high, and 64 filters, it is possible to estimate the imaged area with a certain degree of accuracy. However, in practice, as described in the explanation of the machine learning model above, a better set of parameters can be set using training data appropriate to the usage of the machine learning model. Furthermore, if it is necessary to process three-dimensional or four-dimensional images, the filter kernel size may be extended to three or four dimensions. The estimation method may be performed using only one image and data processing method, or it may be performed by combining two or more image and data processing methods.

[0048] An image quality evaluation engine is a module that outputs an image quality evaluation index for an input image. The image quality evaluation processing method that calculates the image quality evaluation index uses various machine learning algorithms such as deep learning. In addition, this image quality evaluation processing method may also perform existing arbitrary evaluation processes such as an image noise measurement algorithm and matching processing using a database of image quality evaluation indices corresponding to similar images and base images. These evaluation processes may be performed in addition to or instead of the processing using machine learning algorithms.

[0049] For example, image quality evaluation indices can be obtained from a machine learning model constructed using a machine learning algorithm. In this case, the input data pairs that constitute the training data for the machine learning model consist of a set of low-resolution images and a set of high-resolution images taken in advance under various shooting conditions. The output data pairs that constitute the training data for the machine learning model are, for example, a set of image quality evaluation indices set by an examiner performing a medical image diagnosis for each of the input images.

[0050] In this description, the authenticity evaluation engine is a module that evaluates the rendering of an input image and determines, with a certain degree of accuracy, whether or not the image was captured and acquired by the target imaging device. The authenticity evaluation processing method performs processing using various machine learning algorithms such as deep learning. In addition to or instead of processing using machine learning algorithms, the authenticity evaluation processing method may also perform any existing evaluation processing such as knowledge-based processing.

[0051] For example, the authenticity evaluation process can be performed using a machine learning model constructed with a machine learning algorithm. First, let's explain the training data for the machine learning model. The training data includes pairs of high-resolution images taken in advance under various shooting conditions and labels indicating that they were taken and acquired by the target shooting device (hereinafter referred to as "authentic labels"). The training data also includes pairs of high-resolution images generated by inputting low-resolution images into a high-resolution engine (a first-level high-resolution engine) and labels indicating that they were not taken and acquired by the target shooting device (hereinafter referred to as "counterfeit labels"). A machine learning model trained using such training data will output a counterfeit label when it receives a high-resolution image generated by the first-level high-resolution engine as input.

[0052] In particular, Figure 3 shows an example of a CNN configuration for performing authenticity evaluation processing of two-dimensional images. This CNN configuration includes a group of multiple convolutional processing blocks 300, each consisting of a convolutional layer 301, a batch normalization layer 302, and an activation layer 303 using a normalized linear function. This CNN configuration also includes a final convolutional layer 304, a fully connected layer 305, and an output layer 306. The fully connected layer 305 fully connects the output values ​​of the convolutional processing blocks 300. The output layer 306 uses a sigmoid function to output a value of 1 (true) representing the genuine label or a value of 0 (false) representing the counterfeit label for the input image Im 310 as the result (Result) 307 of the authenticity evaluation processing.

[0053] By setting the number of convolutional processing blocks 300 to 16, and the parameters for the convolutional layer group 301 to a filter kernel size of 3 pixels wide and 3 pixels high, and a number of filters to 64, it is possible to obtain the correct authenticity evaluation results with a certain degree of accuracy. However, as described in the above explanation of the machine learning model, in practice, a better set of parameters can be set using training data appropriate to the usage of the machine learning model. Furthermore, if it is necessary to process three-dimensional or four-dimensional images, the filter kernel size may be extended to three or four dimensions.

[0054] The authenticity evaluation engine may output a genuine label when it receives a high-resolution image generated by a second-level image enhancement engine, which enhances image quality to a higher degree than the first-level image enhancement engine. In other words, the authenticity evaluation engine cannot definitively determine whether an image was captured and acquired by a camera, but it can determine whether an image has the characteristics of one captured and acquired by a camera. By utilizing this characteristic, the authenticity evaluation engine can evaluate whether the high-resolution image generated by the image enhancement engine is sufficiently enhanced by inputting a high-resolution image generated by the image enhancement engine.

[0055] Furthermore, the efficiency and accuracy of both engines can be improved by training the machine learning models of the image enhancement engine and the authenticity evaluation engine in coordination. In this case, first, the machine learning model of the image enhancement engine is trained so that when the high-resolution image generated by the image enhancement engine is evaluated by the authenticity evaluation engine, a genuine label is output. In parallel, the machine learning model of the authenticity evaluation engine is trained so that when the image generated by the image enhancement engine is evaluated by the authenticity evaluation engine, a counterfeit label is output. Furthermore, in parallel, the machine learning model of the authenticity evaluation engine is trained so that when the image acquired by the imaging device is evaluated by the authenticity evaluation engine, a genuine label is output. This improves the efficiency and accuracy of both the image enhancement engine and the authenticity evaluation engine.

[0056] <First Embodiment> The medical image processing apparatus according to the first embodiment will be described below with reference to Figures 4 and 5. Figure 4 shows an example of a schematic configuration of the image processing apparatus according to this embodiment.

[0057] The image processing device 400 is connected to the imaging device 10 and the display unit 20 via circuits or networks. Alternatively, the imaging device 10 and the display unit 20 may be directly connected. In this embodiment, these devices are separate devices, but some or all of these devices may be integrated into a single system. Furthermore, these devices may be connected to any other device via circuits or networks, or may be integrated with any other device.

[0058] The image processing device 400 includes an acquisition unit 401, a shooting condition acquisition unit 402, a high-image-quality enhancement feasibility determination unit 403, a high-image-quality enhancement unit 404, and an output unit 405 (display control unit). The image processing device 400 may be composed of multiple devices, each equipped with some of these components. The acquisition unit 401 can acquire various data and images from the imaging device 10 or other devices, and can also acquire input from the examiner via an input device (not shown). The input device may be a mouse, keyboard, touch panel, or any other arbitrary input device. The display unit 20 may also be configured as a touch panel display.

[0059] The shooting condition acquisition unit 402 acquires the shooting conditions of the medical image (input image) acquired by the acquisition unit 401. Specifically, it acquires a set of shooting conditions stored in the data structure that constitutes the medical image, according to the data format of the medical image. If the shooting conditions are not stored in the medical image, the acquisition unit 401 can acquire a set of shooting information, including the shooting conditions, from the shooting device 10 or the image management system.

[0060] The image quality enhancement feasibility determination unit 403 determines whether the medical image can be processed by the image quality enhancement unit 404 using the set of shooting conditions acquired by the shooting condition acquisition unit 402. The image quality enhancement unit 404 enhances the image quality of the medical images that can be processed, generating high-resolution images suitable for image diagnosis. The output unit 405 displays the high-resolution images, input images, and various information generated by the image quality enhancement unit 404 on the display unit 20. The output unit 405 may also store the generated high-resolution images, etc., in a storage device (storage unit) connected to the image processing device 400.

[0061] Next, the image enhancement unit 404 will be described in detail. The image enhancement unit 404 is equipped with an image enhancement engine. The image enhancement method provided by the image enhancement engine in this embodiment performs processing using a machine learning algorithm.

[0062] In this embodiment, training a machine learning model related to a machine learning algorithm uses training data consisting of pairs of input data, which are low-resolution images with specific shooting conditions assumed to be the target of processing, and output data, which are high-resolution images corresponding to the input data. Specifically, the specific shooting conditions include predetermined shooting areas, shooting methods, shooting angles, and image sizes.

[0063] In this embodiment, the input data for training data is a low-resolution image acquired using the same model and settings as the imaging device 10. The output data for training data is a high-resolution image acquired using the same settings and image processing capabilities as the imaging device 10. Specifically, the output data is a high-resolution image (overlay image) obtained by performing an overlay process such as averaging on a group of images (original images) acquired by taking multiple shots. Here, high-resolution and low-resolution images will be explained using OCTA motion contrast data as an example. Motion contrast data is data used in OCTA and the like, obtained by repeatedly photographing the same part of the object being photographed and detecting the temporal change of the object between those shots. At this time, by generating a frontal image using data from a desired range in the depth direction of the object being photographed from the calculated motion contrast data (an example of 3D medical image data), an OCTA En-Face image (motion contrast frontal image) can be generated. In the following, repeatedly photographing OCT data at the same location will be referred to as NOR (Number Of Repeat).

[0064] In this embodiment, two different methods for generating high-resolution and low-resolution images by overlay processing will be explained using Figure 28.

[0065] The first method, as an example of high-resolution images, will be explained using Figure 28(a) regarding motion contrast data generated from OCT data obtained by repeatedly scanning the same area of ​​the subject. In Figure 28(a), Im2810 represents 3D motion contrast data, and Im2811 represents 2D motion contrast data that constitutes the 3D motion contrast data. Im2811-1 to Im2811-3 represent OCT tomographic images (B scans) used to generate Im2811. Here, NOR in Figure 28(a) refers to the number of OCT tomographic images in Im2811-1 to Im2811-3, and in the example shown in the figure, NOR is 3. Im2811-1 to Im2811-3 are scanned at predetermined time intervals (Δt). The same area refers to one line in the frontal direction (XY) of the eye being examined, which corresponds to the area of ​​Im2811 in Figure 28(a). Note that the frontal direction is an example of a direction that intersects the depth direction. Since motion contrast data is data that detects changes over time, at least two NOR operations are required to generate this data. For example, if NOR is 2, one motion contrast data is generated. If NOR is 3, two data is generated when generating motion contrast data using only OCT data from adjacent time intervals (1st and 2nd, 2nd and 3rd). If motion contrast data is generated using OCT data from distant time intervals (1st and 3rd), a total of three data is generated. In other words, increasing NOR to 3, 4, ... increases the number of motion contrast data points at the same location. High-quality motion contrast data can be generated by aligning multiple motion contrast data obtained by repeatedly photographing the same location and performing superposition processing such as additive averaging. For this reason, it is desirable to use at least three NOR operations, and ideally five or more. On the other hand, an example of a corresponding low-quality image would be the motion contrast data before superposition processing such as additive averaging. In this case, it is desirable to use the low-resolution image as the reference image when performing superposition processing such as additive averaging.When performing the overlay process, if the position and shape of the target image are deformed relative to the reference image to perform alignment, there will be almost no spatial misalignment between the reference image and the overlaid image. Therefore, it is easy to create pairs of low-resolution and high-resolution images. Note that the target image that has undergone alignment image deformation can be used as the low-resolution image instead of the reference image. Multiple pairs can be generated by using each of the original image sets (reference image and target image) as input data and the corresponding overlay image as output data. For example, when obtaining one overlay image from a set of 15 original images, a pair of the first original image and the overlay image from the original image set can be generated, and a pair of the second original image and the overlay image from the original image set can be generated. In this way, when obtaining one overlay image from a set of 15 original images, 15 pairs can be generated, each consisting of one image from the original image set and the overlay image. Note that high-resolution 3D data can be generated by repeatedly capturing the same location in the main scan (X) direction and scanning while shifting it in the sub-scan (Y) direction.

[0066] The second method, which involves generating high-quality images by overlaying motion contrast data obtained from multiple images of the same region of the subject, is explained using Figure 28(b). The same region refers to an area such as 3x3mm or 10x10mm in the frontal direction (XY) of the eye under examination, and means acquiring three-dimensional motion contrast data including the depth direction of the tomographic image. When overlaying images of the same region taken multiple times, it is desirable to use two or three NOR (Natural Orbital) settings to shorten the time per shot. Furthermore, to generate high-quality three-dimensional motion contrast data, at least two sets of three-dimensional data from the same region should be acquired. Figure 28(b) shows an example of multiple sets of three-dimensional motion contrast data. Im2820~Im2840 are three-dimensional motion contrast data, similar to those explained in Figure 28(a). Using these two or more sets of three-dimensional motion contrast data, alignment processing is performed in the frontal direction (XY) and depth direction (Z), and after removing artifact data from each set, averaging is performed. This allows for the generation of a single high-resolution 3D motion contrast data set free of artifacts. A high-resolution image can be obtained by generating an arbitrary plane from the 3D motion contrast data. On the other hand, it is desirable that the corresponding low-resolution image be an arbitrary plane generated from the reference data used when performing superposition processing such as averaging. As explained in the first method, there is almost no spatial displacement between the reference image and the image after averaging, so it is easy to create a pair of low-resolution and high-resolution images. Alternatively, the low-resolution image may be an arbitrary plane generated from the target data that has undergone image deformation processing for alignment, rather than the reference data.

[0067] The first method minimizes the burden on the subject because the shooting itself is completed in a single take. However, increasing the number of NOR (Natural Orientation) scans increases the time required for each scan. Also, if artifacts such as eye clouding or eyelashes are present during the scan, a good image is not always obtained. The second method involves taking multiple images, which slightly increases the burden on the subject. However, the time required for each image is short, and even if an artifact is captured in one image, if no artifact is captured in subsequent images, a clean image with fewer artifacts can ultimately be obtained. Considering these characteristics, the appropriate method should be selected when collecting data, depending on the subject's circumstances.

[0068] In this embodiment, motion contrast data was used as an example, but the method is not limited to this. Since OCT data is captured to generate motion contrast data, the same method can be used with OCT data as well. Furthermore, although the tracking process was omitted in this embodiment, it is desirable to perform tracking of the eye while capturing images, as the same location or region of the eye under examination is captured.

[0069] In this embodiment, since pairs of high-resolution and low-resolution 3D data are created, any pair of 2D images can be generated from them. This will be explained using Figure 29. For example, if the target image is an OCTA En-Face image, an OCTA En-Face image is generated from the 3D data within a desired depth range. The desired depth range refers to the range in the Z direction in Figure 28. An example of an OCTA En-Face image generated here is shown in Figure 29(a). Learning is performed using OCTA En-Face images generated in different depth ranges, such as the surface layer (Im2910), deep layer (Im2920), outer layer (Im2930), and choroidal vascular network (Im2940). Note that the types of OCTA En-Face images are not limited to these; the types can be increased by generating OCTA En-Face images with different depth ranges by changing the reference layer and offset values. When performing training, each OCTA En-Face image at a different depth may be trained separately, multiple images from different depth ranges may be combined (for example, separated into surface and deep layers), or OCTA En-Face images from all depth ranges may be trained together. In the case of luminance En-Face images generated from OCT data, training is performed using multiple En-Face images generated from an arbitrary depth range, similar to OCTA En-Face images. For example, consider a case where the image enhancement engine includes a machine learning engine obtained using training data that includes multiple motion contrast frontal images corresponding to different depth ranges of the eye being examined. In this case, the acquisition unit can acquire a motion contrast frontal image corresponding to a part of the depth range within a long depth range that includes different depth ranges as the first image. That is, a motion contrast frontal image corresponding to a depth range different from the multiple depth ranges corresponding to the multiple motion contrast frontal images included in the training data can be used as the input image for image enhancement. Of course, a motion contrast frontal image from the same depth range as during training may also be used as the input image for image enhancement. Furthermore, some depth ranges may be set by the examiner pressing any button on the user interface, or they may be set automatically.Furthermore, the above information is not limited to motion contrast front images; it can also be applied to, for example, luminance En-Face images.

[0070] When the images to be processed are tomographic images, training is performed using OCT tomographic images (B-scans) or tomographic images from motion contrast data. This will be explained using Figure 29(b). In Figure 29(b), Im2951 to Im2953 are OCT tomographic images. The images in Figure 29(b) are different because they show tomographic images from different locations in the subscan (Y) direction. In the case of tomographic images, it is acceptable to train them together without worrying about the difference in location in the subscan direction. However, if the images were taken from different locations (e.g., the macula center, the optic nerve head center), it is acceptable to train each location separately, or to train them together without worrying about the location. Note that OCT tomographic images and tomographic images from motion contrast data have significantly different image features, so it is better to train them separately.

[0071] The superimposed image, created by overlay processing, has pixels that are commonly depicted in the original images emphasized, resulting in a high-quality image suitable for medical imaging. In this case, the resulting high-quality image has a high contrast, with a clear distinction between low-luminance and high-luminance areas, as a result of the emphasis on commonly depicted pixels. Furthermore, in superimposed images, random noise that occurs with each capture can be reduced, and areas that were not well depicted in the original image at a certain point in time can be interpolated by other original images.

[0072] Furthermore, if the input data for a machine learning model needs to consist of multiple images, the required number of original image sets can be selected from the original image set and used as input data. For example, if one superimposed image is obtained from 15 original image sets, and two images are needed as input data for the machine learning model, then 105 (15C2=105) pairs can be generated.

[0073] Furthermore, pairs that do not contribute to image quality improvement can be removed from the training data. For example, if the high-quality images that constitute a pair of training data are of unsuitable quality for medical imaging, the image quality improvement engine trained using that training data may also produce images of unsuitable quality for medical imaging. Therefore, by removing pairs whose output data is of unsuitable quality for medical imaging from the training data, the possibility of the image quality improvement engine generating images of unsuitable quality for medical imaging can be reduced.

[0074] Furthermore, if the average brightness and brightness distribution of a pair of images differ significantly, the image enhancement engine trained using that training data may output an image with a brightness distribution that is significantly different from the low-resolution image, making it unsuitable for image diagnosis. For this reason, pairs of input and output data with significantly different average brightness and brightness distributions can be removed from the training data.

[0075] Furthermore, if the structure and position of the subject depicted in a pair of images differ significantly, the image enhancement engine trained using this training data may output an image unsuitable for diagnostic imaging, where the subject is depicted in a structure and position that differs significantly from the low-resolution image. For this reason, pairs of input and output data where the structure and position of the depicted subject differ significantly can be removed from the training data. In addition, from the perspective of maintaining quality, the image enhancement engine can be configured not to use its own output high-resolution images as training data.

[0076] By using a machine learning-based image enhancement engine in this manner, the image enhancement unit 404 can output a high-quality image that has been enhanced with high contrast and noise reduction through overlay processing when a medical image acquired in a single exposure is input. Therefore, the image enhancement unit 404 can generate a high-quality image suitable for image diagnosis based on the low-quality input image.

[0077] Next, the series of image processing according to this embodiment will be described with reference to the flowchart in Figure 5. Figure 5 is a flowchart of the series of image processing according to this embodiment. First, when the series of image processing according to this embodiment is started, the process moves to step S510.

[0078] In step S510, the acquisition unit 401 acquires an image captured by the imaging device 10, which is connected via a circuit or network, as an input image. The acquisition unit 401 may also acquire an input image in response to a request from the imaging device 10. Such requests may be issued, for example, when the imaging device 10 generates an image, before or after saving the image generated by the imaging device 10 to a storage device provided by the imaging device 10, when displaying the saved image on the display unit 20, or when using a high-resolution image for image analysis processing.

[0079] The acquisition unit 401 may acquire data for generating an image from the imaging device 10, and the image processing device 400 may acquire the image generated based on that data as the input image. In this case, the image processing device 400 may employ any existing image generation method for generating various images.

[0080] In step S520, the shooting condition acquisition unit 402 acquires a set of shooting conditions for the input image. Specifically, the system acquires a set of shooting conditions stored in the data structure that constitutes the input image, depending on the data format of the input image. As mentioned above, if no shooting conditions are stored in the input image, the shooting condition acquisition unit 402 can acquire a set of shooting information, including the shooting conditions, from the shooting device 10 or an image management system (not shown).

[0081] In step S530, the image quality enhancement feasibility determination unit 403 uses the acquired set of shooting conditions to determine whether the input image can be enhanced in quality by the image quality enhancement engine provided in the image quality enhancement unit 404. Specifically, the image quality enhancement feasibility determination unit 403 determines whether the shooting area, shooting method, shooting angle of view, and image size of the input image match the conditions that can be handled by the image quality enhancement engine.

[0082] The image quality enhancement feasibility determination unit 403 determines all shooting conditions, and if it determines that the conditions can be addressed, the process proceeds to step S540. On the other hand, if the image quality enhancement feasibility determination unit 403 determines, based on these shooting conditions, that the image quality enhancement engine cannot address the input image, the process proceeds to step S550.

[0083] Furthermore, depending on the settings and implementation of the image processing device 400, even if it is determined that the input image is unprocessable based on some of the shooting area, shooting method, shooting angle of view, and image size, the image enhancement processing in step S540 may still be performed. For example, if the image enhancement engine is assumed to be able to comprehensively handle any shooting area of ​​the subject and is implemented to handle cases where the input data includes an unknown shooting area, such processing may be performed. In addition, the image enhancement feasibility determination unit 403 may determine, depending on the desired configuration, whether at least one of the shooting area, shooting method, shooting angle of view, and image size of the input image matches the conditions that can be handled by the image enhancement engine.

[0084] In step S540, the image enhancement unit 404 uses an image enhancement engine to enhance the image quality of the input image, generating a higher-quality image more suitable for image diagnosis than the input image. Specifically, the image enhancement unit 404 inputs the input image to the image enhancement engine and generates a high-quality image. The image enhancement engine generates a high-quality image that appears as if it has been overlaid using the input image, based on a machine learning model that has been machine-learned using training data. As a result, the image enhancement engine can generate a higher-quality image that has reduced noise and enhanced contrast than the input image.

[0085] Depending on the settings and implementation configuration of the image processing device 400, the image enhancement unit 404 may input parameters to the image enhancement engine along with the input image according to the shooting conditions, thereby adjusting the degree of image enhancement. Alternatively, the image enhancement unit 404 may input parameters corresponding to the examiner's input to the image enhancement engine along with the input image, thereby adjusting the degree of image enhancement.

[0086] In step S550, if a high-resolution image was generated in step S540, the output unit 405 outputs the high-resolution image and displays it on the display unit 20. On the other hand, if it was determined in step S530 that high-resolution processing was not possible, the output unit 405 outputs the input image and displays it on the display unit 20. Alternatively, instead of displaying the output image on the display unit 20, the output unit 405 may display or store the output image on the imaging device 10 or other devices. Furthermore, depending on the settings and implementation configuration of the image processing device 400, the output unit 405 may process the output image so that it can be used by the imaging device 10 or other devices, or convert the data format so that it can be transmitted to an image management system or the like.

[0087] As described above, the image processing apparatus 400 according to this embodiment comprises an acquisition unit 401 and an image enhancement unit 404. The acquisition unit 401 acquires an input image (first image), which is an image of a predetermined part of the subject. The image enhancement unit 404 uses an image enhancement engine, including a machine learning engine, to generate a high-resolution image (second image) from the input image, which has at least one of noise reduction and contrast enhancement compared to the input image. The image enhancement engine includes a machine learning engine that uses images obtained by overlay processing as training data.

[0088] With this configuration, the image processing device 400 according to this embodiment can output high-quality images from the input image, in which noise is reduced and contrast is enhanced. Therefore, the image processing device 400 can acquire images suitable for diagnostic imaging, such as clearer images or images in which the area or lesion to be observed is emphasized, with less cost than conventional methods, without increasing the invasiveness or effort required of the photographer or the patient.

[0089] Furthermore, the image processing device 400 includes an image enhancement feasibility determination unit 403 that determines whether or not a high-quality image can be generated from an input image using an image enhancement engine. The image enhancement feasibility determination unit 403 makes this determination based on at least one of the following: the area of ​​the input image being captured, the shooting method, the shooting angle of view, and the image size.

[0090] With this configuration, the image processing apparatus 400 according to this embodiment can exclude input images that the image enhancement unit 404 cannot process from the image enhancement process, thereby reducing the processing load and error occurrence of the image processing apparatus 400.

[0091] In this embodiment, the output unit 405 (display control unit) is configured to display the generated high-resolution image on the display unit 20, but the operation of the output unit 405 is not limited to this. For example, the output unit 405 can also output the high-resolution image to other devices connected to the shooting device 10 or the image processing device 400. As a result, the high-resolution image can be displayed on the user interface of these devices, stored in any storage device, used for any image analysis, or transmitted to an image management system.

[0092] In this embodiment, the image quality enhancement feasibility determination unit 403 determines whether the input image is one that can be enhanced in image quality by the image quality enhancement engine, and if it is an input image that can be enhanced in image quality, the image quality enhancement unit 404 performs the enhancement. On the other hand, if the shooting device 10 only takes pictures under shooting conditions that allow for image quality enhancement, the image acquired from the shooting device 10 may be enhanced in image quality unconditionally. In this case, as shown in Figure 6, the processing of steps S520 and S530 can be omitted, and step S540 can be performed after step S510.

[0093] In this embodiment, the output unit 405 is configured to display a high-resolution image on the display unit 20. However, the output unit 405 may also display a high-resolution image on the display unit 20 in response to instructions from the examiner. For example, the output unit 405 may display a high-resolution image on the display unit 20 in response to the examiner pressing any button on the user interface of the display unit 20. In this case, the output unit 405 may switch between the input image and the high-resolution image, or it may display the high-resolution image alongside the input image.

[0094] Furthermore, when the output unit 405 displays a high-resolution image on the display unit 20, it may also display a notification indicating that the displayed image is a high-resolution image generated by processing using a machine learning algorithm, along with the high-resolution image. In this case, the user can easily identify from the notification that the displayed high-resolution image is not the image acquired by photography, thereby reducing misdiagnosis and improving diagnostic efficiency. The notification indicating that the high-resolution image was generated by processing using a machine learning algorithm can take any form as long as it is a notification that can distinguish between the input image and the high-resolution image generated by the processing.

[0095] Furthermore, the output unit 405 may display on the display unit 20 an indication that the image is a high-resolution image generated by processing using a machine learning algorithm, indicating what kind of training data the machine learning algorithm was trained on. This indication may include an explanation of the types of input and output data of the training data, and any indication of the training data such as the imaged body parts included in the input and output data.

[0096] In the high-resolution engine according to this embodiment, superimposed images were used as the output data for training data, but the training data is not limited to this. High-resolution images obtained by performing at least one of the means for obtaining high-resolution images, such as superimposition processing, the processing group described later, or the shooting method described later, may be used as the output data for training data.

[0097] For example, high-resolution images obtained by performing maximum posterior probability estimation (MAP estimation) on the original image set may be used as the output data for training data. In MAP estimation, a likelihood function is calculated from the probability density of each pixel value in multiple low-resolution images, and the true signal value (pixel value) is estimated using the calculated likelihood function.

[0098] High-resolution images obtained through MAP estimation are high-contrast images based on pixel values ​​close to the true signal values. Furthermore, since the estimated signal values ​​are determined based on probability density, randomly occurring noise is reduced in high-resolution images obtained through MAP estimation. Therefore, by using high-resolution images obtained through MAP estimation as training data, the image enhancement engine can generate high-resolution images suitable for medical imaging, with reduced noise and high contrast, from the input image. The method for generating pairs of input and output data for training data may be the same as when using superimposed images as training data.

[0099] Furthermore, a high-resolution image obtained by applying a smoothing filter to the original image may be used as the output data for training data. In this case, the image enhancement engine can generate a high-resolution image with reduced random noise from the input image. Additionally, an image obtained by applying a grayscale conversion process to the original image may be used as the output data for training data. In this case, the image enhancement engine can generate a high-resolution image with enhanced contrast from the input image. The method for generating pairs of input and output data for training data may be the same as when using superimposed images as training data.

[0100] The input data for training data may be images acquired from an imaging device having the same image quality characteristics as imaging device 10. Furthermore, the output data for training data may be high-resolution images obtained through costly processing such as iterative approximation, or high-resolution images acquired by imaging subjects corresponding to the input data with an imaging device more advanced than imaging device 10. Furthermore, the output data may be high-resolution images obtained by performing rule-based noise reduction processing. Here, the noise reduction processing may include, for example, replacing a single high-luminance pixel that is clearly noise and appears in a low-luminance region with the average value of neighboring low-luminance pixels. For this reason, the image enhancement engine may use images captured by an imaging device that is more powerful than the imaging device used to capture the input image, or images acquired in an imaging process that is more complex than the input image capture process, as training data. For example, when the image enhancement engine uses a motion contrast front image as the input image, it may use images obtained by OCTA imaging using an OCT imaging device that is more powerful than the OCT imaging device used to capture the input image, or images obtained in an OCT imaging process that is more complex than the input image OCTA capture process, as training data.

[0101] Although omitted in the description of this embodiment, the high-resolution image generated from multiple images, which is used as the output data for training data, can be generated from multiple images that have already been aligned. For example, as part of the alignment process, one of the multiple images may be selected as a template, the similarity with the other images may be determined while changing the position and angle of the template, the amount of misalignment with the template may be determined, and each image may be corrected based on the amount of misalignment. Alternatively, other existing arbitrary alignment processes may be performed.

[0102] Furthermore, when aligning a three-dimensional image, the three-dimensional image may be decomposed into multiple two-dimensional images, and the alignment of each two-dimensional image may be performed before integrating them. Alternatively, the two-dimensional image may be decomposed into one-dimensional images, and the alignment of each one-dimensional image may be performed before integrating them. These alignment procedures may also be performed on the data used to generate the image, rather than on the image itself.

[0103] In this embodiment, if the image quality enhancement feasibility determination unit 403 determines that the input image can be processed by the image quality enhancement unit 404, the process moves to step S540 and the image quality enhancement process by the image quality enhancement unit 404 is started. Alternatively, the output unit 405 may display the determination result from the image quality enhancement feasibility determination unit 403 on the display unit 20, and the image quality enhancement unit 404 may start the image quality enhancement process in response to instructions from the examiner. In this case, the output unit 405 can display the input image and shooting conditions such as the shooting area acquired for the input image on the display unit 20 along with the determination result. In this case, since the image quality enhancement process is performed after the examiner has determined whether the determination result is correct or not, it is possible to reduce image quality enhancement processing based on misdeterminance.

[0104] Alternatively, the image enhancement feasibility determination unit 403 may not perform a determination, and the output unit 405 may display the input image and the shooting conditions such as the shooting area acquired for the input image on the display unit 20, and the image enhancement unit 404 may start the image enhancement process in response to instructions from the examiner.

[0105] <Second Embodiment> Next, an image processing apparatus according to the second embodiment will be described with reference to Figures 4 and 7. In the first embodiment, the image enhancement unit 404 was equipped with a single image enhancement engine. In contrast, in this embodiment, the image enhancement unit is equipped with multiple image enhancement engines that have performed machine learning using different training data, and generates multiple high-resolution images from an input image.

[0106] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0107] The image enhancement unit 404 according to this embodiment is equipped with two or more image enhancement engines that have been machine-learned using different training data. Here, the method for creating the training data set according to this embodiment will be described. Specifically, first, a group of pairs of original images as input data and superimposed images as output data are prepared, each of which is captured from various shooting areas. Next, a training data set is created by grouping the pair groups by shooting area. For example, a first training data set is created consisting of a pair group obtained by capturing a first shooting area, and a second training data set is created consisting of a pair group obtained by capturing a second shooting area.

[0108] Subsequently, separate image enhancement engines are trained using each set of training data. For example, a group of image enhancement engines is prepared, such as a first image enhancement engine corresponding to a machine learning model trained with the first set of training data, and a second image enhancement engine corresponding to a machine learning model trained with the second set of training data.

[0109] Because each of these image enhancement engines uses different training data for its corresponding machine learning model, the degree to which it can enhance the image quality of an input image varies depending on the shooting conditions of the image input to the engine. Specifically, the first image enhancement engine enhances the image quality to a high degree for input images obtained by shooting a first shooting area, but to a low degree for images obtained by shooting a second shooting area. Similarly, the second image enhancement engine enhances the image quality to a high degree for input images obtained by shooting a second shooting area, but to a low degree for images obtained by shooting a first shooting area.

[0110] Since each training data set is composed of pairs grouped by the area being photographed, the image quality tendencies of the image groups constituting each pair are similar. For this reason, the image quality enhancement engine can enhance image quality more effectively than the image quality enhancement engine according to the first embodiment, provided that the area being photographed corresponds to the same area being photographed. Note that the shooting conditions for grouping the training data pairs are not limited to the area being photographed, but may also be the shooting angle of view, the image resolution, or a combination of two or more of these.

[0111] The series of image processing steps according to this embodiment will now be described with reference to Figure 7. Figure 7 is a flowchart of the series of image processing steps according to this embodiment. Note that the processing in steps S710 and S720 is the same as in steps S510 and S520 of the first embodiment, so the explanation will be omitted. Note that if the input image is to be unconditionally enhanced in image quality, the processing in step S730 may be omitted after the processing in step S720, and the processing may proceed to step S740.

[0112] In step S720, once the shooting conditions for the input image are acquired, the process moves to step S730. In step S730, the image quality enhancement feasibility determination unit 403 uses the set of shooting conditions acquired in step S720 to determine whether any of the image quality enhancement engines provided by the image quality enhancement unit 404 can process the input image.

[0113] If the image quality enhancement feasibility determination unit 403 determines that none of the image quality enhancement engines can process the input image, the process proceeds to step S760. On the other hand, if the image quality enhancement feasibility determination unit 403 determines that any of the image quality enhancement engines can process the input image, the process proceeds to step S740. Depending on the settings and implementation configuration of the image processing device 400, step S740 may also be performed, as in the first embodiment, even if the image quality enhancement engine determines that some shooting conditions cannot be processed.

[0114] In step S740, the image enhancement unit 404 selects an image enhancement engine from the image enhancement engine group based on the shooting conditions of the input image acquired in step S720 and the training data information of the image enhancement engine group. Specifically, for example, for the shooting area among the shooting conditions acquired in step S720, it selects an image enhancement engine that has training data information regarding the shooting area or surrounding shooting areas and has a high degree of image enhancement. In the example above, if the shooting area is the first shooting area, the image enhancement unit 404 selects the first image enhancement engine.

[0115] In step S750, the image enhancement unit 404 generates a high-resolution image by enhancing the image quality of the input image using the image enhancement engine selected in step S740. Then, in step S760, if a high-resolution image was generated in step S750, the output unit 405 outputs the high-resolution image and displays it on the display unit 20. On the other hand, if it was determined in step S730 that image enhancement processing was not possible, the output unit 405 outputs the input image and displays it on the display unit 20. When displaying the high-resolution image on the display unit 20, the output unit 405 may indicate that it is a high-resolution image generated using the image enhancement engine selected by the image enhancement unit 404.

[0116] As described above, the image enhancement unit 404 according to this embodiment includes a plurality of image enhancement engines, each trained using different training data. Here, each of the plurality of image enhancement engines is trained using different training data for at least one of the following: shooting area, shooting angle of view, front images at different depths, and image resolution. The image enhancement unit 404 generates a high-resolution image using an image enhancement engine corresponding to at least one of the following: shooting area, shooting angle of view, front images at different depths, and image resolution of the input image.

[0117] With this configuration, the image processing apparatus 400 according to this embodiment can generate more effective high-resolution images.

[0118] In this embodiment, the image enhancement unit 404 selects an image enhancement engine to be used for image enhancement processing based on the shooting conditions of the input image, but the process of selecting an image enhancement engine is not limited to this. For example, the output unit 405 may display the shooting conditions of the acquired input image and the group of image enhancement engines on the user interface of the display unit 20, and the image enhancement unit 404 may select an image enhancement engine to be used for image enhancement processing in response to instructions from the examiner. The output unit 405 may also display information on the training data used to train each image enhancement engine on the display unit 20 along with the group of image enhancement engines. The manner in which the information on the training data used to train the image enhancement engines is displayed is arbitrary, and for example, the group of image enhancement engines may be displayed using a name related to the training data used for training.

[0119] Furthermore, the output unit 405 may display the image enhancement engine selected by the image enhancement unit 404 on the user interface of the display unit 20 and accept instructions from the examiner. In this case, the image enhancement unit 404 may decide whether or not to ultimately select the image enhancement engine as the image enhancement engine to be used for image enhancement processing, in response to the instructions from the examiner.

[0120] The output unit 405 may output the generated high-resolution images to the imaging device 10 or other devices connected to the image processing device 400, as in the first embodiment. Furthermore, the output data for the training data of the high-resolution engine is not limited to high-resolution images that have undergone overlay processing, as in the first embodiment. That is, high-resolution images obtained by performing at least one of the processing group or imaging method, such as overlay processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, imaging using a high-performance imaging device, high-cost processing, or noise reduction processing, may be used.

[0121] <Third Embodiment> Next, an image processing apparatus according to the third embodiment will be described with reference to Figures 4 and 7. In the first and second embodiments, the shooting condition acquisition unit 402 acquires a set of shooting conditions from the data structure of the input image, etc. In contrast, in this embodiment, the shooting condition acquisition unit uses a shooting location estimation engine to estimate the shooting location or shooting area of ​​the input image based on the input image.

[0122] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the second embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the second embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first and second embodiments, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0123] The shooting condition acquisition unit 402 according to this embodiment is equipped with a shooting location estimation engine that estimates the shooting area or shooting region drawn on the input image acquired by the acquisition unit 401. The shooting location estimation method provided by the shooting location estimation engine according to this embodiment performs estimation processing using a machine learning algorithm.

[0124] In this embodiment, training a machine learning model for a machine learning algorithm-based imaging location estimation method uses training data consisting of pairs of input data (images) and output data (imaging location labels or imaging area labels) corresponding to the input data. Here, input data refers to images with specific imaging conditions assumed to be the target of processing (input images). Preferably, the input data is an image acquired from an imaging device having the same image quality characteristics as imaging device 10, and even better if it is the same model with the same settings as imaging device 10. The types of imaging location labels and imaging area labels in the output data may be imaging locations or imaging areas that are included in at least part of the input data. For example, in the case of OCT, the types of imaging location labels in the output data may be "macula," "optic nerve head," "macula and optic nerve head," and "other."

[0125] The shooting location estimation engine according to this embodiment can output which parts and regions are depicted in the input image by learning using such training data. The shooting location estimation engine can also output the probability that each shooting location label or region is a shooting location or region at the required level of detail. By using the shooting location estimation engine, the shooting condition acquisition unit 402 can estimate the shooting locations and regions of the input image based on the input image and acquire them as shooting conditions for the input image. When the shooting location estimation engine outputs the probability that each shooting location label or region is a shooting location or region, the shooting condition acquisition unit 402 acquires the shooting location or region with the highest probability as the shooting condition for the input image.

[0126] Next, as with the second embodiment, the series of image processing according to this embodiment will be described with reference to the flowchart in Figure 7. Note that the processing in steps S710 and S730 to S760 according to this embodiment is the same as the processing in the second embodiment, so the explanation will be omitted. Note that if the input image is to be unconditionally enhanced in image quality, the processing in step S730 may be omitted after the processing in step S720, and the processing may proceed to step S740.

[0127] When the input image is acquired in step S710, the process moves to step S720. In step S720, the shooting condition acquisition unit 402 acquires the shooting conditions for the input image acquired in step S710.

[0128] Specifically, the system retrieves a set of shooting conditions stored in the data structure that constitutes the input image, depending on the data format of the input image. If the set of shooting conditions does not include information about the shooting area or region, the shooting condition acquisition unit 402 inputs the input image to the shooting location estimation engine and estimates which shooting area the input image was taken from. Specifically, the shooting condition acquisition unit 402 inputs the input image to the shooting location estimation engine, evaluates the probability output for each of the shooting area labels, and sets and acquires the shooting area with the highest probability as the shooting condition for the input image.

[0129] If the input image does not contain any shooting conditions other than the shooting area or shooting region, the shooting condition acquisition unit 402 can acquire a set of shooting information, including a set of shooting conditions, from the shooting device 10 or an image management system (not shown).

[0130] The subsequent processing is the same as the series of image processing according to the second embodiment, so we will omit the explanation.

[0131] As described above, the shooting condition acquisition unit 402 according to this embodiment functions as an estimation unit that estimates at least one of the shooting area and shooting region of the input image. The shooting condition acquisition unit 402 includes a shooting location estimation engine that uses images labeled with the shooting area and shooting region as training data, and estimates the shooting area and shooting region of the input image by inputting the input image to the shooting location estimation engine.

[0132] As a result, the image processing device 400 according to this embodiment can acquire shooting conditions for the shooting area and shooting region of the input image based on the input image.

[0133] In this embodiment, the shooting condition acquisition unit 402 used the shooting location estimation engine to estimate the shooting location and shooting area of ​​the input image when the shooting condition group did not contain information about the shooting location or shooting area. However, the situation in which the shooting location estimation engine is used to estimate the shooting location and shooting area is not limited to this. The shooting condition acquisition unit 402 may also use the shooting location estimation engine to estimate the shooting location and shooting area when the information about the shooting location and shooting area included in the data structure of the input image is insufficient to provide the necessary level of detail.

[0134] Furthermore, regardless of whether the data structure of the input image contains information about the shooting area or shooting region, the shooting condition acquisition unit 402 may use the shooting location estimation engine to estimate the shooting area or shooting region of the input image. In this case, the output unit 405 may display the estimation results output from the shooting location estimation engine and the information about the shooting area or shooting region contained in the data structure of the input image on the display unit 20, and the shooting condition acquisition unit 402 may determine these shooting conditions according to the examiner's instructions.

[0135] The output unit 405 may output the generated high-resolution images to the imaging device 10 or other devices connected to the image processing device 400, as in the first embodiment. Furthermore, the output data for the training data of the high-resolution engine is not limited to high-resolution images that have undergone overlay processing, as in the first embodiment. That is, high-resolution images obtained by performing at least one of the processing group or imaging method, such as overlay processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, imaging using a high-performance imaging device, high-cost processing, or noise reduction processing, may be used.

[0136] <Fourth Embodiment> Next, an image processing apparatus according to the fourth embodiment will be described with reference to Figures 4, 5, 8, and 9. In this embodiment, the image enhancement unit enlarges or reduces the input image so that it becomes an image size that the image enhancement engine can handle. The image enhancement unit also generates a high-resolution image by reducing or enlarging the output image from the image enhancement engine so that the image size of the output image becomes the same as the image size of the input image.

[0137] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0138] The image enhancement unit 404 according to this embodiment is equipped with an image enhancement engine similar to the image enhancement engine according to the first embodiment. However, in this embodiment, as training data used for learning the image enhancement engine, a group of input data and output data pairs is used, which consists of a group of images obtained by enlarging or reducing the input data image and the output data image to a certain image size.

[0139] Here, with reference to Figure 8, the training data for the high-resolution engine according to this embodiment will be described. As shown in Figure 8, for example, consider a case where there is a low-resolution image Im810 and a high-resolution image Im820 that are smaller than a certain image size set for the training data. In this case, the low-resolution image Im810 and the high-resolution image Im820 are enlarged to the certain image size set for the training data. Then, the enlarged low-resolution image Im811 and the enlarged high-resolution image Im821 are used as a pair, and this pair is used as one of the training data.

[0140] Similar to the first embodiment, the input data for the training data uses images with specific shooting conditions assumed to be the target of processing (input images). These specific shooting conditions are predetermined shooting area, shooting method, and shooting angle of view. In other words, unlike the first embodiment, the specific shooting conditions in this embodiment do not include image size.

[0141] In this embodiment, the image enhancement unit 404 uses an image enhancement engine trained with such training data to enhance the image quality of an input image and generate a high-resolution image. At this time, the image enhancement unit 404 generates a distorted image by enlarging or shrinking the input image to a certain image size set for the training data, and inputs the distorted image to the image enhancement engine. The image enhancement unit 404 also reduces or enlarges the output image from the image enhancement engine to the image size of the input image, thereby generating a high-resolution image. Therefore, in this embodiment, the image enhancement unit 404 can enhance the image quality of an input image with the image enhancement engine and generate a high-resolution image even if the image size could not be handled in the first embodiment.

[0142] Next, a series of image processing steps according to this embodiment will be described with reference to Figures 5 and 9. Figure 9 is a flowchart of the image quality enhancement process according to this embodiment. Note that the processes in steps S510, S520, and S550 according to this embodiment are the same as those in the first embodiment, so their explanation will be omitted. Note that if the image quality enhancement is to be performed unconditionally on the input image with respect to shooting conditions other than image size, the process in step S530 may be omitted after the process in step S520, and the process may proceed to step S540.

[0143] In step S520, as in the first embodiment, once the shooting condition acquisition unit 402 acquires a set of shooting conditions for the input image, the process moves to step S530. In step S530, the image quality enhancement feasibility determination unit 403 uses the acquired set of shooting conditions to determine whether the image quality enhancement engine in the image quality enhancement unit 404 can process the input image. Specifically, the image quality enhancement feasibility determination unit 403 determines whether the shooting area, shooting method, and shooting angle of view of the input image are suitable for processing by the image quality enhancement engine. Unlike the first embodiment, the image quality enhancement feasibility determination unit 403 does not determine the image size.

[0144] The image enhancement feasibility determination unit 403 determines the shooting area, shooting method, and shooting angle of view. If it determines that the input image is processable, the process proceeds to step S540. On the other hand, if the image enhancement feasibility determination unit 403 determines, based on these shooting conditions, that the image enhancement engine cannot process the input image, the process proceeds to step S550. Depending on the settings and implementation configuration of the image processing device 400, the image enhancement process in step S540 may be performed even if it is determined that the input image is processable based on some of the shooting area, shooting method, and shooting angle of view.

[0145] When the process moves to step S540, the image enhancement process according to this embodiment, as shown in Figure 9, is started. In the image enhancement process according to this embodiment, first, in step S910, the image enhancement unit 404 enlarges or reduces the input image to a certain image size set for the training data and generates a distorted image.

[0146] Next, in step S920, the image enhancement unit 404 inputs the generated deformed image to the image enhancement engine and obtains a high-quality deformed image that has been enhanced.

[0147] Subsequently, in step S930, the image enhancement unit 404 reduces or enlarges the high-quality deformed image to the image size of the input image to generate a high-quality image. Once the image enhancement unit 404 has generated a high-quality image in step S930, the image enhancement process according to this embodiment is completed, and the process moves on to step S550. The process in step S550 is the same as in step S550 of the first embodiment, so its description is omitted.

[0148] As described above, the image enhancement unit 404 in this embodiment adjusts the image size of the input image to an image size that the image enhancement engine can handle and inputs it to the image enhancement engine. The image enhancement unit 404 also generates a high-resolution image by adjusting the output image from the image enhancement engine to the original image size of the input image. As a result, the image processing device 400 in this embodiment can use the image enhancement engine to enhance the image quality of input images with image sizes that could not be handled in the first embodiment, and generate high-resolution images suitable for image diagnosis.

[0149] The output unit 405 may output the generated high-resolution images to the imaging device 10 or other devices connected to the image processing device 400, as in the first embodiment. Furthermore, the output data for the training data of the high-resolution engine is not limited to high-resolution images that have undergone overlay processing, as in the first embodiment. That is, high-resolution images obtained by performing at least one of the processing group or imaging method, such as overlay processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, imaging using a high-performance imaging device, high-cost processing, or noise reduction processing, may be used.

[0150] <Fifth Embodiment> Next, an image processing apparatus according to a fifth embodiment will be described with reference to Figures 4, 5, 10, and 11. In this embodiment, the image enhancement unit generates a high-resolution image by image enhancement processing based on a certain resolution using an image enhancement engine.

[0151] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0152] The image enhancement unit 404 according to this embodiment is equipped with an image enhancement engine similar to that in the first embodiment. However, in this embodiment, the training data used to train the image enhancement engine is different from the training data in the first embodiment. Specifically, the image group constituting the pair of input data and output data of the training data is enlarged or reduced to an image size such that the resolution of the image group becomes a certain resolution, and then padded to a sufficiently large constant image size. Here, the resolution of the image group refers to, for example, the spatial resolution of the imaging device or the resolution with respect to the imaging area.

[0153] Here, with reference to Figure 10, the training data for the high-resolution engine according to this embodiment will be described. As shown in Figure 10, for example, consider the case where there is a low-resolution image Im1010 and a high-resolution image Im1020, both having a resolution lower than a certain resolution set for the training data. In this case, both the low-resolution image Im1010 and the high-resolution image Im1020 are enlarged to the certain resolution set for the training data. Furthermore, the enlarged low-resolution image Im1010 and the high-resolution image Im1020 are padded to the certain image size set for the training data. Then, the enlarged and padded low-resolution image Im1011 and high-resolution image Im1021 are paired, and this pair is used as one of the training data.

[0154] The specified image size for the training data is the maximum possible image size when the image intended for processing (input image) is enlarged or reduced to a certain resolution. If this specified image size is not sufficiently large, the image size may become unmanageable for the machine learning model when the image input to the image enhancement engine is enlarged.

[0155] Furthermore, the areas where padding is performed are filled with a fixed pixel value, filled with neighboring pixel values, or mirror padding is used, in accordance with the characteristics of the machine learning model, in order to effectively improve image quality. As with the first embodiment, the input data uses images with specific shooting conditions assumed to be the target of processing, but these specific shooting conditions are predetermined shooting area, shooting method, and shooting angle of view. In other words, unlike the first embodiment, the specific shooting conditions in this embodiment do not include image size.

[0156] In this embodiment, the image enhancement unit 404 uses an image enhancement engine trained with such training data to enhance the image quality of the input image and generate a high-resolution image. At this time, the image enhancement unit 404 generates a distorted image by enlarging or reducing the input image to a certain resolution set for the training data. The image enhancement unit 404 also pads the distorted image to a certain image size set for the training data to generate a padded image, and inputs the padded image to the image enhancement engine.

[0157] Furthermore, the image enhancement unit 404 trims the high-quality padded image output from the image enhancement engine by the amount of padding applied, thereby generating a high-quality deformed image. Subsequently, the image enhancement unit 404 reduces or enlarges the generated high-quality deformed image to match the image size of the input image, thereby generating a high-quality image.

[0158] Therefore, the image enhancement unit 404 according to this embodiment can enhance the image quality of input images with an image enhancement engine and generate high-quality images, even if the input image size could not be handled in the first embodiment.

[0159] Next, a series of image processing steps according to this embodiment will be described with reference to Figures 5 and 11. Figure 11 is a flowchart of the image quality enhancement process according to this embodiment. Note that the processes in steps S510, S520, and S550 according to this embodiment are the same as those in the first embodiment, so their explanation is omitted. Note that if the image quality enhancement is to be performed unconditionally on the input image with respect to shooting conditions other than image size, the process in step S530 may be omitted after the process in step S520, and the process may proceed to step S540.

[0160] In step S520, as in the first embodiment, once the shooting condition acquisition unit 402 has acquired the shooting conditions for the input image, the process moves to step S530. In step S530, the image quality enhancement feasibility determination unit 403 uses the acquired shooting conditions to determine whether the image quality enhancement engine in the image quality enhancement unit 404 can process the input image. Specifically, the image quality enhancement feasibility determination unit 403 determines whether the shooting area, shooting method, and shooting angle of view of the input image are suitable for processing by the image quality enhancement engine. Unlike the first embodiment, the image quality enhancement feasibility determination unit 403 does not determine the image size.

[0161] The image enhancement feasibility determination unit 403 determines the shooting area, shooting method, and shooting angle of view. If it determines that the input image is processable, the process proceeds to step S540. On the other hand, if the image enhancement feasibility determination unit 403 determines, based on these shooting conditions, that the image enhancement engine cannot process the input image, the process proceeds to step S550. Depending on the settings and implementation configuration of the image processing device 400, the image enhancement process in step S540 may be performed even if it is determined that the input image is processable based on some of the shooting area, shooting method, and shooting angle of view.

[0162] When the process moves to step S540, the image enhancement process according to this embodiment, as shown in Figure 11, is started. In the image enhancement process according to this embodiment, first, in step S1110, the image enhancement unit 404 enlarges or reduces the input image to a certain resolution set for the training data, and generates a distorted image.

[0163] Next, in step S1120, the image enhancement unit 404 pads the generated deformed image to create a padded image so that it matches the image size set for the training data. At this time, the image enhancement unit 404 fills the padded area with a fixed pixel value, fills it with neighboring pixel values, or performs mirror padding, in accordance with the characteristics of the machine learning model, so as to effectively enhance the image quality.

[0164] In step S1130, the image enhancement unit 404 inputs the padding image to the image enhancement engine and obtains a high-quality padding image. Next, in step S1140, the image enhancement unit 404 trims the high-quality padding image by the amount of padding performed in step S1120 to generate a high-quality deformed image.

[0165] Subsequently, in step S1150, the image enhancement unit 404 reduces or enlarges the high-quality deformed image to the image size of the input image to generate a high-quality image. Once the image enhancement unit 404 has generated a high-quality image in step S1130, the image enhancement process according to this embodiment is completed, and the process moves on to step S550. The process in step S550 is the same as in step S550 of the first embodiment, so its description is omitted.

[0166] As described above, the image enhancement unit 404 in this embodiment adjusts the image size of the input image so that the resolution of the input image becomes a predetermined resolution. The image enhancement unit 404 also generates a padded image by padding the adjusted image size of the input image so that the adjusted image size becomes an image size that can be handled by the image enhancement engine, and inputs the padded image to the image enhancement engine. Subsequently, the image enhancement unit 404 trims the output image from the image enhancement engine by the amount of padding applied. Then, the image enhancement unit 404 generates a high-resolution image by adjusting the image size of the trimmed image to the original image size of the input image.

[0167] As a result, the image enhancement unit 404 of this embodiment can enhance the image quality of input images with the image enhancement engine, even if the input image size could not be handled by the first embodiment, and generate high-quality images. Furthermore, by using an image enhancement engine trained on resolution-based training data, it may be possible to enhance the image quality of input images more efficiently than the image enhancement engine of the fourth embodiment, which simply processes images of the same size.

[0168] The output unit 405 may output the generated high-resolution images to the imaging device 10 or other devices connected to the image processing device 400, as in the first embodiment. Furthermore, the output data for the training data of the high-resolution engine is not limited to high-resolution images that have undergone overlay processing, as in the first embodiment. That is, high-resolution images obtained by performing at least one of the processing group or imaging method, such as overlay processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, imaging using a high-performance imaging device, high-cost processing, or noise reduction processing, may be used.

[0169] <Sixth Embodiment> Next, an image processing apparatus according to the sixth embodiment will be described with reference to Figures 4, 5, 12, and 13. In this embodiment, the image enhancement unit generates a high-resolution image by enhancing the image quality of the input image in areas of a certain image size.

[0170] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0171] The image enhancement unit 404 according to this embodiment is equipped with an image enhancement engine similar to that in the first embodiment. However, in this embodiment, the training data used to train the image enhancement engine is different from the training data in the first embodiment. Specifically, the training data consists of pairs of input data, which are low-resolution images, and output data, which are high-resolution images, and is composed of rectangular region images of a certain image size whose positional relationship corresponds to that of the low-resolution image and the high-resolution image. Note that the rectangular region is just one example of a sub-region and does not need to be rectangular; it can be any shape.

[0172] Now, with reference to Figure 12, the training data for the image enhancement engine according to this embodiment will be described. As shown in Figure 12, consider a case where one of the pairs constituting the training data includes, for example, a low-resolution original image Im1210 and a high-resolution superimposed image Im1220. In this case, in the first embodiment, the input data for the training data is Im1210 and the output data is Im1220.

[0173] In contrast, in this embodiment, the rectangular region image R1211 from the original image Im1210 is used as input data, and the rectangular region image R1221, which has the same shooting area as the rectangular region image R1211 in the superimposed image Im1220, is used as output data. Then, the input data rectangular region image R1211 and the output data rectangular region image R1221 constitute a training data pair (hereinafter referred to as the first rectangular region image pair). Here, the rectangular region images R1211 and R1221 are assumed to be images of a fixed image size. The original image Im1210 and the superimposed image Im1220 may be aligned by any method. Furthermore, the corresponding positional relationship between the rectangular region images R1211 and R1221 may be determined by any method such as template matching. Depending on the design of the high-image-quality engine, the respective image sizes and dimensionalities of the input data and output data may differ. For example, if the image to be processed is an OCT image, and the input data is a part of a B-scan image (two-dimensional image), the output data may be a part of an A-scan image (one-dimensional image).

[0174] A fixed image size for the rectangular region images R1211 and R1221 can be determined, for example, from the common divisor of the pixel count groups in each dimension of the set of image sizes of the image to be processed (input image). In this case, it is possible to prevent the positional relationships of the rectangular region images output by the image quality enhancement engine from overlapping. Specifically, consider the case where the image to be processed is a two-dimensional image, and the first image size in the set of image sizes is 500 pixels wide and 500 pixels high, and the second image size is 100 pixels wide and 100 pixels high. Here, a fixed image size for the rectangular region images R1211 and R1221 is selected from the common divisor of each side. In this case, for example, the fixed image size can be selected from 100 pixels wide and 100 pixels high, 50 pixels wide and 50 pixels high, 25 pixels wide and 25 pixels high, etc.

[0175] If the image to be processed is three-dimensional, the number of pixels is determined for width, height, and depth. Multiple rectangular regions can be set for one pair of low-resolution images corresponding to the input data and high-resolution images corresponding to the output data. For example, rectangular region image R1212 from the original image Im1210 is used as input data, and rectangular region image R1222, which has the same capture area as rectangular region image R1212 in the superimposed image Im1220, is used as output data. Then, a training data pair is formed using the input rectangular region image R1212 and the output rectangular region image R1222. This allows for the creation of a rectangular region image pair different from the first pair.

[0176] Furthermore, by creating numerous pairs of rectangular region images by transforming the rectangular region images into images with different coordinates, the set of pairs constituting the training data can be enriched, and efficient image enhancement can be expected using an image enhancement engine trained with these training pairs. However, pairs that do not contribute to image enhancement of the machine learning model can be excluded from the training data. For example, if the rectangular region image created from the high-resolution image that constitutes the output data of a pair is of unsuitable quality for diagnosis, the image enhancement engine trained with such training data may also output an image of unsuitable quality for image diagnosis. Therefore, pairs containing such high-resolution images can be removed from the training data.

[0177] Furthermore, for example, if a pair of rectangular region images—one created from a low-resolution image and the other from a high-resolution image—have significantly different average brightness or brightness distributions, such pairs can be removed from the training data. Training with such training data may result in the image enhancement engine outputting images with brightness distributions significantly different from the input images, making them unsuitable for image diagnosis.

[0178] Furthermore, consider a case where, for example, the structure and position of the subject depicted in a pair of rectangular region images—one created from a low-resolution image and the other from a high-resolution image—differ significantly. In this case, a high-resolution enhancement engine trained using such training data may output an image unsuitable for diagnostic imaging, where the subject is depicted with a structure and position that differs greatly from the input image. Therefore, such pairs can be removed from the training data.

[0179] Similar to the first embodiment, the input data for the training data uses images with specific shooting conditions that are assumed to be the target of processing. These specific shooting conditions are predetermined shooting area, shooting method, and shooting angle of view. In other words, unlike the first embodiment, the specific shooting conditions in this embodiment do not include image size.

[0180] The image enhancement unit 404 according to this embodiment enhances the image quality of the input image to generate a high-resolution image using an image enhancement engine trained with such training data. In this process, the image enhancement unit 404 divides the input image into a group of rectangular region images of a fixed size set for the training data, which are continuous and without gaps. The image enhancement unit 404 enhances the image quality of each of the divided rectangular region image groups using the image enhancement engine to generate a high-resolution group of rectangular region images. Subsequently, the image enhancement unit 404 arranges and combines the generated high-resolution rectangular region image groups according to the positional relationship of the input image to generate a high-resolution image. Here, during training, if the positional relationship between the input data and output data, which are paired images, corresponds, each rectangular region may be cut out (extracted) from any location in the low-resolution image and the high-resolution image. On the other hand, during image enhancement, the input image may be divided into a group of rectangular region images that are continuous and without gaps. Furthermore, the image size of each paired image during training and the image size of each rectangular region image during image enhancement may be set to correspond to each other (for example, be the same). These measures improve learning efficiency while preventing problems such as unnecessary calculations or missing information that prevent the image from being generated.

[0181] Thus, the image enhancement unit 404 of this embodiment enhances the image quality of the input image in rectangular region units and combines the enhanced images, thereby enabling the enhancement of image sizes that could not be handled in the first embodiment and generating high-quality images.

[0182] Next, a series of image processing steps according to this embodiment will be described with reference to Figures 5, 13, and 14. Figure 13 is a flowchart of the image quality enhancement process according to this embodiment. Note that the processes in steps S510, S520, and S550 according to this embodiment are the same as those in the first embodiment, so their explanation will be omitted. Note that if image quality enhancement is to be performed unconditionally on the input image with respect to shooting conditions other than image size, the process in step S530 may be omitted after the process in step S520, and the process may proceed to step S540.

[0183] In step S520, as in the first embodiment, once the shooting condition acquisition unit 402 has acquired the shooting conditions for the input image, the process moves to step S530. In step S530, the image quality enhancement feasibility determination unit 403 uses the acquired shooting conditions to determine whether the image quality enhancement engine in the image quality enhancement unit 404 can process the input image. Specifically, the image quality enhancement feasibility determination unit 403 determines whether the shooting area, shooting method, and shooting angle of view of the input image are suitable for processing by the image quality enhancement engine. Unlike the first embodiment, the image quality enhancement feasibility determination unit 403 does not determine the image size.

[0184] The image enhancement feasibility determination unit 403 determines the shooting area, shooting method, and shooting angle of view. If it determines that the input image is processable, the process proceeds to step S540. On the other hand, if the image enhancement feasibility determination unit 403 determines, based on these shooting conditions, that the image enhancement engine cannot process the input image, the process proceeds to step S550. Depending on the settings and implementation configuration of the image processing device 400, the image enhancement process in step S540 may be performed even if it is determined that the input image is processable based on some of the shooting area, shooting method, and shooting angle of view.

[0185] When the process moves to step S540, the image enhancement process according to this embodiment, as shown in Figure 13, is started. This will be explained using Figure 14. In the image enhancement process according to this embodiment, first, in step S1310, the input image is divided into a group of rectangular region images of a fixed image size (shown in R1411) set for the training data, without gaps, as shown in Figure 14(a). Here, Figure 14(a) shows an example in which the input image Im1410 is divided into a group of rectangular region images R1411 to R1426 of a fixed image size. As mentioned above, depending on the design of the image enhancement engine, the image size and number of dimensions of the input image and output image of the image enhancement engine may differ. In this case, the division positions of the input image can be adjusted by overlapping or separating them so that there are no defects in the combined high-resolution image generated in step S1320. Figure 14(b) shows an example of overlapping division positions. In Figure 14(b), R1411' and R1412' show overlapping regions. Although not shown in the diagram for simplicity, R1413 to R1426 also have similar overlapping regions R1413' to R1426'. Note that the rectangular region size set for the training data in the case of Figure 14(b) is the size shown in R1411'. Since there is no data outside the image (top, bottom, left, and right edges) of the input image Im1410, it is filled with a certain pixel value, filled with the value of a neighboring pixel, or mirror padded. Also, depending on the image enhancement engine, the accuracy of image enhancement may decrease in the peripheral areas (top, bottom, left, and right edges) inside the image due to filtering. For this reason, as shown in Figure 14(b), rectangular region images may be set with overlapping division positions, and a portion of the rectangular region images may be cropped and combined to form the final image. The size of the rectangular region should be set according to the characteristics of the image enhancement engine. While Figures 14(a) and (b) show examples of OCT tomographic images, as shown in Figures 14(c) and (d), the input image (Im1450) can also be a frontal image such as an OCTA En-Face image, and similar processing is possible. The size of the rectangular region image should be set appropriately depending on the target image and the type of image enhancement engine.

[0186] Next, in step S1320, the image enhancement unit 404 enhances the image quality of each of the rectangular region images R1411 to R1426, or, if overlapping regions are set, the rectangular region images R1411' to R1426', using the image enhancement engine to generate a group of high-quality rectangular region images.

[0187] Then, in step S1330, the image enhancement unit 404 combines the generated high-quality rectangular region images by arranging each of them in the same positional relationship as the rectangular region images R1411 to R1426 groups that were divided for the input image, thereby generating a high-quality image. If overlapping regions are set, the rectangular region images R1411 to R1426 are cut out and combined after being arranged in the same positional relationship as the rectangular region images R1411' to R1426', thereby generating a high-quality image. Alternatively, the luminance values ​​of the rectangular region images R1411' to R1426' may be corrected using the overlapping regions. For example, a reference rectangular region image can be arbitrarily set. Then, by measuring the luminance values ​​at the same coordinate points in adjacent rectangular images that have overlapping regions with the reference rectangular image, the difference (ratio) in luminance values ​​between adjacent images can be determined. Similarly, by determining the difference (ratio) in luminance values ​​in the overlapping regions for all images, it becomes possible to correct the overall luminance values ​​to eliminate unevenness. Furthermore, it is not necessary to use the entire overlapping area for brightness value correction; some of the overlapping area (a few pixels in the peripheral area) may be omitted.

[0188] As described above, the image enhancement unit 404 according to this embodiment divides the input image into a plurality of rectangular region images (third images) R1411 to R1426 of a predetermined image size. Then, the image enhancement unit 404 inputs the divided plurality of rectangular region images R1411 to R1426 to the image enhancement engine to generate a plurality of fourth images, and generates a high-resolution image by integrating the plurality of fourth images. If the positional relationship between the rectangular region groups overlaps during integration, the pixel value groups of the rectangular region groups can be integrated or overwritten.

[0189] As a result, the image quality enhancement unit 404 of this embodiment can enhance the image quality of input images with an image quality enhancement engine, even if the input image size could not be handled in the first embodiment, and generate a high-quality image. Furthermore, if the training data is created from multiple images obtained by dividing low-quality and high-quality images into predetermined image sizes, a large amount of training data can be created from a small number of images. Therefore, in this case, the number of low-quality and high-quality images required to create the training data can be reduced.

[0190] The output unit 405 may output the generated high-resolution images to the imaging device 10 or other devices connected to the image processing device 400, as in the first embodiment. Furthermore, the output data for the training data of the high-resolution engine is not limited to high-resolution images that have undergone overlay processing, as in the first embodiment. That is, high-resolution images obtained by performing at least one of the processing group or imaging method, such as overlay processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, imaging using a high-performance imaging device, high-cost processing, or noise reduction processing, may be used.

[0191] <Seventh Embodiment> Next, an image processing apparatus according to the seventh embodiment will be described with reference to Figures 15-17. In this embodiment, the image quality evaluation unit selects the highest quality image from among multiple high-quality images output from multiple high-quality engines, in accordance with the examiner's instructions.

[0192] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment.

[0193] Figure 15 shows a schematic configuration of the image processing apparatus 1500 according to this embodiment. The image processing apparatus 1500 according to this embodiment is provided with an acquisition unit 401, a shooting condition acquisition unit 402, a high-quality image enhancement feasibility determination unit 403, a high-quality image enhancement unit 404, and an output unit 405, in addition to a quality evaluation unit 1506. Note that the image processing apparatus 1500 may be composed of multiple devices, each equipped with some of these components. Here, the acquisition unit 401, the shooting condition acquisition unit 402, the high-quality image enhancement feasibility determination unit 403, the high-quality image enhancement unit 404, and the output unit 405 are the same as those in the image processing apparatus according to the first embodiment, so the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0194] Furthermore, the image processing device 1500 may be connected to the imaging device 10, the display unit 20, and other devices (not shown) via any circuits or networks, similar to the image processing device 400 according to the first embodiment. These devices may also be connected to other devices via circuits or networks, or they may be configured integrally with other devices. In this embodiment, these devices are shown as separate devices, but some or all of these devices may be configured integrally.

[0195] The image enhancement unit 404 according to this embodiment is equipped with two or more image enhancement engines that have been machine-learned using different training data. Here, the method for creating the training data set according to this embodiment will be described. Specifically, first, a group of pairs of input data, which are low-resolution images, and output data, which are high-resolution images, are prepared, each taken under various shooting conditions. Next, a training data set is created by grouping the pair groups according to an arbitrary combination of shooting conditions. For example, a first training data set is created consisting of a pair group acquired by a first combination of shooting conditions, and a second training data set is created consisting of a pair group acquired by a second combination of shooting conditions.

[0196] Subsequently, separate image enhancement engines are trained using each set of training data. For example, a group of image enhancement engines is prepared, such as a first image enhancement engine corresponding to a machine learning model trained with the first set of training data, and another first image enhancement engine corresponding to a machine learning model trained with the first set of training data.

[0197] Because each of these image enhancement engines uses different training data for its corresponding machine learning model, the degree to which it can enhance the image quality of an input image varies depending on the shooting conditions of the image input to the engine. Specifically, the first image enhancement engine enhances the image quality to a high degree for input images captured using the first combination of shooting conditions, and to a low degree for images captured using the second combination of shooting conditions. Similarly, the second image enhancement engine enhances the image quality to a high degree for input images captured using the second combination of shooting conditions, and to a low degree for images captured using the first combination of shooting conditions.

[0198] Each training data set is composed of pairs grouped by combinations of shooting conditions, resulting in similar image quality tendencies among the image groups constituting each pair. Therefore, the image quality enhancement engine can enhance image quality more effectively than the image quality enhancement engine according to the first embodiment, provided that the combinations of shooting conditions are compatible. The combinations of shooting conditions for grouping the training data pairs are arbitrary and may include, for example, two or more combinations of shooting area, shooting angle of view, and image resolution. Alternatively, the training data may be grouped based on a single shooting condition, similar to the second embodiment.

[0199] The image quality evaluation unit 1506 selects the image with the highest image quality from among the multiple high-resolution images generated by the image quality enhancement unit 404 using multiple image quality enhancement engines, in accordance with the examiner's instructions.

[0200] The output unit 405 can display the high-resolution image selected by the image quality evaluation unit 1506 on the display unit 20 or output it to another device. The output unit 405 can display multiple high-resolution images generated by the image quality enhancement unit 404 on the display unit 20, and the image quality evaluation unit 1506 can select the image with the highest quality according to instructions from the examiner who has viewed the display unit 20.

[0201] As a result, the image processing device 1500 can output the highest quality image among multiple high-quality images generated using multiple high-quality image enhancement engines, according to the examiner's instructions.

[0202] The series of image processing according to this embodiment will now be described with reference to Figures 16 and 17. Figure 16 is a flowchart of the series of image processing according to this embodiment. Note that the processing in steps S1610 and S1620 according to this embodiment is the same as the processing in steps S510 and S520 in the first embodiment, so the explanation will be omitted. Note that if the image quality is to be increased unconditionally for the input image regardless of the shooting conditions, the processing in step S1630 may be omitted after the processing in step S1620, and the processing may proceed to step S1640.

[0203] In step S1620, as in the first embodiment, once the shooting condition acquisition unit 402 has acquired a set of shooting conditions for the input image, the process moves to step S1630. In step S1630, the image quality enhancement feasibility determination unit 403, as in the second embodiment, uses the acquired set of shooting conditions to determine whether any of the image quality enhancement engines provided in the image quality enhancement unit 404 can process the input image.

[0204] If the image quality enhancement feasibility determination unit 403 determines that none of the image quality enhancement engines can process the input image, the process proceeds to step S1660. On the other hand, if the image quality enhancement feasibility determination unit 403 determines that any of the image quality enhancement engines can process the input image, the process proceeds to step S1640. Depending on the settings and implementation configuration of the image processing device 400, step S1640 may also be performed, as in the first embodiment, even if the image quality enhancement engine determines that some shooting conditions cannot be processed.

[0205] In step S1640, the image enhancement unit 404 inputs the input image acquired in step S1610 to each of the image enhancement engines and generates a group of high-resolution images.

[0206] In step S1650, the image quality evaluation unit 1506 selects the highest quality image from the group of high-quality images generated in step S1640. Specifically, first, the output unit 405 displays the group of high-quality images generated in step S1640 on the user interface of the display unit 20.

[0207] Here, Figure 17 shows an example of the interface. The interface displays the input image Im1710 and the high-resolution images Im1720, Im1730, Im1740, and Im1750 output by each of the image enhancement engines. The examiner operates an arbitrary input device (not shown) to indicate the image with the highest image quality, i.e., the image most suitable for diagnostic imaging, from the image group (high-resolution images Im1720 to Im1750). Note that input images that have not been enhanced by the image enhancement engine may also be more suitable for diagnostic imaging, so input images may be added to the image group to be indicated by the examiner.

[0208] Subsequently, the image quality evaluation unit 1506 selects the high-quality image indicated by the examiner as the highest quality image.

[0209] In step S1660, the output unit 405 displays the image selected in step S1650 on the display unit 20 or outputs it to another device. However, if it is determined in step S1630 that the input image is unprocessable, the output unit 405 outputs the input image as the output image. The output unit 405 may also indicate on the display unit 20 that the output image is the same as the input image if the examiner specifies an input image or if the input image is unprocessable.

[0210] As described above, the image enhancement unit 404 according to this embodiment generates multiple high-resolution images from an input image using multiple image enhancement engines, and the output unit 405 of the image processing device 1500 outputs at least one of the multiple high-resolution images according to the examiner's instructions. In particular, in this embodiment, the output unit 405 outputs the highest quality image according to the examiner's instructions. As a result, the image processing device 1500 can output a high-resolution image with the highest image quality according to the examiner's instructions from among multiple high-resolution images generated using multiple high-resolution engines.

[0211] The output unit 405 may output the generated high-resolution images to the imaging device 10 or other devices connected to the image processing device 1500, as in the first embodiment. Furthermore, the output data for the training data of the high-resolution engine is not limited to high-resolution images that have undergone overlay processing, as in the first embodiment. That is, high-resolution images obtained by performing at least one of the processing group or imaging method, such as overlay processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, imaging using a high-performance imaging device, high-cost processing, or noise reduction processing, may also be used.

[0212] <Eighth Embodiment> Next, an image processing apparatus according to the eighth embodiment will be described with reference to Figures 15 and 16. In this embodiment, the image quality evaluation unit uses an image quality evaluation engine to select the highest quality image from among multiple high-quality images output from multiple high-quality engines.

[0213] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 1500 according to the seventh embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the seventh embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the seventh embodiment, the same reference numerals are used to indicate the configuration shown in Figure 15, and their explanation is omitted.

[0214] The image quality evaluation unit 1506 according to this embodiment is equipped with an image quality evaluation engine that evaluates the image quality of an input image. The image quality evaluation engine outputs an image quality evaluation index for the input image. The image quality evaluation processing method that calculates the image quality evaluation index in the image quality evaluation engine according to this embodiment uses a machine learning model constructed using a machine learning algorithm. The input data pairs that constitute the training data for the machine learning model are image sets consisting of a group of low-resolution images and a group of high-resolution images taken in advance under various shooting conditions. The output data pairs that constitute the training data for the machine learning model are, for example, a group of image quality evaluation indices set by an examiner performing a medical image diagnosis for each of the images in the input data.

[0215] Next, with reference to Figure 16, a series of image processing steps according to this embodiment will be described. Note that the processing in steps S1610, S1620, S1630, and S1660 according to this embodiment is the same as the processing in the seventh embodiment, so the description will be omitted. Note that if the image quality is to be increased unconditionally for the input image regardless of the shooting conditions, the processing in step S1630 may be omitted after the processing in step S1620, and the processing may proceed to step S1640.

[0216] In step S1630, similar to the seventh embodiment, if the image quality enhancement feasibility determination unit 403 determines that any of the image quality enhancement engines can handle the input image, the process proceeds to step S1640. Depending on the settings and implementation of the image processing device 400, step S1640 may also be performed, similar to the first embodiment, even if the image quality enhancement engine determines that some shooting conditions cannot be handled.

[0217] In step S1640, the image enhancement unit 404 inputs the input image acquired in step S1610 to each of the image enhancement engines and generates a group of high-resolution images.

[0218] In step S1650, the image quality evaluation unit 1506 selects the highest quality image from the high-quality image group generated in step S1640. Specifically, first, the image quality evaluation unit 1506 inputs the high-quality image group generated in step S1640 into the image quality evaluation engine. The image quality evaluation engine calculates an image quality evaluation index for each input high-quality image based on learning. The image quality evaluation unit 1506 selects the high-quality image with the highest calculated image quality evaluation index. Note that input images that have not been enhanced by the image quality enhancement engine may be more suitable for image diagnosis, so the image quality evaluation unit 1506 may also input the input images into the image quality evaluation engine and include the image quality evaluation index for the input images in the selection. Step S1660 is the same as step S1660 in the seventh embodiment, so the explanation is omitted.

[0219] As described above, the image processing device 1500 according to this embodiment further includes an image quality evaluation unit 1506 for evaluating the image quality of high-resolution images. The image quality enhancement unit 404 generates multiple high-resolution images from an input image using multiple image quality enhancement engines, and the output unit 405 of the image processing device 1500 outputs at least one image from the multiple high-resolution images according to the evaluation result by the image quality evaluation unit 1506. In particular, the image quality evaluation unit 1506 according to this embodiment includes an image quality evaluation engine that uses evaluation values ​​from a predetermined evaluation method as training data. The image quality evaluation unit 1506 selects the high-resolution image from the multiple high-resolution images that has the highest evaluation result using the image quality evaluation engine of the image quality evaluation unit 1506. The output unit 405 outputs the high-resolution image with the highest evaluation value selected by the image quality evaluation unit 1506.

[0220] As a result, the image processing device 1500 according to this embodiment can easily output the highest quality image most suitable for image diagnosis from among multiple high-quality images based on the output of the image quality evaluation engine.

[0221] In this embodiment, the image quality evaluation unit 1506 selects the high-resolution image with the highest image quality evaluation index among those output by the image quality evaluation engine, and the output unit 405 displays the selected high-resolution image on the display unit 20. However, the configuration of the image quality evaluation unit 1506 is not limited to this. For example, the image quality evaluation unit 1506 may select high-resolution images with the top several image quality evaluation indices among those output by the image quality evaluation engine, and the output unit 405 may display the selected high-resolution image on the display unit 20. Alternatively, the output unit 405 may display the image quality evaluation index output by the image quality evaluation engine along with the corresponding high-resolution image on the display unit 20, and the image quality evaluation unit 1506 may select the highest-resolution image according to the examiner's instructions.

[0222] The output unit 405 may output the generated high-resolution images to the imaging device 10 or other devices connected to the image processing device 1500, as in the first embodiment. Furthermore, the output data for the training data of the high-resolution engine is not limited to high-resolution images that have undergone overlay processing, as in the first embodiment. That is, high-resolution images obtained by performing at least one of the processing group or imaging method, such as overlay processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, imaging using a high-performance imaging device, high-cost processing, or noise reduction processing, may also be used.

[0223] <Ninth Embodiment> Next, an image processing apparatus according to the ninth embodiment will be described with reference to Figures 18 and 19. In this embodiment, the authenticity evaluation unit uses the authenticity evaluation engine to evaluate whether the high-resolution image generated by the high-resolution enhancement unit 404 is sufficiently high-resolution.

[0224] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment.

[0225] Figure 18 shows a schematic configuration of the image processing apparatus 1800 according to this embodiment. The image processing apparatus 1800 according to this embodiment is provided with an acquisition unit 401, a shooting condition acquisition unit 402, a high-quality image enhancement feasibility determination unit 403, a high-quality image enhancement unit 404, and an output unit 405, as well as an authenticity evaluation unit 1807. Note that the image processing apparatus 1800 may be composed of multiple devices, each equipped with some of these components. Here, the acquisition unit 401, the shooting condition acquisition unit 402, the high-quality image enhancement feasibility determination unit 403, the high-quality image enhancement unit 404, and the output unit 405 are the same as those in the image processing apparatus according to the first embodiment, so the same reference numerals are used for the components shown in Figure 4, and their explanation is omitted.

[0226] Furthermore, the image processing device 1800 may be connected to the imaging device 10, the display unit 20, and other devices (not shown) via any circuits or networks, similar to the image processing device 400 according to the first embodiment. These devices may also be connected to other devices via circuits or networks, or they may be configured integrally with other devices. In this embodiment, these devices are treated as separate devices, but some or all of these devices may be configured integrally.

[0227] The authenticity evaluation unit 1807 is equipped with an authenticity evaluation engine. The authenticity evaluation unit 1807 uses the authenticity evaluation engine to evaluate whether the high-resolution image generated by the high-resolution enhancement engine is sufficiently high-resolution. The authenticity evaluation processing method in the authenticity evaluation engine according to this embodiment uses a machine learning model constructed using a machine learning algorithm.

[0228] The training data for training the machine learning model includes pairs of high-resolution images taken under various shooting conditions beforehand and labels indicating that the images were taken and acquired by the target shooting device (hereinafter referred to as "genuine labels"). The training data also includes pairs of high-resolution images generated by inputting low-resolution images into a high-resolution engine with poor image enhancement accuracy and labels indicating that the images were not taken and acquired by the target shooting device (hereinafter referred to as "fake labels").

[0229] A genuineness evaluation engine trained using such training data cannot definitively determine whether an input image was captured and acquired by a camera, but it can evaluate whether an image has the characteristics of an image captured and acquired by a camera. Utilizing this characteristic, the genuineness evaluation unit 1807 can input a high-resolution image generated by the high-resolution enhancement unit 404 to the genuineness evaluation engine, thereby evaluating whether the high-resolution image generated by the high-resolution enhancement unit 404 is sufficiently high-resolution.

[0230] If the authenticity evaluation unit 1807 determines that the high-resolution image generated by the image enhancement unit 404 is sufficiently enhanced, the output unit 405 displays the high-resolution image on the display unit 20. On the other hand, if the authenticity evaluation unit 1807 determines that the high-resolution image generated by the image enhancement unit 404 is not sufficiently enhanced, the output unit 405 displays the input image on the display unit 20. When displaying the input image, the output unit 405 can also display on the display unit 20 that the high-resolution image generated by the image enhancement unit 404 was not sufficiently enhanced and that the displayed image is the input image.

[0231] The series of image processing steps according to this embodiment will now be described with reference to Figure 19. Figure 19 is a flowchart of the series of image processing steps according to this embodiment. Note that the processing steps S1910 to S1940 according to this embodiment are the same as the processing steps S510 to S540 in the first embodiment, so the explanation will be omitted. Note that if the image quality is to be increased unconditionally for the input image regardless of the shooting conditions, the processing in step S1930 may be omitted after the processing in step S1920, and the processing may proceed to step S1940.

[0232] In step S1940, once the image enhancement unit 404 has generated a group of high-resolution images, the process moves to step S1950. In step S1950, the authenticity evaluation unit 1807 inputs the high-resolution images generated in step S1940 into the authenticity evaluation engine and performs an authenticity evaluation based on the output of the authenticity evaluation engine. Specifically, if the authenticity evaluation engine outputs a genuine label (true), the authenticity evaluation unit 1807 evaluates that the generated high-resolution images are sufficiently high-resolution. On the other hand, if the authenticity evaluation engine outputs a counterfeit label (fake), the authenticity evaluation unit 1807 evaluates that the generated high-resolution images are not sufficiently high-resolution.

[0233] In step S1960, if the authenticity evaluation unit 1807 determines that the high-resolution image generated by the image enhancement unit 404 is sufficiently high-resolution, the output unit 405 displays the high-resolution image on the display unit 20. On the other hand, if the authenticity evaluation unit 1807 determines that the high-resolution image generated by the image enhancement unit 404 is not sufficiently high-resolution, the output unit 405 displays the input image on the display unit 20.

[0234] As described above, the image processing apparatus 1800 according to this embodiment further comprises an authenticity evaluation unit 1807 for evaluating the quality of high-resolution images, and the authenticity evaluation unit 1807 includes an authenticity evaluation engine for evaluating the authenticity of images. The authenticity evaluation engine includes a machine learning engine that uses images generated by an image enhancement engine with lower (worse) accuracy in image enhancement processing than the image enhancement engine of the image enhancement unit 404 as training data. The output unit 405 of the image processing apparatus 1800 outputs a high-resolution image when the output from the authenticity evaluation engine of the authenticity evaluation unit is true.

[0235] As a result, with the image processing apparatus 1800 according to this embodiment, the examiner can efficiently check a high-resolution image that has been sufficiently enhanced.

[0236] Furthermore, the efficiency and accuracy of both engines may be improved by training the machine learning models of the image enhancement engine and the authenticity evaluation engine in coordination.

[0237] In this embodiment, the image enhancement unit 404 generates one high-resolution image, and the authenticity evaluation unit 1807 evaluates the generated high-resolution image. However, the evaluation by the authenticity evaluation unit 1807 is not limited to this. For example, as in the second embodiment, if the image enhancement unit 404 generates multiple high-resolution images using multiple image enhancement engines, the authenticity evaluation unit 1807 may be configured to evaluate at least one of the generated high-resolution images. In this case, for example, the authenticity evaluation unit 1807 may evaluate all of the generated high-resolution images, or it may evaluate only the image indicated by the examiner among the multiple high-resolution images.

[0238] Furthermore, the output unit 405 may display on the display unit 20 the result of the judgment made by the authenticity evaluation unit 1807 as to whether or not the high-resolution image has been sufficiently enhanced, and may output the high-resolution image according to the inspector's instructions.

[0239] The output unit 405 may output the generated high-resolution images to the imaging device 10 or other devices connected to the image processing device 1800, as in the first embodiment. Furthermore, the output data for the training data of the high-resolution engine is not limited to high-resolution images that have undergone overlay processing, as in the first embodiment. That is, high-resolution images obtained by performing at least one of the processing group or imaging method, such as overlay processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, imaging using a high-performance imaging device, high-cost processing, or noise reduction processing, may be used.

[0240] <Tenth Embodiment> Next, an image processing apparatus according to the tenth embodiment will be described with reference to Figures 4 and 5. In this embodiment, the image enhancement unit divides a three-dimensional input image into a plurality of two-dimensional images and inputs them to the image enhancement engine, and generates a three-dimensional high-resolution image by combining the output images from the image enhancement engine.

[0241] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0242] The acquisition unit 401 according to this embodiment acquires a three-dimensional image composed of structurally continuous two-dimensional image groups. Specifically, the three-dimensional image is, for example, a three-dimensional OCT volume image composed of a group of OCT B-scan images (tomographic images). Alternatively, it may be, for example, a three-dimensional CT volume image composed of a group of axial tomographic images.

[0243] The image enhancement unit 404 is equipped with an image enhancement engine, similar to the first embodiment. The input data and output data pairs, which serve as the training data for the image enhancement engine, consist of a group of two-dimensional images. The image enhancement unit 404 divides the acquired three-dimensional image into multiple two-dimensional images and inputs each two-dimensional image into the image enhancement engine. As a result, the image enhancement unit 404 can generate multiple two-dimensional high-resolution images.

[0244] The output unit 405 combines multiple two-dimensional high-resolution images generated for each two-dimensional image of the three-dimensional image by the image enhancement unit 404, and outputs a three-dimensional high-resolution image.

[0245] Next, with reference to Figure 5, a series of image processing steps according to this embodiment will be described. Note that the processing in steps S510 to S530 and step S550 according to this embodiment is the same as the processing in the first embodiment, so the description will be omitted. However, in step S510, the acquisition unit 401 acquires a three-dimensional image. Note that if the image quality is to be improved unconditionally for the input image regardless of the shooting conditions, the processing in step S530 may be omitted after the processing in step S520, and the processing may proceed to step S540.

[0246] In step S530, if the image quality improvement availability determination unit 403 determines that the input image can be processed by the image quality improvement engine, the process proceeds to step S540. The image quality improvement availability determination unit 403 may perform the determination based on the imaging conditions of the three-dimensional image, or may perform the determination based on the imaging conditions for a plurality of two-dimensional images constituting the three-dimensional image. In step S540, the image quality improvement unit 404 divides the acquired three-dimensional image into a plurality of two-dimensional images. The image quality improvement unit 404 inputs each of the plurality of divided two-dimensional images to the image quality improvement engine, and generates a plurality of two-dimensional high-quality images. The image quality improvement unit 404 combines the generated plurality of two-dimensional high-quality images based on the acquired three-dimensional image to generate a three-dimensional high-quality image.

[0247] In step S550, the output unit 405 causes the display unit 20 to display the generated three-dimensional high-quality image. The display mode of the three-dimensional high-quality image may be arbitrary.

[0248] As described above, the image quality improvement unit 404 according to the present embodiment divides a three-dimensional input image into a plurality of two-dimensional images and inputs the divided images to the image quality improvement engine. The image quality improvement unit 404 combines the plurality of two-dimensional high-quality images output from the image quality improvement engine to generate a three-dimensional high-quality image.

[0249] Accordingly, the image quality improvement unit 404 according to the present embodiment can improve the image quality of a three-dimensional image using the image quality improvement engine trained with two-dimensional image teacher data.

[0250] Note that, as in the first embodiment, the output unit 405 may output the generated high-quality image to the imaging device 10 or another device connected to the image processing apparatus 400. In addition, as in the first embodiment, the output data of the teacher data for the image quality improvement engine is not limited to high-quality images subjected to superposition processing. That is, a high-quality image obtained by performing at least one of processing groups and imaging methods including superposition processing, MAP estimation processing, smoothing filter processing, gradation conversion processing, imaging using a high-performance imaging device, high-cost processing, and noise reduction processing may be used.

[0251] <Embodiment 11> Next, an image processing apparatus according to the eleventh embodiment will be described with reference to Figures 4 and 5. In this embodiment, the image enhancement unit divides a three-dimensional input image into a plurality of two-dimensional images, enhances the image quality of the plurality of two-dimensional images in parallel using a plurality of image enhancement engines, and generates a three-dimensional high-resolution image by combining the output images from the image enhancement engines.

[0252] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the tenth embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the tenth embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first and tenth embodiments, the configuration shown in Figure 4 will be indicated using the same reference numerals and its description will be omitted.

[0253] The image enhancement unit 404 according to this embodiment is equipped with multiple image enhancement engines similar to those in the tenth embodiment. The group of multiple image enhancement engines provided in the image enhancement unit 404 may be implemented so as to be distributed processing across two or more device groups via circuits or networks, or they may be implemented in a single device.

[0254] Similar to the tenth embodiment, the image enhancement unit 404 divides the acquired three-dimensional image into multiple two-dimensional images. The image enhancement unit 404 then uses multiple image enhancement engines to enhance the image quality of the multiple two-dimensional images in a divided (parallel) manner, generating multiple high-resolution two-dimensional images. The image enhancement unit 404 combines multiple two-dimensional high-resolution images output from multiple image enhancement engines based on the three-dimensional image to be processed, thereby generating a three-dimensional high-resolution image.

[0255] Next, with reference to Figure 5, a series of image processing steps according to this embodiment will be described. Note that the processing in steps S510 to S530 and step S550 according to this embodiment is the same as the processing in the 10th embodiment, so the description will be omitted. Note that if the image quality is to be increased unconditionally for the input image regardless of the shooting conditions, the processing in step S530 may be omitted after the processing in step S520, and the processing may proceed to step S540.

[0256] In step S530, if the image quality enhancement feasibility determination unit 403 determines that the input image can be processed by the image quality enhancement engine, the process proceeds to step S540. The image quality enhancement feasibility determination unit 403 may make this determination based on the shooting conditions of the three-dimensional image, or it may make this determination based on the shooting conditions of multiple two-dimensional images that constitute the three-dimensional image.

[0257] In step S540, the image enhancement unit 404 divides the acquired three-dimensional image into multiple two-dimensional images. The image enhancement unit 404 inputs each of the divided two-dimensional images into multiple image enhancement engines and processes them in parallel to generate multiple two-dimensional high-resolution images. Based on the acquired three-dimensional image, the image enhancement unit 404 combines the generated multiple two-dimensional high-resolution images to generate a three-dimensional high-resolution image.

[0258] In step S550, the output unit 405 displays the generated three-dimensional high-resolution image on the display unit 20. The display mode of the three-dimensional high-resolution image may be arbitrary.

[0259] As described above, the image enhancement unit 404 according to this embodiment includes multiple image enhancement engines. The image enhancement unit 404 divides a three-dimensional input image into multiple two-dimensional images and generates multiple two-dimensional high-resolution images using multiple image enhancement engines in parallel. The image enhancement unit 404 generates a three-dimensional high-resolution image by integrating the multiple two-dimensional high-resolution images.

[0260] As a result, the image enhancement unit 404 according to this embodiment can enhance the image quality of three-dimensional images using an image enhancement engine that has been trained using training data of two-dimensional images. Furthermore, it is possible to enhance the image quality of three-dimensional images more efficiently compared to the tenth embodiment.

[0261] Furthermore, the training data for multiple image enhancement engines may differ depending on the processing target that each engine is working on. For example, the first image enhancement engine may be trained using training data for a first shooting area, and the second image enhancement engine may be trained using training data for a second shooting area. In this case, each image enhancement engine can perform two-dimensional image enhancement with greater accuracy.

[0262] Furthermore, the output unit 405 may output the generated high-resolution images to the imaging device 10 or other devices connected to the image processing device 400, similar to the first embodiment. Also, the output data for the training data of the high-resolution engine is not limited to high-resolution images that have undergone overlay processing, similar to the first embodiment. That is, high-resolution images obtained by performing at least one of the processing group or imaging method, such as overlay processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, imaging using a high-performance imaging device, high-cost processing, or noise reduction processing, may be used.

[0263] <Twelfth Embodiment> Next, an image processing apparatus according to the twelfth embodiment will be described with reference to Figures 5 and 20. In this embodiment, the acquisition unit 401 acquires input images from the image management system 2000 rather than from the imaging device.

[0264] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus 400 according to the first embodiment, the same reference numerals are used for the configuration shown in Figure 4 and their explanation is omitted.

[0265] Figure 20 shows a schematic configuration of the image processing device 400 according to this embodiment. The image processing device 400 according to this embodiment is connected to the image management system 2000 and the display unit 20 via arbitrary circuits and networks. The image management system 2000 is a device and system that receives and stores images captured by any imaging device or images that have been processed. Furthermore, the image management system 2000 can transmit images in response to requests from connected devices, perform image processing on stored images, and request image processing from other devices. The image management system can include, for example, a Picture Archiving and Communication System (PACS).

[0266] The acquisition unit 401 according to this embodiment can acquire an input image from the image management system 2000 connected to the image processing device 400. The output unit 405 can output the high-resolution image generated by the high-resolution enhancement unit 404 to the image management system 2000.

[0267] Next, with reference to Figure 5, a series of image processing steps according to this embodiment will be described. Note that the processing in steps S520 to S540 according to this embodiment is the same as the processing in the first embodiment, so the description will be omitted. Note that if the image quality is to be increased unconditionally for the input image regardless of the shooting conditions, the processing in step S530 may be omitted after the processing in step S520, and the processing may proceed to step S540.

[0268] In step S510, the acquisition unit 401 acquires, as an input image, an image stored in an image management system 2000 from the image management system 2000 connected via a circuit or a network. Note that the acquisition unit 401 may acquire the input image in response to a request from the image management system 2000. Such a request may be issued, for example, when the image management system 2000 stores an image, before transmitting the stored image to another device, or when displaying the stored image on the display unit 20. Furthermore, the request may be issued, for example, when a user operates the image management system 2000 to request high-image-quality processing, or when a high-quality image is used for an image analysis function included in the image management system 2000.

[0269] The processing from step S520 to step S540 is the same as the processing in the first embodiment. After the image quality enhancement unit 404 generates a high-quality image in step S540, the processing proceeds to step S550. In step S550, if a high-quality image has been generated in step S540, the output unit 405 outputs the high-quality image as an output image to the image management system 2000. If no high-quality image has been generated in step S540, the output unit 405 outputs the aforementioned input image as an output image to the image management system 2000. Note that, depending on the settings and implementation of the image processing apparatus 400, the output unit 405 may process the output image so that the image management system 2000 can use it, or convert the data format of the output image.

[0270] As described above, the acquisition unit 401 according to this embodiment acquires input images from the image management system 2000. Therefore, the image processing device 400 of this embodiment can output high-resolution images suitable for image diagnosis based on images stored in the image management system 2000, without increasing the invasiveness or effort required for the photographer or the patient. Furthermore, the output high-resolution images can be stored in the image management system 2000 or displayed on the user interface provided by the image management system 2000. In addition, the output high-resolution images can be used in the image analysis function provided by the image management system 2000 or transmitted via the image management system 2000 to other devices connected to the image management system 2000.

[0271] The image processing device 400, the image management system 2000, and the display unit 20 may be connected to other devices (not shown) via circuits or networks. Furthermore, although these devices are separate in this embodiment, some or all of these devices may be integrated into a single unit.

[0272] Furthermore, the output unit 405 may output the generated high-resolution image to the image management system 2000 or other devices connected to the image processing device 400, similar to the first embodiment.

[0273] <13th Embodiment> Next, an image processing apparatus according to the thirteenth embodiment will be described with reference to Figures 4, 5, 21A, and 21B. In this embodiment, the image enhancement unit takes multiple images as input images and generates a single high-resolution image.

[0274] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0275] The acquisition unit 401 according to this embodiment acquires multiple images from the imaging device 10 or other devices as input data to be processed.

[0276] The image enhancement unit 404 according to this embodiment is equipped with an image enhancement engine similar to that in the first embodiment. The training data may also be the same as in the first embodiment. The image enhancement unit 404 inputs each of the multiple images acquired by the acquisition unit 401 into the image enhancement engine, and processes the output multiple high-resolution images by superimposing them to generate a final high-resolution image. The image enhancement unit 404 may align the multiple high-resolution images using any method before overlaying them.

[0277] The output unit 405 displays the final high-resolution image generated by the image enhancement unit 404 on the display unit 20. The output unit 405 may also display multiple input images on the display unit 20 along with the final high-resolution image. Furthermore, the output unit 405 may display multiple generated high-resolution images on the display unit 20 along with the final high-resolution image and input images.

[0278] Next, a series of image processing steps according to this embodiment will be described with reference to Figures 5 and 21A. Figure 21A is a flowchart of the image enhancement process according to this embodiment. Note that the processing steps S510 to S530 according to this embodiment are the same as those in the first embodiment, so their explanation will be omitted.

[0279] However, in step S510, the acquisition unit 401 acquires multiple images, and in steps S520 and S530, the shooting conditions are acquired for each of the multiple images, and it is determined whether or not they can be processed by the image quality enhancement engine. Note that if the image quality is to be enhanced unconditionally for the input image based on the shooting conditions, the processing in step S530 may be omitted after the processing in step S520, and the processing may proceed to step S540. Also, if it is determined that some of the multiple images cannot be processed by the image quality enhancement engine, those images can be excluded from subsequent processing.

[0280] In step S530, if the image quality enhancement feasibility determination unit 403 determines that the image quality enhancement engine can handle the multiple input images, the process proceeds to step S540. When the process proceeds to step S540, the image quality enhancement process according to this embodiment, as shown in Figure 21A, begins. In the image quality enhancement process according to this embodiment, first, in step S2110, the image quality enhancement unit 404 inputs each of the multiple input images to the image quality enhancement engine and generates a group of high-resolution images.

[0281] Next, in step S2120, the image enhancement unit 404 overlays the generated high-resolution images to produce a final high-resolution image. The overlay process may be performed by averaging, such as additive averaging, or by any other existing process. Furthermore, when overlaying, the image enhancement unit 404 may align multiple high-resolution images using any method before overlaying them. Once the image enhancement unit 404 has produced the final high-resolution image, the process proceeds to step S550.

[0282] In step S550, the output unit 405 displays the final high-resolution image that has been generated on the display unit 20.

[0283] As described above, the image enhancement unit 404 according to this embodiment generates a single final high-resolution image from multiple input images. Since the image enhancement by the image enhancement engine is based on the input images, for example, if a lesion or the like is not properly displayed in one input image, the high-resolution image created by enhancing that input image will have a low pixel value. On the other hand, in other input images taken of the same area, the lesion or the like may be properly displayed, and the high-resolution image created by enhancing those other input images may have a high pixel value. Therefore, by superimposing these high-resolution images, the areas with low or high pixel values ​​can be properly displayed, and a high-contrast high-resolution image can be generated. Furthermore, by using fewer input images than the number required for conventional superimposition, the costs associated with longer shooting times and other drawbacks can be minimized.

[0284] This effect is particularly noticeable when using input images that utilize motion contrast data such as OCTA.

[0285] Motion contrast data detects temporal changes in a subject over time intervals when the same point on the subject is repeatedly photographed. For example, in a given time interval, only slight movement of the subject may be detected. Conversely, if photographed at a different time interval, the subject's movement may be detected as larger. Therefore, by overlaying high-resolution images of the motion contrast from each case, it is possible to interpolate motion contrast that was not present or was only slightly detected at a particular time. Thus, this processing can generate a motion contrast image in which more of the subject's movement is enhanced, allowing the examiner to grasp the subject's state more accurately.

[0286] Therefore, when using images that depict areas that change over time, such as OCTA images, as input images, it is possible to image a specific area of ​​the subject in more detail by superimposing high-resolution images acquired at different times.

[0287] In this embodiment, a high-resolution image is generated from each of multiple input images, and then the high-resolution images are superimposed to produce a single final high-resolution image. However, the method for generating a single high-resolution image from multiple input images is not limited to this. For example, in another example of the image enhancement process of this embodiment shown in Figure 21B, when the image enhancement process is started in step S540, in step S2130, the image enhancement unit 404 superimposes the input image group to generate a single superimposed input image.

[0288] Subsequently, in step S2140, the image enhancement unit 404 inputs the superimposed input images to the image enhancement engine and generates a single high-resolution image. Even with this image enhancement process, similar to the image enhancement process described above, areas with low or high pixel values ​​in multiple input images can be appropriately displayed, and a high-contrast, high-resolution image can be generated. This process is particularly effective when motion contrast images such as the OCTA image are used as input images.

[0289] Furthermore, when performing this high-resolution processing, the same number of superimposed images of the multiple input images to be processed are used as the training data input for the high-resolution engine. This allows the high-resolution engine to perform appropriate high-resolution processing.

[0290] Furthermore, the image enhancement process according to this embodiment and the other image enhancement process described above are not limited to superimposing images. For example, a single image may be generated by applying MAP estimation processing to these image groups. Alternatively, a single image may be generated by synthesizing the image groups or input image groups.

[0291] When generating a single image by combining a group of high-resolution images or input images, for example, one might use an image with a wide tonal range in the high-luminance region and an image with a wide tonal range in the low-luminance region as input images. In this case, for example, an image enhanced from the image with a wide tonal range in the high-luminance region and an image enhanced from the image with a wide tonal range in the low-luminance region are combined. This makes it possible to generate an image that can express a wider range of brightness (dynamic range). In this case, the input data for the image enhancement engine's training data can be the images to be processed, which have a wide tonal range in the high-luminance region and low-resolution images with a wide tonal range in the low-luminance region. The output data for the image enhancement engine's training data can be the high-resolution images corresponding to the input data.

[0292] Alternatively, an image with a wide tonal range in the high-luminance region and an image with a wide tonal range in the low-luminance region may be combined, and the combined image may be enhanced in quality using a high-quality image enhancement engine. In this case as well, an image capable of representing a wider range of brightness can be generated. In this case, the input data for the high-quality image enhancement engine's training data can be an image created by combining a low-quality image with a wide tonal range in the high-luminance region and a low-quality image with a wide tonal range in the low-luminance region, which are to be processed. The output data for the high-quality image enhancement engine's training data can be a high-quality image corresponding to the input data.

[0293] In these cases, by using a high-resolution engine, it is possible to enhance the image quality of images that can express a wider range of brightness, allowing processing with fewer images compared to conventional methods, and providing images suitable for image analysis at a lower cost.

[0294] Furthermore, any method may be used to capture images with a wide tonal range in the high-luminance region and images with a wide tonal range in the low-luminance region, such as shortening or lengthening the exposure time of the imaging device. Also, the division of the tonal range is not limited to low-luminance and high-luminance regions, but may be arbitrary.

[0295] Furthermore, in the image quality enhancement process according to this embodiment, multiple image quality enhancement engines may be used to process multiple input images in parallel. The output unit 405 may output the generated high-resolution images to the imaging device 10 or other devices connected to the image processing device 400, as in the first embodiment. Also, the output data of the training data for the image quality enhancement engine is not limited to high-resolution images that have undergone superposition processing, as in the first embodiment. That is, high-resolution images obtained by performing at least one of the processing group or imaging method, such as superposition processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, imaging using a high-performance imaging device, high-cost processing, and noise reduction processing, may be used.

[0296] <Embodiment 14> Next, an image processing apparatus according to the 14th embodiment will be described with reference to Figures 4 and 5. In this embodiment, the image enhancement unit takes a medium-resolution image generated from a plurality of low-resolution images as an input image and generates a high-resolution image.

[0297] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0298] In this embodiment, the acquisition unit 401 acquires a medium-resolution image obtained by superimposing multiple low-resolution images from the imaging device 10 or other devices as input data to be processed. Note that arbitrary alignment processing may be performed when superimposing the low-resolution images.

[0299] The image enhancement unit 404 according to this embodiment is equipped with an image enhancement engine similar to that of the first embodiment. However, the image enhancement engine of this embodiment is designed to take a medium-resolution image as input and output a high-resolution image. A medium-resolution image is a superimposed image generated by superimposing multiple low-resolution images. A high-resolution image is an image with higher resolution than a medium-resolution image. Furthermore, the pairs that constitute the training data used to train the image enhancement engine are also medium-resolution images generated in the same way as the medium-resolution images, while the output data is a high-resolution image.

[0300] The output unit 405 displays the high-resolution image generated by the high-resolution image processing unit 404 on the display unit 20. The output unit 405 may also display the input image on the display unit 20 along with the high-resolution image. In this case, the output unit 405 may also display on the display unit 20 that the input image was generated from multiple low-resolution images.

[0301] Next, a series of image processing steps according to this embodiment will be described with reference to Figure 5. Note that the processing steps S520 to S550 according to this embodiment are the same as those in the first embodiment, so their description will be omitted.

[0302] In step S510, the acquisition unit 401 acquires a medium-resolution image as an input image from the imaging device 10 or other devices. The acquisition unit 401 may also acquire a medium-resolution image generated by the imaging device 10 as an input image in response to a request from the imaging device 10. Such requests may be issued, for example, when the imaging device 10 generates an image, before or after the imaging device 10 saves the generated image to the storage device provided by the imaging device 10, when displaying the saved image on the display unit 20, or when using a high-resolution image for image analysis processing.

[0303] The subsequent processing is the same as in the first embodiment, so the explanation will be omitted.

[0304] As described above, the acquisition unit 401 according to this embodiment acquires a medium-resolution image, which is an image generated using multiple images of a predetermined part of the subject, as the input image. In this case, since the input image is a clearer image, the high-resolution engine can generate a high-resolution image with greater accuracy. Note that the number of low-resolution images used to generate the medium-resolution image may be less than the number of images used to generate a conventional superimposed image.

[0305] Furthermore, a medium-resolution image is not limited to an image created by superimposing multiple low-resolution images; for example, it may be an image to which MAP estimation processing has been applied to multiple low-resolution images, or an image created by combining multiple low-resolution images. When combining multiple low-resolution images, images with different tonal ranges may be combined.

[0306] Furthermore, the output unit 405 may output the generated high-resolution images to the imaging device 10 or other devices connected to the image processing device 400, similar to the first embodiment. Also, the output data for the training data of the high-resolution engine is not limited to high-resolution images that have undergone overlay processing, similar to the first embodiment. That is, high-resolution images obtained by performing at least one of the processing group or imaging method, such as overlay processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, imaging using a high-performance imaging device, high-cost processing, or noise reduction processing, may be used.

[0307] <Embodiment 15> Next, an image processing apparatus according to the 15th embodiment will be described with reference to Figures 4 and 5. In this embodiment, the image quality enhancement unit performs image quality enhancement according to the first embodiment and the like, as well as increasing the image size (enlarging) of the input image.

[0308] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0309] The acquisition unit 401 according to this embodiment acquires a low-size image as an input image. A low-size image is an image with fewer pixels than a high-size image (high-size image) output by the high-image-quality engine described later. Specifically, for example, if the image size of a high-size image is 1024 pixels wide, 1024 pixels high, and 1024 pixels deep, then the image size of a low-size image is 512 pixels wide, 512 pixels high, and 512 pixels deep. In this regard, high-size image enhancement as used herein refers to the process of increasing the number of pixels per image and enlarging the image size.

[0310] The image enhancement unit 404 according to this embodiment is equipped with an image enhancement engine, similar to the first embodiment. However, the image enhancement engine in this embodiment is configured to reduce noise and enhance contrast in the input image, as well as to increase the image size of the input image. Therefore, the image enhancement engine in this embodiment is configured to take a low-size image as input and output a high-size image.

[0311] In this regard, for the pair group that constitutes the training data of the high-image-quality engine, the input data for each pair is a low-size image, and the output data is a high-size image. The high-size images used for output data can be acquired from a device with higher performance than the imaging device that acquired the low-size images, or by changing the settings of the imaging device. If a group of high-size images already exists, the low-size image group to be used as input data may be obtained by reducing the image size of the high-size image group to the image size of the image expected to be acquired from the imaging device 10. Furthermore, the high-size images are obtained by superimposing low-size images, as in the first embodiment.

[0312] Furthermore, the image size enlargement of the input image by the image quality enhancement unit 404 in this embodiment is obtained by acquiring training data from a device with higher performance than the shooting device 10, or by changing the settings of the shooting device 10, and therefore differs from simple image enlargement. Specifically, the image size enlargement process of the input image by the image quality enhancement unit 404 in this embodiment can reduce the degradation of resolution compared to simply enlarging an image.

[0313] With this configuration, the image enhancement unit 404 according to this embodiment can generate a high-quality image with reduced noise, enhanced contrast, and increased image size from the input image.

[0314] Next, with reference to Figure 5, a series of image processing steps according to this embodiment will be described. Note that the processing in steps S520, S530, and S550 according to this embodiment is the same as the processing in the first embodiment, so the description will be omitted. Note that if the image quality is to be increased unconditionally for the input image regardless of the shooting conditions, the processing in step S530 may be omitted after the processing in step S520, and the processing may proceed to step S540.

[0315] In step S510, the acquisition unit 401 acquires a low-size image from the imaging device 10 or other devices as input data to be processed. The acquisition unit 401 may also acquire a low-size image generated by the imaging device 10 as an input image in response to a request from the imaging device 10. Such requests may be issued, for example, when the imaging device 10 generates an image, before or after the imaging device 10 saves the generated image to the storage device provided by the imaging device 10, when displaying the saved image on the display unit 20, or when using a high-resolution image for image analysis processing.

[0316] The processing in steps S520 and S530 is the same as in the first embodiment, so a description is omitted. In step S540, the image enhancement unit 404 inputs the input image to the image enhancement engine, and generates a high-resolution image in which noise reduction and contrast enhancement are performed, and the image size is increased. The subsequent processing is the same as in the first embodiment, so a description is omitted.

[0317] As described above, the image enhancement unit 404 according to this embodiment generates a high-resolution image in which at least one of noise reduction and contrast enhancement is performed compared to the input image, and the image size is enlarged. As a result, the image processing device 400 according to this embodiment can output a high-resolution image suitable for diagnostic imaging without increasing the invasiveness or effort required of the photographer or the patient.

[0318] In this embodiment, a high-resolution image was generated by performing high-resolution processing and high-image-quality processing according to the first embodiment, etc., using a single high-resolution processing engine. However, the configuration for performing these processing is not limited to this. For example, the high-resolution processing unit may include a high-resolution processing engine that performs high-resolution processing according to the first embodiment, etc., and another high-resolution processing engine that performs high-image-quality resizing processing.

[0319] In this case, the image enhancement engine that performs the image enhancement processing according to the first embodiment can use a machine learning model that has been trained in the same way as the image enhancement engine according to the first embodiment. Furthermore, the high-resolution images generated by the image enhancement engine according to the first embodiment are used as the input data for the training data of the image enhancement engine that performs the image resizing processing. Furthermore, the high-resolution images generated by the image enhancement engine according to the first embodiment are used as the output data for the training data of the image enhancement engine, using images acquired by a high-performance imaging device. As a result, the image enhancement engine that performs the image resizing processing can generate a final high-resolution image with a larger image size from the high-resolution image that has undergone the image enhancement processing according to the first embodiment.

[0320] Furthermore, the image resizing process performed by the image enhancement engine may be performed before the image enhancement process performed by the image enhancement engine according to the first embodiment, etc. In this case, the training data for the image enhancement engine that performs the image resizing process consists of a group of pairs of input data, which are low-size images acquired by the imaging device, and output data, which are high-size images. In addition, the training data for the image enhancement engine that performs the image enhancement process according to the first embodiment, etc. consists of a group of pairs of input data, which are high-size images, and output data, which are images obtained by superimposing the high-size images.

[0321] Even with this configuration, the image processing device 400 can generate a high-quality image in which at least one of noise reduction and contrast enhancement is performed compared to the input image, and the image size is enlarged.

[0322] In this embodiment, the image enhancement process according to the first embodiment and the like describes a configuration in which superimposed images are used as output data for training data. However, as with the first embodiment, the output data is not limited to this. That is, high-quality images obtained by performing at least one of the processing group or shooting method, such as superimposition processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, shooting using a high-performance shooting device, high-cost processing, or noise reduction processing, may also be used.

[0323] The output unit 405 may also output the generated high-resolution image to the imaging device 10 or other devices connected to the image processing device 400, similar to the first embodiment.

[0324] <Embodiment 16> Next, an image processing apparatus according to the 16th embodiment will be described with reference to Figures 4 and 5. In this embodiment, the image quality enhancement unit performs high spatial resolution enhancement along with image quality enhancement according to the first embodiment and the like.

[0325] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0326] The acquisition unit 401 according to this embodiment acquires a low spatial resolution image as an input image. A low spatial resolution image is an image with a lower spatial resolution than the high spatial resolution image output by the image enhancement unit 404.

[0327] The image enhancement unit 404 is equipped with an image enhancement engine, similar to the first embodiment. However, the image enhancement engine in this embodiment is configured to reduce noise and enhance contrast in the input image, as well as to increase the spatial resolution of the input image. Therefore, the image enhancement engine according to this embodiment is configured to take a low spatial resolution image as input and output a high spatial resolution image.

[0328] In this regard, the pairs that constitute the training data for the image enhancement engine also consist of low spatial resolution images as input data and high spatial resolution images as output data for each pair. High spatial resolution images can be obtained from a more powerful imaging device than the one that acquired the low spatial resolution images, or by changing the settings of the imaging device. Furthermore, for high spatial resolution images, images obtained by superimposing low spatial resolution images are used, similar to the first embodiment.

[0329] With this configuration, the image enhancement unit 404 according to this embodiment can generate a high-quality image with reduced noise, enhanced contrast, and high spatial resolution from the input image.

[0330] Next, with reference to Figure 5, a series of image processing steps according to this embodiment will be described. Note that the processing in steps S520, S530, and S550 according to this embodiment is the same as the processing in the first embodiment, so the description will be omitted. Note that if the image quality is to be increased unconditionally for the input image regardless of the shooting conditions, the processing in step S530 may be omitted after the processing in step S520, and the processing may proceed to step S540.

[0331] In step S510, the acquisition unit 401 acquires a low spatial resolution image from the imaging device 10 or other devices as input data to be processed. The acquisition unit 401 may also acquire a low spatial resolution image generated by the imaging device 10 as an input image in response to a request from the imaging device 10. Such requests may be issued, for example, when the imaging device 10 generates an image, before or after saving an image generated by the imaging device 10 to a storage device provided by the imaging device 10, when displaying a saved image on the display unit 20, or when using a high-resolution image for image analysis processing.

[0332] The processing in steps S520 and S530 is the same as in the first embodiment, so a description is omitted. In step S540, the image enhancement unit 404 inputs the input image to the image enhancement engine, and generates a high-resolution image in which noise reduction and contrast enhancement are performed, as well as an image with high spatial resolution. The subsequent processing is the same as in the first embodiment, so a description is omitted.

[0333] As described above, the image enhancement unit 404 according to this embodiment generates a high-quality image in which at least one of noise reduction and contrast enhancement is performed compared to the input image, and the spatial resolution is improved. As a result, the image processing device 400 according to this embodiment can output a high-quality image suitable for diagnostic imaging without increasing the invasiveness or effort required of the photographer or the patient.

[0334] In this embodiment, a high-resolution image was generated by performing both the high-resolution processing and high-image-quality processing according to the first embodiment, etc., using a single high-resolution processing engine. However, the configuration for performing these processing is not limited to this. For example, the high-resolution processing unit may include a high-resolution processing engine that performs the high-resolution processing according to the first embodiment, etc., and another high-resolution processing engine that performs the high-resolution processing.

[0335] In this case, the image enhancement engine that performs the image enhancement processing according to the first embodiment, etc., can use a machine learning model that has been trained in the same way as the image enhancement engine according to the first embodiment, etc. Furthermore, the high-resolution images generated by the image enhancement engine according to the first embodiment, etc., are used as the input data for the training data of the image enhancement engine that performs the high-resolution processing. Furthermore, the high-resolution images generated by the image enhancement engine according to the first embodiment, etc., from images acquired with a high-performance imaging device are used as the output data for the training data of the said image enhancement engine. As a result, the image enhancement engine that performs high spatial resolution processing can generate a final high-resolution image with high spatial resolution applied to the high-resolution image that has undergone the high-resolution processing according to the first embodiment, etc.

[0336] Furthermore, the high spatial resolution processing performed by the high-image-quality engine may be performed before the high-image-quality processing performed by the high-image-quality processing engine according to the first embodiment, etc. In this case, the training data for the high-image-quality engine that performs the high spatial resolution processing consists of a group of pairs of input data, which are low spatial resolution images acquired by the imaging device, and output data, which are high spatial resolution images. Furthermore, the training data for the image enhancement engine that performs image enhancement processing according to the first embodiment, etc., consists of a group of pairs of input data consisting of a high spatial resolution image and output data consisting of an image obtained by superimposing the high spatial resolution image.

[0337] Even with this configuration, the image processing device 400 can generate a high-quality image in which at least one of noise reduction and contrast enhancement is performed compared to the input image, and the spatial resolution is improved.

[0338] In this embodiment, the image enhancement process according to the first embodiment and the like describes a configuration in which superimposed images are used as output data for training data. However, as with the first embodiment, the output data is not limited to this. That is, high-quality images obtained by performing at least one of the processing group or shooting method, such as superimposition processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, shooting using a high-performance shooting device, high-cost processing, or noise reduction processing, may also be used.

[0339] Furthermore, the image enhancement unit 404 may use an image enhancement engine to perform image enhancement processing according to the 15th embodiment in addition to high spatial resolution processing. In this case, at least one of noise reduction and contrast enhancement is performed compared to the input image, and an image with a larger image size and higher spatial resolution compared to the input image can be generated as a high-quality image. As a result, the image processing device 400 according to this embodiment can output high-quality images suitable for diagnostic imaging without increasing the invasiveness or effort required of the photographer or the patient.

[0340] The output unit 405 may also output the generated high-resolution image to the imaging device 10 or other devices connected to the image processing device 400, similar to the first embodiment.

[0341] <Embodiment 17> Next, an image processing apparatus according to the 17th embodiment will be described with reference to Figures 22 and 23. In this embodiment, the analysis unit performs image analysis on the high-resolution image generated by the image enhancement unit.

[0342] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment.

[0343] Figure 22 shows a schematic configuration of the image processing apparatus 2200 according to this embodiment. The image processing apparatus 2200 according to this embodiment is provided with an acquisition unit 401, a shooting condition acquisition unit 402, a high-quality image enhancement feasibility determination unit 403, a high-quality image enhancement unit 404, and an output unit 405, in addition to an analysis unit 2208. Note that the image processing apparatus 2200 may be composed of multiple devices, each equipped with some of these components. Here, the acquisition unit 401, the shooting condition acquisition unit 402, the high-quality image enhancement feasibility determination unit 403, the high-quality image enhancement unit 404, and the output unit 405 are the same as those in the image processing apparatus according to the first embodiment, so the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0344] The analysis unit 2208 applies predetermined image analysis processing to the high-resolution image generated by the image enhancement unit 404. Image analysis processing includes, for example, existing arbitrary image analysis processing in the ophthalmology field, such as retinal layer segmentation, layer thickness measurement, three-dimensional optic disc shape analysis, cribriform lamina analysis, vascular density measurement of OCTA images, and corneal shape analysis for images acquired by OCT. Furthermore, image analysis processing is not limited to ophthalmology, but also includes existing arbitrary analysis processing in the radiology field, such as diffusion tensor analysis and VBL (Voxel-based Morphometry) analysis.

[0345] The output unit 405 can display the high-resolution image generated by the image enhancement unit 404 on the display unit 20, and can also display the analysis results of the image analysis processing performed by the analysis unit 2208. The output unit 405 may display only the image analysis results from the analysis unit 2208 on the display unit 20, or it may output the image analysis results to the shooting device 10, image management system, or other devices. The display format of the analysis results may be arbitrary depending on the image analysis processing performed by the analysis unit 2208, and may be displayed as an image, numerical value, or text. Furthermore, the display format of the analysis results may be an image (for example, a 2D map) obtained by blending the analysis results obtained by analyzing the high-resolution image with the high-resolution image with an arbitrary transparency.

[0346] Hereinafter, with reference to Figure 23, a series of image processing steps according to this embodiment will be explained using an OCTA En-Face image as an example. Figure 23 is a flowchart of the series of image processing steps according to this embodiment. Note that the processing in steps S2310 to S2340 according to this embodiment is the same as the processing in steps S510 to S540 in the first embodiment, so the explanation will be omitted. Note that if the image quality is to be improved unconditionally for the input image regardless of the shooting conditions, the processing in step S2330 may be omitted after the processing in step S2320, and the processing may proceed to step S2340.

[0347] In step S2340, the image enhancement unit 404 enhances the image quality of the OCTA En-Face image, and the processing proceeds to step S2350. In step S2350, the analysis unit 2208 performs image analysis on the high-resolution image generated in step S2340. For image analysis of the enhanced OCTA En-Face image, by applying an arbitrary binarization process, areas equivalent to blood vessels (vascular regions) can be detected from the image. By determining the proportion of the image occupied by the detected areas equivalent to blood vessels, the area density can be analyzed. Alternatively, by thinning the binarized areas equivalent to blood vessels, an image with a line width of 1 pixel can be created, and the proportion occupied by blood vessels that do not depend on thickness (also called skeleton density) can be determined. These images may be used to analyze the area and shape (circularity, etc.) of the avascular zone (FAZ). As for the analysis method, the above values ​​may be calculated from the entire image, or a user interface (not shown) may be used to calculate values ​​for a specified region of interest (ROI) based on the examiner's (user's) instructions. The ROI setting is not necessarily specified by the examiner; a predetermined area may be automatically selected. Here, the various parameters described above are just examples of analysis results related to blood vessels, and any parameters related to blood vessels may be used. The analysis unit 2208 may perform multiple image analysis processes. That is, although an example of analysis on OCTA En-Face images is shown here, it may also perform retinal layer segmentation, layer thickness measurement, optic disc three-dimensional shape analysis, lamina cribriform analysis, etc., on images acquired simultaneously by OCT. In this regard, the analysis unit 2208 may perform some or all of the multiple image analysis processes in response to instructions from the examiner via any input device.

[0348] In step S2360, the output unit 405 displays the high-resolution image generated by the image enhancement unit 404 and the analysis results from the analysis unit 2208 on the display unit 20. The output unit 405 may output the high-resolution image and the analysis results to separate display units or devices. Alternatively, the output unit 405 may display only the analysis results on the display unit 20. Furthermore, if the analysis unit 2208 outputs multiple analysis results, the output unit 405 may output some or all of the multiple analysis results to the display unit 20 or other devices. For example, the analysis results regarding blood vessels in the OCTA En-Face image may be displayed on the display unit 20 as a two-dimensional map. Alternatively, values ​​indicating the analysis results regarding blood vessels in the OCTA En-Face image may be superimposed on the OCTA En-Face image and displayed on the display unit 20.

[0349] As described above, the image processing apparatus 2200 according to this embodiment further includes an analysis unit 2208 that performs image analysis on high-resolution images, and the output unit 405 displays the analysis results from the analysis unit 2208 on the display unit 20. In this way, the image processing apparatus 2200 according to this embodiment uses high-resolution images for image analysis, thereby improving the accuracy of the analysis.

[0350] Furthermore, the output unit 405 may output the generated high-resolution images to the imaging device 10 or other devices connected to the image processing device 2200, similar to the first embodiment. Also, the output data for the training data of the high-resolution engine is not limited to high-resolution images that have undergone overlay processing, similar to the first embodiment. That is, high-resolution images obtained by performing at least one of the processing group or imaging method, such as overlay processing, MAP estimation processing, smoothing filter processing, grayscale conversion processing, imaging using a high-performance imaging device, high-cost processing, or noise reduction processing, may be used.

[0351] <Embodiment 18> Next, with reference to Figure 4, an image processing apparatus according to the 18th embodiment will be described. In this embodiment, an example will be described in which the image enhancement unit generates a high-resolution image by adding noise to the image during training and learning the noise component.

[0352] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0353] The acquisition unit 401 according to this embodiment acquires images from the imaging device 10 or other devices as input data to be processed. An example of the CNN configuration in the image enhancement unit according to this embodiment will be explained using Figure 24. Figure 24 shows an example of the machine learning model configuration in the image enhancement unit 404. The configuration shown in Figure 24 consists of multiple layers that are responsible for processing the input value group and outputting it. The types of layers included in the above configuration are, as shown in Figure 24, a convolution layer, a downsampling layer, an upsampling layer, and a merger layer. The convolution layer is a layer that performs convolution processing on the input value group according to parameters such as the kernel size of the set filter, the number of filters, the stride value, and the dilation value. The dimensionality of the kernel size of the filter may also be changed according to the dimensionality of the input image. The downsampling layer is a process that reduces the number of output value groups to less than the number of input value groups by decimating or merging the input value group. Specifically, there is a Max Pooling process, for example. An upsampling layer is a process that increases the number of output values ​​to more than the number of input values ​​by duplicating the input value set or adding interpolated values ​​from the input value set. Specifically, there is a linear interpolation process, for example. A synthesis layer is a layer that takes value sets, such as the output value set of a certain layer or the pixel value set that makes up an image, as input from multiple sources and synthesizes them by concatenating or adding them. In such a configuration, the pixel value set that makes up the input image Im2410 is synthesized in the synthesis layer with the value set output after the convolutional processing block. After that, the synthesized pixel value set is shaped into a high-resolution image Im2420 in the final convolutional layer. Although not shown in the diagram, as an example of changing the CNN configuration, for example, a batch normalization layer or an activation layer using a Rectifier Linear Unit can be incorporated after the convolutional layer.

[0354] The image enhancement engine of this embodiment takes low-resolution images obtained from the imaging device 10 or other devices with a first noise component added as input, and uses images obtained from the imaging device 10 or other devices with a second noise component added as output data to train the engine as high-resolution images. In other words, the training images used during training in this embodiment are the same images for both the low-resolution and high-resolution images, but the noise components in each image are different. Since the same images are used, alignment is not required when pairing the images.

[0355] The noise components added include Gaussian noise and noise modeled after noise specific to the target image. However, the first and second noises will be different types of noise. Different noise refers to noise that is added in different spatial locations (pixel positions) or has different noise values. For example, in the case of OCT, noise specific to a given image can be estimated based on data acquired without a model eye or the patient's eye, and these can be used as noise models. In the case of OCTA, noise appearing in the avascular zone (FAZ) or noise appearing in images of a model eye that schematically reproduces blood flow can be used as noise models.

[0356] In the case of Gaussian noise, the standard deviation or variance is defined as the magnitude of the noise, and noise is randomly applied to the image based on these values. The average value of the image as a result of applying random noise may remain unchanged. That is, the average value of the noise added to each pixel of an image should be 0. Here, it is not necessary for the average value to be 0; it is sufficient if different patterns of noise are applied to the input data and the output data. Also, it is not necessary to apply noise to both the input data and the output data; noise can be applied to only one of them. Here, if no noise is applied, for example, false images of blood vessels may appear in the high-resolution image, but this can be thought to occur when the difference between the image before and after high-resolution is relatively large. For this reason, it may be possible to reduce the difference between the image before and after high-resolution. In this case, during training, two images obtained by applying different patterns of noise to a low-resolution image and a high-resolution image may be used as paired images, or two images obtained by applying different patterns of noise to a high-resolution image may be used as paired images.

[0357] The output unit 405 displays the high-resolution image generated by the high-resolution image processing unit 404 on the display unit 20. The output unit 405 may also display the input image on the display unit 20 along with the high-resolution image.

[0358] The subsequent processing is the same as in the first embodiment, so the explanation will be omitted.

[0359] In this embodiment, high-resolution images were generated using images obtained by adding a first noise component and a second noise component different from the first noise component to low-resolution images acquired from the imaging device 10 or other devices. However, the configuration for performing these processes is not limited to this. For example, the image to which noise is added may be a high-resolution image that has undergone the superposition process shown in the first embodiment, to which the first and second noise components are added. That is, the image obtained by adding the first noise component to the superposition image may be learned as a low-resolution image, and the image obtained by adding the second noise component to the superposition image may be learned as a high-resolution image.

[0360] Furthermore, although this embodiment describes an example of learning using first and second noise components, it is not limited to this. For example, the first noise component may be added only to the low-resolution image, and learning may be performed without adding a noise component to the high-resolution image. In this case, the images may be images obtained from the shooting device 10 or other devices, or images obtained by superimposing such images may be used as the target.

[0361] Furthermore, the magnitude of the noise component may be dynamically changed depending on the type of input image or for each rectangular region image being trained. Specifically, adding a large noise value will increase the noise reduction effect, while adding a small noise value will decrease the noise reduction effect. Therefore, for example, the noise added may be adjusted according to the conditions and type of the entire image or rectangular region image, such as decreasing the value of the noise component when the image is dark and increasing the value of the noise component when the image is bright, and then the training may be performed accordingly.

[0362] Although the image capture conditions are not specified in this embodiment, training should be performed using images with various shooting ranges and scan counts, as well as frontal images of different shooting locations and depths.

[0363] The above describes images obtained from the imaging device 10 or other devices, noised images obtained by adding noise to those images, superimposed images, and superimposed images obtained by adding noise to the superimposed images. However, these combinations are not limited to those described above, and low-resolution and high-resolution images can be combined in any way.

[0364] <Embodiment 19> Next, with reference to Figures 25 and 26, an image processing apparatus according to the 19th embodiment will be described. In this embodiment, the image enhancement unit is equipped with multiple image enhancement engines and generates multiple high-resolution images from an input image. Then, an example will be described in which the synthesis unit 2505 synthesizes the multiple high-resolution images output from the multiple image enhancement engines.

[0365] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0366] The acquisition unit 401 according to this embodiment acquires images from the imaging device 10 or other devices as input data to be processed.

[0367] The image enhancement unit 404 according to this embodiment is equipped with multiple image enhancement engines, similar to the second embodiment. Here, each of the multiple image enhancement engines is trained using different training data for at least one of the following: shooting area, shooting angle of view, front images at different depths, noise components, and image resolution. The image enhancement unit 404 generates a high-resolution image using multiple image enhancement engines corresponding to at least one of the following: shooting area, shooting angle of view, front images at different depths, noise components, and image resolution of the input image.

[0368] Figure 26 is a flowchart of a series of image processing steps according to this embodiment. Note that the processing in steps S2610 and S2620 according to this embodiment is the same as the processing in steps S510 and S520 in the first embodiment, so a description is omitted. Note that if the image quality is to be increased unconditionally for the input image regardless of the shooting conditions, the processing in step S2630 may be omitted after the processing in step S2620, and the processing may proceed to step S2640.

[0369] In step S2620, as in the first embodiment, once the shooting condition acquisition unit 402 has acquired a set of shooting conditions for the input image, the process moves to step S2630. In step S2630, the image quality enhancement feasibility determination unit 403, as in the second embodiment, uses the acquired set of shooting conditions to determine whether any of the image quality enhancement engines provided in the image quality enhancement unit 404 can process the input image.

[0370] If the image quality enhancement feasibility determination unit 403 determines that none of the image quality enhancement engines can process the input image, the process proceeds to step S2660. On the other hand, if the image quality enhancement feasibility determination unit 403 determines that any of the image quality enhancement engines can process the input image, the process proceeds to step S2640. Depending on the settings and implementation configuration of the image processing device 400, step S2640 may also be performed, as in the first embodiment, even if some shooting conditions are determined to be unprocessable by the image quality enhancement engines.

[0371] In step S2640, the image enhancement unit 404 inputs the input image acquired in step S2610 to each of the image enhancement engines and generates a group of high-resolution images.

[0372] In step S2650, the synthesis unit 2405 synthesizes several high-resolution images from the high-resolution image group generated in step S2640. Specifically, for example, it synthesizes the results of two high-resolution images: a first high-resolution engine trained using paired images of low-resolution images acquired from the imaging device 10 and high-resolution images obtained by superimposing a group of images acquired by taking multiple low-resolution images, as shown in the first embodiment; and a second high-resolution engine trained using paired images to which noise has been added, as shown in the 18th embodiment. The synthesis method can be performed using methods such as averaging or weighted averaging.

[0373] In step S2660, the output unit 405 displays the image synthesized in step S2650 on the display unit 20 or outputs it to another device. However, if it is determined in step S2630 that the input image is unprocessable, the output unit 405 outputs the input image as the output image. The output unit 405 may also indicate on the display unit 20 that the output image is the same as the input image if the examiner specifies an input image or if the input image is unprocessable.

[0374] <20th Embodiment> Next, with reference to Figure 4, an image processing apparatus according to the 20th embodiment will be described. In this embodiment, an example will be described in which the image enhancement unit uses the output result of the first image enhancement engine to generate a high-resolution image in which the second image enhancement engine generates a high-resolution image.

[0375] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0376] The acquisition unit 401 according to this embodiment acquires images from the imaging device 10 or other devices as input data to be processed.

[0377] The image enhancement unit 404 according to this embodiment is equipped with multiple image enhancement engines, similar to those in the first embodiment. The image enhancement unit of this embodiment includes a first image enhancement engine that learns low-resolution images acquired as input data from the imaging device 10 or other devices, and medium-resolution images generated from multiple low-resolution images, as output data. Furthermore, it includes a second image enhancement engine that learns images output from the first image enhancement engine and images of higher resolution than the medium-resolution images as output data. Note that the medium-resolution images are the same as in the 14th embodiment, so their explanation will be omitted.

[0378] The output unit 405 displays the high-resolution image generated by the high-resolution image processing unit 404 on the display unit 20. The output unit 405 may also display the input image on the display unit 20 along with the high-resolution image. In this case, the output unit 405 may also display on the display unit 20 that the input image was generated from multiple low-resolution images.

[0379] Next, a series of image processing steps according to this embodiment will be described with reference to Figure 5. Note that the processing steps S510 to S530 according to this embodiment are the same as those in the first embodiment, so their description will be omitted.

[0380] In step S540, the image enhancement unit 404 uses an image enhancement engine to enhance the image quality of the input image, generating a higher-quality image more suitable for image diagnosis than the input image. Specifically, the image enhancement unit 404 inputs the input image to a first image enhancement engine to generate a first high-quality image. Furthermore, it inputs the first high-quality image to a second image enhancement engine to obtain a second high-quality image. The image enhancement engine generates a high-quality image that appears as if it has been overlaid using the input image, based on a machine learning model that has been machine-learned using training data. As a result, the image enhancement engine can generate a higher-quality image that has reduced noise and enhanced contrast than the input image.

[0381] The subsequent processing is the same as in the first embodiment, so the explanation will be omitted.

[0382] In this embodiment, high-resolution images were generated using a first image enhancement engine that learned pairs of low-resolution and medium-resolution images obtained from the imaging device 10 or other devices, and a second image enhancement engine that learned pairs of the first high-resolution image and another high-resolution image. However, the configuration for these processes is not limited to this. For example, the image pairs learned by the first image enhancement engine may be those of an engine that learns noise as described in the 18th embodiment, and the second image enhancement engine may learn pairs of the first high-resolution image and another high-resolution image. Conversely, the first image enhancement engine may learn pairs of low-resolution and medium-resolution images, and the second image enhancement engine may learn images in which noise has been added to the first high-resolution image.

[0383] Furthermore, both the first and second image enhancement engines may be noise-learning engines as described in the 18th embodiment. In this case, for example, the first image enhancement engine learns pairs of images in which the first and second noises are added to the high-resolution images generated by the superimposed image processing, and the second image enhancement engine learns pairs of images in which the first and second noises are added to the first high-resolution image generated by the first image enhancement engine. In this embodiment, two image enhancement engines have been described, but the system is not limited to these, and a third, fourth, and other engines may be added to further link the processing. By improving the quality of the images used for learning, a network is constructed that can more easily generate smoother and sharper images.

[0384] <21st Embodiment> Next, an image processing apparatus according to the 21st embodiment will be described with reference to Figures 4 and 27. In the first embodiment, the image enhancement unit 404 was equipped with one image enhancement engine. In contrast, in this embodiment, the image enhancement unit includes multiple image enhancement engines that have performed machine learning using different training data, and generates multiple high-resolution images from the input image.

[0385] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the second embodiment. Therefore, the following description will focus on the differences between the image processing apparatus according to this embodiment and the image processing apparatus according to the first and second embodiments. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first and second embodiments, the configuration shown in Figure 4 will be indicated using the same reference numerals, and its explanation will be omitted.

[0386] The image enhancement unit 404 according to this embodiment is equipped with two or more image enhancement engines that have been machine-learned using different training data. Here, the method for creating the training data set according to this embodiment will be explained. First, a group of pairs of original images as input data and superimposed images as output data are prepared, each captured with images of various shooting ranges and different scan counts. To explain using OCT and OCTA as examples, for instance, there is a first pair of image sets captured in a 3x3mm area with 300 A scans and 300 B scans, and a second pair of image sets captured in a 10x10mm area with 500 A scans and 500 B scans. In this case, the scan density of the first pair of image sets and the second pair of image sets differs by a factor of two. Therefore, these image sets are grouped separately. If there is an image set captured in a 6x6mm area with 600 A scans and 600 B scans, it is grouped with the first image set. In other words, here, image sets with the same or approximately the same scan density (with an error of about 10%) are grouped together.

[0387] Next, we create training data sets by grouping pairs according to scan density. For example, a training data set is created such that the first training data consists of pairs acquired by scanning at a first scan density, and the second training data consists of pairs acquired by scanning at a second scan density.

[0388] Subsequently, separate image enhancement engines are trained using each set of training data. For example, a group of image enhancement engines is prepared, such as a first image enhancement engine corresponding to a machine learning model trained with the first set of training data, and a second image enhancement engine corresponding to a machine learning model trained with the second set of training data.

[0389] Because each of these image enhancement engines uses different training data for its corresponding machine learning model, the degree to which it can enhance the image quality of an input image varies depending on the shooting conditions of the image input to the engine. Specifically, the first image enhancement engine enhances the image quality to a high degree for input images acquired at the first scan density, and to a low degree for images acquired at the second scan density. Similarly, the second image enhancement engine enhances the image quality to a high degree for input images acquired at the second scan density, and to a low degree for images acquired at the first scan density.

[0390] On the other hand, during training, it may not be possible to collect a sufficient number of images with various shooting ranges and scan densities as training data. In such cases, a high-image-quality enhancement engine that has learned the noise component is prepared for those image sets, as shown in the 18th embodiment.

[0391] The image enhancement engine, which has learned the noise component, is less affected by the scan density during shooting. Therefore, when an image with a scan density that has not been trained is input, this engine is applied.

[0392] Since each training data set is composed of pairs grouped by scan density, the image quality characteristics of the image groups constituting each pair are similar. Therefore, the image quality enhancement engine can enhance image quality more effectively than the image quality enhancement engine according to the first embodiment, provided that the scan density is the same. Note that the shooting conditions for grouping the training data pairs are not limited to scan density, but may also be the shooting area, images at different depths in the case of frontal images, or a combination of two or more of these.

[0393] The series of image processing steps according to this embodiment will now be described with reference to Figure 27. Figure 27 is a flowchart of the series of image processing steps according to this embodiment. Note that the processing in steps S2710 and S2720 is the same as in steps S510 and S520 of the first embodiment, so the explanation will be omitted.

[0394] In step S2720, once the shooting conditions for the input image are acquired, the process moves to step S2730. In step S2730, the image quality enhancement feasibility determination unit 403 uses the set of shooting conditions acquired in step S2720 to determine whether any of the image quality enhancement engines provided by the image quality enhancement unit 404 can process the input image.

[0395] If the image quality enhancement feasibility determination unit 403 determines that the shooting conditions are outside the specified range, the process proceeds to step S2770. On the other hand, if the image quality enhancement feasibility determination unit 403 determines that the shooting conditions are within the specified range, the process proceeds to step S2740.

[0396] In step S2740, the image enhancement unit 404 selects an image enhancement engine from the image enhancement engine group based on the shooting conditions of the input image acquired in step S2720 and the training data information of the image enhancement engine group. Specifically, for example, with respect to the scan density among the shooting conditions acquired in step S2720, it selects an image enhancement engine that has training data information regarding scan density and has a high degree of image enhancement. In the above example, if the scan density is the first scan density, the image enhancement unit 404 selects the first image enhancement engine.

[0397] On the other hand, in step S2770, the image enhancement unit 404 selects an image enhancement engine that has learned the noise components.

[0398] In step S2750, the image enhancement unit 404 generates a high-resolution image by enhancing the image quality of the input image using the image enhancement engine selected in steps S2740 and S2770. Then, in step S2760, the output unit 405 outputs the high-resolution image from step S2750 and displays it on the display unit 20. When displaying the high-resolution image on the display unit 20, the output unit 405 may also indicate that it is a high-resolution image generated using the image enhancement engine selected by the image enhancement unit 404.

[0399] As described above, the image enhancement unit 404 according to this embodiment includes a plurality of image enhancement engines, each trained using different training data. Here, each of the plurality of image enhancement engines is trained using different training data for at least one of the following: shooting area, shooting angle of view, frontal images at different depths, and image resolution. Furthermore, for data for which sufficient correct data (output data) could not be collected, training was performed using noise components. The image enhancement unit 404 generates a high-resolution image using an image enhancement engine corresponding to at least one of these.

[0400] With this configuration, the image processing apparatus 400 according to this embodiment can generate more effective high-resolution images.

[0401] <22nd Embodiment> Next, an image processing apparatus according to the 22nd embodiment will be described with reference to Figures 30 to 32. In this embodiment, a wide-angle image generation unit generates a wide-angle image (panoramic image) using a plurality of high-resolution images generated by the high-resolution enhancement unit.

[0402] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0403] Figure 31(a) is a flowchart of a series of image processing steps according to this embodiment. In step S3110, the acquisition unit 401 acquires multiple images (at least two) as input data from the imaging device 10 or other devices. The multiple images are images of different locations on the same subject (such as the eye under examination), and are images that do not completely overlap with the subject, but rather capture areas where parts of the images overlap. To explain using the case of photographing the eye under examination as an example, by changing the position of the fixation lamp during shooting and having the eye under examination fixate on that fixation lamp, it is possible to acquire images of different locations on the same eye under examination. When taking images, it is desirable to change the position of the fixation lamp so that at least 20% of the overlapping area between adjacent images is the same. Figure 32(a) shows an example of an OCTA En-Face image taken by changing the position of the fixation lamp so that parts of adjacent images overlap. Figure 32(a) shows an example of taking five images of different locations by changing the position of the fixation lamp. Note that while Figure 32 shows five images as an example, any two or more images are acceptable, not just five.

[0404] Note that the processing in step S3120 according to this embodiment is the same as the processing in step S520 in the first embodiment, so the explanation is omitted. If the image quality is to be unconditionally improved for the input image regardless of the shooting conditions, the processing in step S3130 may be omitted after the processing in step S3120, and the processing may proceed to step S3140.

[0405] In step S3120, as in the first embodiment, once the shooting condition acquisition unit 402 has acquired a set of shooting conditions for the input image, the process moves to step S3130. In step S3130, the image quality enhancement feasibility determination unit 403, as in the first embodiment, uses the acquired set of shooting conditions to determine whether the image quality enhancement engine in the image quality enhancement unit 404 can process the input image.

[0406] If the image quality enhancement feasibility determination unit 403 determines that the image quality enhancement engine cannot handle multiple input images, the process proceeds to step S3160. On the other hand, if the image quality enhancement feasibility determination unit 403 determines that the image quality enhancement engine can handle multiple input images, the process proceeds to step S3140. Depending on the settings and implementation configuration of the image processing device 400, step S3140 may also be performed, as in the first embodiment, even if the image quality enhancement engine determines that some shooting conditions cannot be handled.

[0407] In step S3140, the image enhancement unit 404 performs processing on the multiple input images acquired in step S3110 to generate multiple high-resolution images.

[0408] In step S3150, the wide-angle image generation unit 3005 synthesizes several high-resolution images from the high-resolution image group generated in step S3140. Specifically, this will be explained using OCTA En-Face images as an example. The multiple images are OCTA En-Face images taken so that they do not completely overlap, but adjacent images partially overlap each other. Therefore, the wide-angle image generation unit 3005 detects the overlapping regions from the multiple OCTA En-Face images and performs alignment using the overlapping regions. By deforming the OCTA En-Face images based on the alignment parameters and synthesizing the images, it is possible to generate an OCTA En-Face image with a wider range than a single OCTA En-Face image. At this time, since the multiple input OCTA En-Face images have been enhanced in quality in step S3140, the wide-angle OCTA En-Face image output in step S3150 is already enhanced in quality. Figure 32(b) shows an example of a wide-angle OCTA En-Face image generated by the wide-angle image generation unit 3005. Figure 32(b) shows an example of images generated by aligning the five images shown in Figure 32(a). Figure 32(c) shows the positional correspondence between Figure 32(a) and Figure 32(b). As shown in Figure 32(c), Im3210 is at the center, with Im3220~3250 arranged around it. Note that multiple OCTA En-Face images can be generated by setting different depth ranges from 3D motion contrast data. Therefore, although Figure 32 shows an example of a wide-angle surface image, it is not limited to this. For example, the surface OCTA En-Face image (Im2910) shown in Figure 29 could be used for alignment, and the En-Face images of OCTAs in other depth ranges could be deformed using the parameters obtained there. Alternatively, the input image for alignment can be a color image, and a composite color image can be generated using the RG component of the RGB components as the surface OCTA En-Face image and the B component as the OCTA En-Face image to be aligned. Then, alignment can be performed on the composite color OCTA En-Face image, which is created by combining layers of multiple depth ranges into a single image. By doing so, if only the B component is extracted from the aligned color OCTA En-Face image, a wide-angle OCTA En-Face image with the target OCTA En-Face image already aligned can be obtained. Note that the target for image quality enhancement is not limited to 2D OCTA En-Face images; 3D OCT and 3D motion contrast data themselves can also be used. In that case, alignment can be performed on the 3D data to generate wide-range 3D data. By cutting out an arbitrary cross-section (any XYZ plane is possible) or an arbitrary depth range (range in the Z direction) from the wide-range 3D data, a high-quality wide-angle image can be generated.

[0409] In step S3160, the output unit 405 displays the image synthesized from multiple images in step S3150 on the display unit 20 or outputs it to another device. However, if it is determined in step S3130 that the input image is unprocessable, the output unit 405 outputs the input image as the output image. The output unit 405 may also indicate on the display unit 20 that the output image is the same as the input image if the examiner specifies an input image or if the input image is unprocessable.

[0410] In this embodiment, a high-resolution image is generated from multiple input images, and then the high-resolution images are aligned to produce a single high-resolution wide-angle image. However, the method for generating a single high-resolution image from multiple input images is not limited to this. For example, in another example of the image enhancement process of this embodiment shown in Figure 31(b), a single wide-angle image may be generated first, and then the image enhancement process may be performed on the wide-angle image to produce a single high-resolution wide-angle image.

[0411] This process will be explained using Figure 31(b), but the part of the process that is the same as in Figure 31(a) will be omitted from the explanation.

[0412] In step S3121, the wide-angle image generation unit 3005 synthesizes multiple images acquired in step S3110. The wide-angle image generation is the same as described in step S3150, but the difference is that the input images are images acquired from the imaging device 10 or other devices, and are images before high-resolution processing.

[0413] In step S3151, the image enhancement unit 404 performs processing on the high-resolution image generated by the wide-angle image generation unit 3005 to generate a single high-resolution wide-angle image.

[0414] With this configuration, the image processing device 400 according to this embodiment can generate wide-angle, high-resolution images.

[0415] In the first to 22 embodiments described above, the display of high-resolution images on the display unit 20 by the output unit 405 is basically performed automatically in accordance with the generation of high-resolution images by the image enhancement unit 404 and the output of analysis results by the analysis unit 2208. However, the display of high-resolution images may also be done in accordance with instructions from the examiner. For example, the output unit 405 may display on the display unit 20 an image selected from the high-resolution images generated by the image enhancement unit 404 and the input images in accordance with instructions from the examiner. In addition, the output unit 405 may switch the display on the display unit 20 from the captured image (input image) to the high-resolution image in accordance with instructions from the examiner. That is, the output unit 405 may change the display of a low-resolution image to the display of a high-resolution image in accordance with instructions from the examiner. In addition, the output unit 405 may change the display of a high-resolution image to the display of a low-resolution image in accordance with instructions from the examiner. Furthermore, the image enhancement unit 404 may initiate image enhancement processing by the image enhancement engine (input of an image to the image enhancement engine) in response to instructions from the examiner, and the output unit 405 may display the high-resolution image generated by the image enhancement unit 404 on the display unit 20. Alternatively, when an input image is captured by the imaging device 10, the image enhancement engine may automatically generate a high-resolution image based on the input image, and the output unit 405 may display the high-resolution image on the display unit 20 in response to instructions from the examiner. These processes can also be performed similarly for the output of analysis results. That is, the output unit 405 may change the display of the analysis results for low-resolution images to the display of the analysis results for high-resolution images in response to instructions from the examiner. Also, the output unit 405 may change the display of the analysis results for high-resolution images to the display of the analysis results for low-resolution images in response to instructions from the examiner. Of course, the output unit 405 may also change the display of the analysis results for low-resolution images to the display of low-resolution images in response to instructions from the examiner. Furthermore, the output unit 405 may change the display of the low-resolution image to the display of the analysis results of the low-resolution image, in response to instructions from the examiner. Also, the output unit 405 may change the display of the analysis results of the high-resolution image to the display of the high-resolution image, in response to instructions from the examiner. Furthermore, the output unit 405 may change the display of the high-resolution image to the display of the analysis results of the high-resolution image, in response to instructions from the examiner.Furthermore, the output unit 405 may change the display of the analysis results for low-resolution images to the display of other types of analysis results for low-resolution images, in response to instructions from the examiner. The output unit 405 may also change the display of the analysis results for high-resolution images to the display of other types of analysis results for high-resolution images, in response to instructions from the examiner. Here, the display of the analysis results for high-resolution images may be a superimposed display of the analysis results for high-resolution images with an arbitrary degree of transparency. Similarly, the display of the analysis results for low-resolution images may be a superimposed display of the analysis results for low-resolution images with an arbitrary degree of transparency. In this case, the change to displaying the analysis results may, for example, be a change to a state where the analysis results are superimposed on the displayed image with an arbitrary degree of transparency. Alternatively, the change to displaying the analysis results may be a change to displaying an image obtained by blending the analysis results and the image with an arbitrary degree of transparency (e.g., a 2D map). Furthermore, the image processing device may be configured to start processing by the shooting location estimation engine, image quality evaluation engine, authenticity evaluation engine, and evaluation unit in response to instructions from the examiner. Regarding the embodiments 1 to 22 described above, the display mode in which the output unit 405 displays the high-resolution image on the display unit 20 is arbitrary. For example, the output unit 405 may display the input image and the high-resolution image side by side, or switch between them. The output unit 405 may also display the input image and the high-resolution image in order according to the shooting location, shooting date and time, facility where the shooting took place, etc. Similarly, the output unit 405 may display the image analysis results using the high-resolution image in order according to arbitrary shooting conditions of the high-resolution image and the input image corresponding to the high-resolution image. Furthermore, the output unit 405 may display the image analysis results using the high-resolution image in order for each analysis item.

[0416] <Embodiment 23> Next, an image processing apparatus according to the 23rd embodiment will be described with reference to Figures 4, 29, and 33. In this embodiment, training is performed using training data consisting of a group of pairs of output data, which are high-resolution images corresponding to input data. At that time, one high-resolution engine is generated using multiple high-resolution output data generated by multiple high-resolution engines.

[0417] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment. Since the configuration of the image processing apparatus according to this embodiment is the same as that of the image processing apparatus according to the first embodiment, the same reference numerals are used to indicate the configuration shown in Figure 4, and their explanation is omitted.

[0418] The acquisition unit 401 according to this embodiment acquires images from the imaging device 10 or other devices as input data to be processed. The generation of the image quality enhancement engine in the image quality enhancement unit 404 according to this embodiment will be explained using Figures 29 and 33. First, the first learning in this embodiment will be explained using Figure 33(a). Figure 33(a) shows an example of a group of multiple input data and output data pairs and multiple image quality enhancement engines. Im3311 and Im3312 represent a group of input data and output data pairs. For example, this pair is the same as the surface layer (Im2910) pair shown in Figure 29. 3313 represents an image quality enhancement engine that has been trained using the Im3311 and Im3312 pair. Note that the learning in Figure 33(a) may use a method that uses high-resolution images generated by superposition processing as described in the first embodiment, or a method that learns noise components as described in the 18th embodiment. Alternatively, a combination of these may be used. Im3321 and Im3322 are pairs of input and output data, for example, the deep layer (Im2920) pair shown in Figure 29. Then, 3323 represents the image enhancement engine trained on the Im3321 and Im3322 pair. Similarly, Im3331 and Im3332 are pairs of input and output data, for example, the outer layer (Im2930) pair shown in Figure 29. Then, 3333 represents the image enhancement engine trained on the Im3331 and Im3332 pair. In other words, in Figure 33(a), training is performed for each image. Therefore, for example, in the case of noise components described in the 18th embodiment, training can be performed with noise parameters appropriate for each image. In this case, the image enhancement engine may include a machine learning engine obtained using training data in which noise corresponding to the state of at least a portion of the medical image is added to that at least portion of the region. Here, the noise corresponding to the above conditions may be, for example, noise of a magnitude corresponding to the pixel values ​​in at least some areas. Also, the noise corresponding to the above conditions may be small noise if, for example, there are few features in at least some areas (e.g., small pixel values, low contrast, etc.).Furthermore, the noise corresponding to the above conditions may be large if, for example, there are many features in at least some areas (e.g., large pixel values, high contrast, etc.). The image enhancement engine may also include a machine learning engine obtained using training data that includes multiple frontal images to which noise of different magnitudes has been added to at least two of the multiple depth ranges. In this case, for example, in the depth range corresponding to a frontal image with few features (e.g., small pixel values), frontal images with small noise may be used as training data. Also, for example, in the depth range corresponding to a frontal image with many features (e.g., large pixel values), frontal images with large noise may be used as training data. In the depth range corresponding to a frontal image with a moderate amount of features, frontal images with moderate noise may be used as training data. Here, the multiple depth ranges may have overlapping portions of two adjacent depth ranges in the depth direction.

[0419] Next, we will explain the image inference in this embodiment using Figure 33(b). Figure 33(b) shows the generation of images using the high-resolution engines 3313 to 3333 that were trained in Figure 33(a). For example, when a low-resolution surface image Im3310 is input to the high-resolution engine 3313, which has been trained using multiple surface images, a high-resolution surface image Im3315 is output. Similarly, when a low-resolution deep image Im3320 is input to the high-resolution engine 3323, which has been trained using multiple deep images, a high-resolution deep image Im3325 is output. In the same way, when a low-resolution outer layer image Im3330 is input to the high-resolution engine 3333, which has been trained using multiple outer layer images, a high-resolution outer layer image Im3335 is output.

[0420] Next, the second learning process in this embodiment will be explained using Figure 33(c). Figure 33(c) shows how a single image enhancement engine 3300 is trained using multiple sets of different image pairs. Im3310 represents a low-resolution surface image, Im3315 represents a set of high-resolution surface image pairs, Im3320 represents a low-resolution deep image, Im3325 represents a set of high-resolution deep image pairs, Im3330 represents a low-resolution outer image, and Im3335 represents a set of high-resolution outer image pairs. In other words, the image enhancement engine 3300 is generated using training data consisting of sets of pairs of output data, which are high-resolution images generated using the image enhancement engine trained in the first learning process, and low-resolution input data. As a result, the image enhancement engine 3300 can generate high-resolution images suitable for image diagnosis, with reduced noise and high contrast, from various types of input images.

[0421] The output unit 405 displays the high-resolution image generated by the image enhancement unit 404 on the display unit 20. The output unit 405 may also display the input image on the display unit 20 along with the high-resolution image.

[0422] The subsequent processing is the same as in the first embodiment, so the explanation will be omitted.

[0423] In this embodiment, OCTA En-Face images were described using three layers of different depths, but the types of images are not limited to these. The number of types can be increased by generating OCTA En-Face images with different depth ranges by changing the reference layer and offset values. The types of images are not limited to differences in depth, but may also be differences by body part. For example, they may be from different locations, such as the anterior and posterior segments of the eye. Furthermore, the images are not limited to OCTA En-Face images, but may also be luminance En-Face images generated from OCT data. In the first training, the images may be trained separately, and in the second training, these OCTA En-Face images and luminance En-Face images may be trained together. Moreover, not only En-Face images, but also tomographic images, SLO images, fundus photographs, fluorescein photographs, and other images from different imaging devices may be used.

[0424] Although the second learning process described an example where there is only one image enhancement engine, it is not necessarily required to have only one. Any configuration of the image enhancement engine that learns using pairs of output data generated in the first learning process and low-resolution input data is acceptable. Furthermore, in the second learning process, Figure 33(c) shows an example where multiple pairs of different types of images are used for simultaneous learning, but transfer learning is also acceptable. For example, the network could be trained using pairs of surface images Im3310 and Im3315, and then that network could be used to train pairs of deep images Im3320 and Im3325, ultimately generating the image enhancement engine 3300.

[0425] With this configuration, the image enhancement unit 404 according to this embodiment can generate more effective high-resolution images for various types of images.

[0426] <24th Embodiment> Next, with reference to Figure 34, an image processing apparatus according to the 24th embodiment will be described. In this embodiment, an example will be described in which the output unit 405 displays the processing results of the image enhancement unit 404 on the display unit 20. In this embodiment, Figure 34 will be used for explanation, but the display screen is not limited to this. The image enhancement processing can also be applied to display screens that display multiple images obtained at different dates and times side by side, such as progress observation. Furthermore, the image enhancement processing can also be applied to display screens that allow the examiner to confirm the success or failure of an image immediately after it has been taken, such as a shooting confirmation screen.

[0427] Unless otherwise specified, the configuration and processing of the image processing apparatus according to this embodiment are the same as those of the image processing apparatus 400 according to the first embodiment. Therefore, the following description of the image processing apparatus according to this embodiment will focus on the differences from the image processing apparatus according to the first embodiment.

[0428] The output unit 405 can display multiple high-resolution images generated by the image enhancement unit 404, as well as low-resolution images that have not undergone enhancement, on the display unit 20. This allows for the output of low-resolution and high-resolution images, respectively, according to the examiner's instructions.

[0429] An example of the interface 3400 is shown below with reference to Figure 34. 3400 represents the entire screen, 3401 represents the patient tab, 3402 represents the imaging tab, 3403 represents the report tab, and 3404 represents the settings tab. The diagonal lines in the report tab (3403) indicate the active state of the report screen. In this embodiment, an example of displaying the report screen will be described. Im3405 is an SLO image, and Im3406 displays the OCTA En-Face image shown in Im3407 superimposed on the SLO image Im3405. Here, an SLO image is a frontal image of the fundus acquired by an SLO (Scanning Laser Ophthalmoscope) optical system (not shown). Im3407 and Im3408 are OCTA En-Face images, Im3409 is a luminance En-Face image, and Im3411 and Im3412 are tomographic images. Images 3413 and 3414 superimpose the upper and lower boundary lines of the OCTA En-Face images, as shown in Im3407 and Im3408, onto the tomographic image. Button 3420 is used to specify the execution of high-resolution processing. Of course, as will be described later, button 3420 may also be used to instruct the display of the high-resolution image.

[0430] In this embodiment, the image enhancement process is performed by specifying button 3420, or by determining whether or not to perform it based on information stored in the database. First, an example of switching between displaying high-resolution and low-resolution images by specifying button 3420 in response to instructions from the examiner will be described. The target images for the image enhancement process will be described as OCTA's En-Face images. When the examiner selects the report tab 3403 and transitions to the report screen, low-resolution OCTA En-Face images Im3407 and Im3408 are displayed. Subsequently, when the examiner selects button 3420, the image enhancement unit 404 performs image enhancement on the images Im3407 and Im3408 displayed on the screen. After the image enhancement process is completed, the output unit 405 displays the high-resolution images generated by the image enhancement unit 404 on the report screen. Furthermore, since Im3406 displays Im3407 superimposed on the SLO image Im3405, Im3406 also displays an image that has undergone high-resolution processing. The display of button 3420 is then changed to an active state, indicating that high-resolution processing has been performed. Here, the execution of processing in the high-resolution unit 404 does not need to be limited to the timing when the examiner specifies button 3420. Since the types of OCTA En-Face images Im3407 and Im3408 to be displayed when opening the report screen are known in advance, high-resolution processing may be performed when transitioning to the report screen. Also, the output unit 405 may be configured to display the high-resolution image on the report screen at the timing when button 3420 is pressed. Moreover, the types of images for which high-resolution processing is performed in response to instructions from the examiner or when transitioning to the report screen do not need to be limited to two types. The system may process multiple OCTA En-Face images, such as the surface (Im2910), deep (Im2920), outer layer (Im2930), and choroidal vascular network (Im2940) images shown in Figure 29, which are likely to be displayed. In this case, the images obtained after the image enhancement processing may be temporarily stored in memory or in a database.

[0431] Next, we will explain the case where image enhancement processing is performed based on information stored in the database. If the database has saved a state where image enhancement processing is to be performed, the high-resolution image obtained by performing the processing will be displayed by default when transitioning to the report screen. Furthermore, by displaying button 3420 as active by default, the examiner can be informed that the high-resolution image obtained by performing the image enhancement processing is being displayed. If the examiner wants to display the low-resolution image before the enhancement processing, they can deactivate button 3420 to display the low-resolution image. To return to the high-resolution image, the examiner selects button 3420. Whether or not to perform image enhancement processing on the database shall be specified hierarchically, either for all data stored in the database or for each image data (each examination). For example, if the state for performing image enhancement processing is saved for the entire database, and the examiner saves a state not to perform image enhancement processing for individual image data (individual examinations), the next time that image data is displayed, it will be displayed without image enhancement processing. To save the execution state of image enhancement processing for each image data (each examination), an unillustrated user interface (e.g., a save button) may be used. In addition, when transitioning to other image data (other examinations) or other patient data (e.g., changing to a display screen other than the report screen in response to instructions from the examiner), the state for performing image enhancement processing may be saved based on the display state (e.g., the state of button 3420). This allows processing to be performed based on the information specified for the entire database when the execution of image enhancement processing is not specified for each image data (each examination), and to be performed individually based on the information specified for each image data (each examination).

[0432] In this embodiment, examples are shown where Im3407 and Im3408 are displayed as OCTA En-Face images. However, the displayed OCTA En-Face image can be changed by the examiner. Therefore, the image change when high-image-quality processing is specified (button 3420 is active) will be explained.

[0433] Image changes are made using a user interface (e.g., a combo box) not shown. For example, when the examiner changes the image type from the superficial layer to the choroidal vascular network, the image enhancement unit 404 performs image enhancement processing on the choroidal vascular network image, and the output unit 405 displays the high-resolution image generated by the image enhancement unit 404 on the report screen. That is, the output unit 405 may change the display of the high-resolution image of the first depth range to the display of the high-resolution image of the second depth range, which is at least partially different from the first depth range, in response to instructions from the examiner. In this case, the output unit 405 may change the display of the high-resolution image of the first depth range to the display of the high-resolution image of the second depth range, as the first depth range is changed to the second depth range in response to instructions from the examiner. As mentioned above, for images that are likely to be displayed when transitioning to the report screen, if a high-resolution image has already been generated, the output unit 405 may display the generated high-resolution image. The method for changing the image type is not limited to those described above. It is also possible to generate OCTA En-Face images with different depth ranges by changing the reference layer and offset values. In this case, when the reference layer or offset value is changed, the image enhancement unit 404 performs image enhancement processing on the arbitrary OCTA En-Face image, and the output unit 405 displays the high-quality image on the report screen. The reference layer and offset value can be changed using a user interface (not shown) (e.g., a combo box or text box). Furthermore, the generation range of the OCTA En-Face image can be changed by dragging either boundary line 3413 or 3414 superimposed on the tomographic images Im3411 and Im3412 (moving the layer boundary). When changing the boundary line by dragging, the execution command for image enhancement processing is executed continuously. Therefore, the image enhancement unit 404 may always perform processing in response to the execution command, or it may perform processing only after the layer boundary has been changed by dragging. Alternatively, the image enhancement process can be executed sequentially, but the previous command can be canceled and the latest command executed when the next command arrives. Note that the image enhancement process may take a relatively long time. Therefore, regardless of when the command is executed as described above, it may take a relatively long time for the high-resolution image to be displayed. Thus, from the time the depth range for generating the OCTA En-Face image is set in response to the examiner's instructions until the high-resolution image is displayed, the OCTA En-Face image (low-resolution image) corresponding to the set depth range may be displayed. That is, when the depth range is set, the OCTA En-Face image (low-resolution image) corresponding to the set depth range is displayed, and when the high-resolution processing is completed, the display of the OCTA En-Face image (low-resolution image) is changed to the display of the high-resolution image. Furthermore, information indicating that the high-resolution processing is being performed may be displayed from the time the depth range is set until the high-resolution image is displayed. These measures can be applied not only when the execution of the high-resolution processing has already been specified (button 3420 is active), but also, for example, when the execution of the high-resolution processing is instructed in response to the examiner's instructions, until the high-resolution image is displayed.

[0434] In this embodiment, we have shown an example in which different layers are displayed on Im3407 and Im3408 as OCTA En-Face images, and low-resolution and high-resolution images are displayed in a switchable manner, but this is not limited to this. For example, a low-resolution OCTA En-Face image may be displayed side by side on Im3407, and a high-resolution OCTA En-Face image may be displayed side by side on Im3408. When images are switched, it is easy to compare the parts that have changed because the images are switched at the same location, and when images are displayed side by side, it is easy to compare the entire image because the images can be displayed simultaneously.

[0435] Next, we will explain the execution of image enhancement processing during screen transitions using Figures 34(a) and (b). Figure 34(b) is an example screen showing an enlarged view of the OCTA En-Face image Im3407 in Figure 34(a). In Figure 34(b), the button 3420 is displayed in the same way as in Figure 34(a). The screen transition from Figure 34(a) to Figure 34(b) is performed, for example, by double-clicking the OCTA En-Face image Im3407, and the transition from Figure 34(b) to Figure 34(a) is performed by clicking the close button 3430. Note that the method of screen transition is not limited to the method shown here, and a user interface not shown may also be used. If the execution of image enhancement processing is specified during a screen transition (button 3420 is active), that state is maintained even during the screen transition. That is, if a high-resolution image is displayed on the screen of Figure 34(a) and the screen transitions to the screen of Figure 34(b), the high-resolution image will also be displayed on the screen of Figure 34(b). Then, button 3420 is activated. The same applies when transitioning from Figure 34(b) to Figure 34(a). In Figure 34(b), it is also possible to switch the display to a low-resolution image by specifying button 3420. Regarding screen transitions, not limited to the screens shown here, if the transition is to a screen that displays the same shooting data, such as a progress observation screen or a panoramic image screen, the transition will be performed while maintaining the display state of the high-resolution image. In other words, the image displayed on the screen after the transition corresponds to the state of button 3420 on the screen before the transition. For example, if button 3420 is active on the screen before the transition, a high-resolution image will be displayed on the screen after the transition. Also, for example, if button 3420 is deactivated on the screen before the transition, a low-resolution image will be displayed on the screen after the transition. Furthermore, when button 3420 on the observation screen becomes active, multiple images obtained at different dates (different examination dates) displayed side by side on the observation screen may be switched to a high-resolution image. In other words, when button 3420 on the progress observation display screen becomes active, it may be configured to apply the results to multiple images obtained at different dates and times simultaneously. An example of the progress observation display screen is shown in Figure 38. When tab 3801 is selected in response to instructions from the examiner, the progress observation display screen is displayed as shown in Figure 38. At this time, the depth range of the En-Face image can be changed by the examiner selecting from the default depth range sets (3802 and 3803) displayed in the list box. For example, the retinal surface is selected in list box 3802, and the retinal depth is selected in list box 3803. The upper display area shows the analysis results of the En-Face image of the retinal surface, and the lower display area shows the analysis results of the En-Face image of the retinal depth. In other words, when a depth range is selected, the display is changed simultaneously to a parallel display of the analysis results of multiple En-Face images within the selected depth range for multiple images taken at different dates and times. At this time, if the display of analysis results is deselected, the display may be changed simultaneously to a parallel display of multiple En-Face images taken at different dates and times. Then, when button 3420 is selected in response to instructions from the examiner, the display of multiple En-Face images is changed to the display of multiple high-resolution images all at once. Also, if the display of analysis results is selected, when button 3420 is selected in response to instructions from the examiner, the display of analysis results for multiple En-Face images is changed to the display of analysis results for multiple high-resolution images all at once. Here, the display of analysis results may be a display of the analysis results superimposed on the image with an arbitrary transparency. In this case, the change to display of analysis results may be, for example, a change to a state in which the analysis results are superimposed on the displayed image with an arbitrary transparency. Alternatively, the change to display of analysis results may be, for example, a change to display of an image obtained by blending the analysis results and the image with an arbitrary transparency (e.g., a 2D map). Furthermore, the type of layer boundary and the offset position used to specify the depth range can be changed all at once from a user interface such as 3805 and 3806, respectively.Furthermore, the depth range of multiple En-Face images from different dates and times may be changed simultaneously by displaying tomographic images together and moving the layer boundary data superimposed on the tomographic images according to the examiner's instructions. In this case, multiple tomographic images from different dates and times may be displayed side by side, and if the above movement is performed on one tomographic image, the layer boundary data may be similarly moved on the other tomographic images. In addition, the image projection method and the presence or absence of projection artifact suppression processing may be changed by selecting them from a user interface such as a context menu. Alternatively, the selection button 3807 may be selected to display a selection screen, and the image selected from the image list displayed on the selection screen may be displayed. Note that the arrow 3804 displayed at the top of Figure 38 indicates that the currently selected examination is the one being performed, and the baseline examination is the examination selected during follow-up imaging (the leftmost image in Figure 38). Of course, a mark indicating the baseline examination may also be displayed on the display unit. Furthermore, if the "Show Difference" checkbox 3808 is selected, the measured value distribution (map or sector map) relative to the baseline image will be displayed on the baseline image. Furthermore, in this case, a difference measurement value map is displayed in the area corresponding to the other inspection days, showing the difference between the measurement value distribution calculated for the reference image and the measurement value distribution calculated for the image displayed in that area. As a measurement result, a trend graph (a graph of measurement values ​​for the image on each inspection day obtained by measuring changes over time) may also be displayed on the report screen. In other words, time-series data (e.g., a time-series graph) of multiple analysis results corresponding to multiple images taken at different dates and times may be displayed. In this case, analysis results for dates and times other than those corresponding to the displayed multiple images may also be displayed as time-series data in a manner that makes them distinguishable from the analysis results corresponding to the displayed multiple images (e.g., the color of each point on the time-series graph differs depending on whether an image is displayed or not). Furthermore, the regression line (curve) of the trend graph and the corresponding mathematical formula may be displayed on the report screen.

[0436] In this embodiment, the description has focused on OCTA En-Face images, but is not limited to these. The images used for display, image enhancement, and image analysis in this embodiment may also be luminance En-Face images. Furthermore, the images used may not be limited to En-Face images, but may be tomographic images, SLO images, fundus photographs, or fluorescein photographs, or other types of images. In that case, the user interface for performing image enhancement processing may include one that instructs the user to perform image enhancement processing on multiple images of different types, or one that allows the user to select any image from multiple images of different types and instructs the user to perform image enhancement processing.

[0437] With this configuration, the output unit 405 can display the image processed by the image enhancement unit 404 according to this embodiment on the display unit 20. At this time, as described above, if at least one of the multiple conditions related to the display of high-resolution images, the display of analysis results, the depth range of the displayed front image, etc. is selected, the selected state may be maintained even if the display screen is transitioned. Also, as described above, if at least one of the multiple conditions is selected, the state in which at least one is selected may be maintained even if the other conditions are changed to a selected state. For example, if the display of analysis results is selected, the output unit 405 may change the display of the analysis results of low-resolution images to the display of the analysis results of high-resolution images in response to instructions from the examiner (for example, when button 3420 is specified). Also, if the display of analysis results is selected, the output unit 405 may change the display of the analysis results of high-resolution images to the display of the analysis results of low-resolution images in response to instructions from the examiner (for example, when button 3420 is deselected). Furthermore, if the display of high-resolution images is not selected, the output unit 405 may change the display of the analysis results for low-resolution images to the display of low-resolution images in response to instructions from the examiner (for example, when the specification for displaying analysis results is canceled). Furthermore, if the display of high-resolution images is not selected, the output unit 405 may change the display of the analysis results for low-resolution images in response to instructions from the examiner (for example, when the display of analysis results is specified). Furthermore, if the display of high-resolution images is selected, the output unit 405 may change the display of the analysis results for high-resolution images to the display of high-resolution images in response to instructions from the examiner (for example, when the specification for displaying analysis results is canceled). Furthermore, if the display of high-resolution images is selected, the output unit 405 may change the display of the analysis results for high-resolution images in response to instructions from the examiner (for example, when the display of analysis results is specified). Let's also consider the case where the display of high-resolution images is not selected and the display of the first type of analysis result is selected. In this case, the output unit 405 may change the display of the first type of analysis result for the low-resolution image to the display of the second type of analysis result for the low-resolution image, in response to instructions from the examiner (for example, if the display of the second type of analysis result is specified).Furthermore, consider the case where the display of high-resolution images is selected and the display of the first type of analysis results is selected. In this case, the output unit 405 may change the display of the first type of analysis results for high-resolution images to the display of the second type of analysis results for high-resolution images in response to instructions from the examiner (for example, when the display of the second type of analysis results is specified). Note that, as described above, the display screen for monitoring progress may be configured to reflect these display changes collectively for multiple images obtained at different dates and times. Here, the display of analysis results may be a display of the analysis results superimposed on the image with an arbitrary degree of transparency. In this case, the change to displaying analysis results may be, for example, a change to a state in which the analysis results are superimposed on the ...

Claims

1. An acquisition unit acquires a first image, which is a medical image of a predetermined part of the subject, The system includes a display control unit that controls the display unit to display either a second image obtained by inputting the first image as an input image to the trained model, or the first image. In a first display screen that displays either the first image or the second image, and a second display screen that displays either the first image or the second image, the display screen is changed from one display screen to the other in response to instructions from the examiner, If the first image is displayed on one of the display screens, the modification is made so that the first image is displayed on the other display screen, and A medical image processing apparatus, wherein if the second image is displayed on one of the display screens, the modification is made so that the second image is displayed on the other display screen.

2. A user interface is displayed that allows the user to specify whether to display the first image or the second image in response to an instruction from the examiner; if the first image and a user interface indicating that the first image has been selected are displayed on one of the display screens, the modification is made so that the first image and a user interface indicating that the first image has been selected are displayed on the other display screen; and if the second image and a user interface indicating that the second image has been selected are displayed on one of the display screens, the modification is made so that the second image and a user interface indicating that the second image has been selected are displayed on the other display screen.

3. The medical image processing apparatus according to claim 2, wherein the display control unit controls the display unit to display either the analysis result of the first image or the analysis result of the second image, and controls the display unit to change from the display of one of the analysis results of the first image or the analysis result of the second image to the display of the other in response to instructions from an examiner via the displayed user interface, and when the analysis result of the first image is displayed on one of the display screens, the change is made so that the analysis result of the first image is displayed on the other display screen, and when the analysis result of the second image is displayed on one of the display screens, the change is made so that the analysis result of the second image is displayed on the other display screen.

4. The display control unit controls the display unit so that a plurality of medical images corresponding to a plurality of different depth ranges of the predetermined part in the three-dimensional medical image data of the predetermined part are displayed side by side on the first display screen as one of the first image and the second image; the display control unit controls the display unit so that the display of the plurality of medical images is simultaneously changed from one of the first image and the second image to the other image in response to instructions from the examiner via the displayed user interface; and when one of the plurality of medical images is selected in response to instructions from the examiner, the display unit controls the display unit so that the display screen is changed from the first display screen to the second display screen in which the selected medical image is displayed enlarged.

5. The medical image processing apparatus according to any one of claims 1 to 4, wherein the second image displayed on the first display screen and the second display screen is a composite image obtained by combining the pixel values ​​of each corresponding pixel in the first image and the second image at a ratio obtained using information about at least a portion of a region in at least one of the first image and the second image, or at a ratio that can be changed according to instructions from the examiner.

6. The medical image processing apparatus according to any one of claims 1 to 5, wherein the trained model is a trained model obtained using training data in which noise is added to at least a portion of the medical image.

7. The medical image processing apparatus according to any one of claims 1 to 6, wherein the trained model is a trained model obtained using training data in which noise corresponding to the state of at least a portion of the medical image is added to the at least portion of the medical image.

8. The medical image processing apparatus according to any one of claims 1 to 7, wherein the trained model is a trained model obtained using training data in which noise of a magnitude corresponding to the pixel values ​​of at least a portion of the medical image is added to the at least portion of the region.

9. The medical image processing apparatus according to any one of claims 1 to 8, wherein the trained model is a trained model obtained using training data that includes multiple medical images as paired images to which different patterns of noise have been added.

10. The medical image processing apparatus according to any one of claims 1 to 9, wherein the trained model is a trained model obtained using training data that includes a plurality of medical images obtained as paired images by adding noise of different patterns to medical images obtained by superimposing processing.

11. The system further includes a designation means for specifying a portion of the depth range of the predetermined part in the three-dimensional medical image data of the predetermined part, in response to instructions from the examiner. The acquisition unit acquires a medical image corresponding to the specified depth range as the first image. The medical image processing apparatus according to any one of claims 1 to 10, wherein the trained model is a trained model obtained using training data that includes multiple medical images corresponding to multiple depth ranges of a predetermined part of a subject.

12. The medical image processing apparatus according to claim 11, wherein the trained model is a trained model obtained using training data in which noise of a magnitude corresponding to information about the pixel values ​​in at least a portion of each of the multiple medical images corresponding to the multiple depth ranges is added to at least a portion of each of the multiple depth ranges.

13. The medical image processing apparatus according to claim 11 or 12, further comprising a generation unit that generates a wide-angle image using a plurality of second images obtained from a plurality of first images, wherein a plurality of adjacent medical images corresponding to a specified depth range overlap by photographing different positions of the predetermined part in a direction intersecting the depth direction of the predetermined part, and the first image apparatus according to claim 11 or 12, wherein a plurality of first images are obtained by photographing different positions of the predetermined part in a direction intersecting the depth direction of the predetermined part.

14. A medical image processing apparatus according to any one of claims 1 to 13, wherein the first image is divided into a plurality of two-dimensional images and input into the trained model, and the second image is generated by integrating the plurality of output images from the trained model.

15. The aforementioned trained model is a trained model obtained using training data that includes multiple medical images as paired images, where the relative positions of each image correspond. The medical image processing apparatus according to claim 14, wherein the first image is divided into a plurality of two-dimensional images with an image size corresponding to the image size of the paired images and input to the trained model.

16. The medical image processing apparatus according to claim 14 or 15, wherein the trained model is a trained model obtained using training data that includes images of a plurality of subregions set such that parts of adjacent subregions overlap with each other, for a region including a medical image and the surrounding area outside the medical image.

17. The medical image processing apparatus according to any one of claims 1 to 16, wherein the trained model is a trained model obtained using training data that includes an image obtained by octaurating an image using an OCT imaging apparatus that is more advanced than the OCT imaging apparatus used for octaurating the first image, or an image obtained by an OCT imaging process that has more steps than the octaurating process for the first image.

18. The medical image processing apparatus according to any one of claims 1 to 17, wherein the trained model is a trained model obtained using training data including medical images obtained by superimposing.

19. The medical image processing apparatus according to any one of claims 1 to 18, wherein the medical image is a frontal image of OCTA or a tomographic image of OCT.

20. The medical image processing apparatus according to any one of claims 1 to 19, wherein the second image relating to a partial image of a predetermined area that has not yet been selected by an examiner is acquired before the input of an instruction to select the partial image is received.

21. A method for operating a medical image processing device, The medical image processing device acquires a first image, which is a medical image of a predetermined part of the subject, The medical image processing device includes controlling the display unit to display either a second image obtained by inputting the first image into the trained model as an input image for the trained model, or the first image. In a first display screen that displays either the first image or the second image, and a second display screen that displays either the first image or the second image, the display screen is changed from one display screen to the other in response to instructions from the examiner, If the first image is displayed on one of the display screens, the modification is made so that the first image is displayed on the other display screen, and A method for operating a medical image processing device, wherein, when the second image is displayed on one of the display screens, the modification is made so that the second image is displayed on the other display screen.

22. A program that, when executed by a processor, causes the processor to perform each step of the operating method of the medical image processing apparatus described in claim 21.

Citation Information

Patent Citations

  • Ophthalmologic image display device and ophthalmologic imaging apparatus

    JP2016198447A

  • Information processing device

    JP2017158962A

  • Medical image processing method and medical image processing device

    JP2018005841A

  • Denoising medical images by learning sparse image representations with a deep unfolding approach

    US20180240219A1