Image processing method, machine learning method, image processing device, image processing program, non-volatile recording medium with image processing program recorded thereon, and computer program product with image processing program recorded thereon

WO2026159805A1PCT designated stage Publication Date: 2026-07-30OLYMPUS MEDICAL SYST CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
OLYMPUS MEDICAL SYST CORP
Filing Date
2025-01-22
Publication Date
2026-07-30

Smart Images

  • Figure JP2025001899_30072026_PF_FP_ABST
    Figure JP2025001899_30072026_PF_FP_ABST
Patent Text Reader

Abstract

In this image processing method, an image characteristics calculation unit calculates a first image characteristic, which is an image characteristic in a first region of a teacher image, and calculates a second image characteristic, which is an image characteristic in a second region of the teacher image, the second image being different from the first region. A low-pass filter selection unit selects a first low-pass filter corresponding to the first image characteristic and a second low-pass filter corresponding to the second image characteristic from a low-pass filter group comprising a plurality of types of low-pass filters. A low-pass filter processing unit generates a student image for machine learning from the teacher image by performing image processing on the teacher image such that the first low-pass filter is applied in the first region and by performing image processing on the teacher image such that the second low-pass filter is applied in the second region.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method, machine learning method, image processing apparatus, image processing program, non-volatile recording medium recording the image processing program, and computer program product recording the image processing program

[0001] The present invention relates to an image processing method, a machine learning method, an image processing apparatus, an image processing program, a non-volatile recording medium recording the image processing program, and a computer program product recording the image processing program, which are suitable for machine learning of AI for high image quality processing.

[0002] In recent years, systems for realizing high image quality processing using AI (artificial intelligence) have been developed. For example, there is an AI high image quality system having an inference model that obtains a high image quality image from an input low image quality image by performing deep learning on a dataset of low image quality student images and high image quality teacher images of the same subject with the same configuration. Also in the medical field, an AI high image quality system may be adopted to improve the image quality of endoscopic images.

[0003] It is difficult to obtain low image quality student images and high image quality teacher images of the same subject with the same configuration by endoscopic imaging. Therefore, a technique of generating a low image quality image (student image) by performing degradation processing on a high image quality image (teacher image) obtained by imaging with a high image quality endoscope capable of high quality imaging may be adopted. The degradation processing is performed based on the optical system and image sensor information of the high image quality endoscope and the optical system and image sensor information of the low image quality endoscope that captures a low resolution image during actual use.

[0004] For example, in International Publication No. 2021 / 090469, a target endoscopic image is used as a teacher image, and the entire teacher image is convolved with one correction PSF (Point Spread Function), blurred to generate a student image, and a technique of generating an inference model by deep learning of a learning image set of this teacher image and student image is disclosed. By applying the inference model to the image of the target endoscope, the image quality of the endoscopic image is improved by deep learning and inference.

[0005] Furthermore, Japanese Patent Publication No. 2010-066925 discloses an invention for improving the image quality of an image sensor by applying different filters according to the pixel position to generate student images from teacher images.

[0006] International Publication No. 2021 / 090469, Japanese Patent Publication No. 2010-066925

[0007] However, the technology described in the patent document cannot add scene-adaptive blurring when generating student images from training images. Therefore, even if an inference model obtained by training these training and student images is used, scene-adaptive image quality enhancement cannot be achieved. Furthermore, it is difficult to create a filter with ideal frequency response characteristics across the entire bandwidth as a filter for generating low-quality student images from training images in the first place.

[0008] The present invention aims to provide an image processing method, a machine learning method, an image processing apparatus, an image processing program, a non-volatile recording medium storing the image processing program, and a computer program product storing the image processing program, which enable scene-adaptive high-quality image processing by using an inference model constructed with student images created by giving each localized training image a scene-adaptively varying degree of blur.

[0009] An image processing method according to one aspect of the present invention involves an image characteristic calculation unit calculating a first image characteristic which is the image characteristic of a first region of a training image, and calculating a second image characteristic which is the image characteristic of a second region of the training image different from the first region, a low-pass filter selection unit selecting a first low-pass filter corresponding to the first image characteristic and a second low-pass filter corresponding to the second image characteristic from a group of low-pass filters consisting of a plurality of types of low-pass filters, and a low-pass filter processing unit performing image processing on the first region of the training image using the first low-pass filter and on the second region of the training image using the second low-pass filter, thereby generating student images for machine learning from the training image.

[0010] A machine learning method according to one aspect of the present invention involves feeding the training image and the student image generated by the above image processing method to a neural network for training to obtain an inference model for image quality enhancement processing that converts a low-resolution input image into a high-resolution image.

[0011] An image processing device according to one aspect of the present invention includes artificial intelligence trained by the above-described machine learning method and an image input unit that inputs received image information to the artificial intelligence.

[0012] In another aspect of the present invention, an image processing method is provided in which at least one processor calculates a first image characteristic, which is an image characteristic of a first region of a training image; calculates a second image characteristic, which is an image characteristic of a second region of the training image different from the first region; selects a first low-pass filter corresponding to the first image characteristic from a group of low-pass filters consisting of a plurality of types of low-pass filters; selects a second low-pass filter corresponding to the second image characteristic; performs image processing on the first region of the training image using the first low-pass filter; and performs image processing on the second region of the training image using the second low-pass filter, thereby generating student images for machine learning from the training image.

[0013] An image processing program according to one aspect of the present invention causes at least one processor to calculate a first image characteristic, which is the image characteristic of a first region of a training image; to calculate a second image characteristic, which is the image characteristic of a second region of the training image that is different from the first region; to select a first low-pass filter corresponding to the first image characteristic from a group of low-pass filters consisting of a plurality of types of low-pass filters; to select a second low-pass filter corresponding to the second image characteristic; to perform image processing on the first region of the training image using the first low-pass filter; to perform image processing on the second region of the training image using the second low-pass filter; and to generate student images for machine learning from the training image.

[0014] A non-volatile recording medium storing an image processing program according to one aspect of the present invention contains an image processing program which causes at least one processor to calculate a first image characteristic, which is the image characteristic of a first region of a training image; calculate a second image characteristic, which is the image characteristic of a second region of the training image that is different from the first region; select a first low-pass filter corresponding to the first image characteristic from a group of low-pass filters consisting of a plurality of types of low-pass filters; select a second low-pass filter corresponding to the second image characteristic; perform image processing on the first region of the training image using the first low-pass filter; and perform image processing on the second region of the training image using the second low-pass filter, thereby generating student images for machine learning from the training image.

[0015] A computer program product recording an image processing program according to one aspect of the present invention records an image processing program that causes at least one computer to calculate a first image characteristic, which is the image characteristic of a first region of a training image; calculate a second image characteristic, which is the image characteristic of a second region of the training image that is different from the first region; select a first low-pass filter corresponding to the first image characteristic from a group of low-pass filters consisting of multiple types of low-pass filters; select a second low-pass filter corresponding to the second image characteristic; perform image processing on the first region of the training image using the first low-pass filter; and perform image processing on the second region of the training image using the second low-pass filter; thereby generating student images for machine learning from the training image.

[0016] According to the present invention, an inference model constructed using student images created by applying scene-adaptively varying degrees of blurring to localized areas of the training image enables scene-adaptively improving image quality.

[0017] This is a block diagram of an image processing apparatus according to the first embodiment of the present invention. This is a graph showing the PSF of the target endoscope and the image endoscope, with the frequency band of the image on the horizontal axis and the contrast on the vertical axis. This is a flowchart for explaining the operation of the first embodiment. This is an explanatory diagram for explaining the change in the characteristic value of contrast based on the difference in images. This is a graph showing an example of the corrected PSF of the LPF based on the difference in contrast, with the frequency band on the horizontal axis and the contrast on the vertical axis. This is a graph showing the frequency band of the region of interest, with the frequency band on the horizontal axis and the contrast on the vertical axis. This is an explanatory diagram showing the characteristics A11 to A13 of three filters as examples of filters generated by the filter generation unit 5. This is an explanatory diagram showing an example of an endoscope image. This is an explanatory diagram showing the characteristics A21 to A23 of three filters as examples of filters generated by the filter generation unit 5. This is a block diagram of the second embodiment of the present invention. This is an explanatory diagram for explaining the target frequency response characteristics obtained by the target frequency response characteristic setting unit 4. This is a graph for explaining filter generation, with the frequency band on the horizontal axis and the contrast on the vertical axis. This is a graph for explaining filter generation, with the frequency band on the horizontal axis and the contrast on the vertical axis. This is a block diagram of an AI image enhancement system using teacher images and student images.

[0018] Embodiments of the present invention will be described in detail below with reference to the drawings.

[0019] (First Embodiment) Figure 1 is a block diagram showing an image processing apparatus according to the first embodiment of the present invention. The image processing apparatus in Figure 1 realizes the image processing method of the embodiment. In this method, when generating a low-resolution student image by applying appropriate blurring to a high-resolution teacher image, a filter is applied that applies different amounts of blurring (different blur differences or different strengths of blur) according to the image characteristics of the target pixel and its surrounding area in the teacher image, such as brightness, color, contrast, frequency, edge, noise, subject distance, etc., to generate the target pixel of the student image. In other words, in this embodiment, a filter with different blur differences can be used according to the scene characteristics of the teacher image, and it is possible to generate a student image that has a scene-adaptive difference in the degree of blurring from the teacher image (different blur differences), thereby enabling the construction of an inference model that achieves scene-adaptive high image quality.

[0020] In this embodiment, we will describe an example in which a high-resolution endoscope (hereinafter referred to as the target endoscope) is used as the imaging device to obtain the target high-resolution image (hereinafter referred to as the target imaging device), and a low-resolution endoscope (hereinafter referred to as the target endoscope), which has relatively lower image quality compared to the high-resolution endoscope, is used as the imaging device to obtain the low-resolution image to be used for inference (hereinafter referred to as the target imaging device). However, various imaging devices other than endoscopes can be used as imaging devices to obtain high-resolution and low-resolution images. Furthermore, the images to be used for inference are not limited to images of the inside of a body.

[0021] Figure 2 is a graph showing the PSF of the target endoscope and the sample-scale field (PSF) of the target endoscope, with the frequency band of the image on the horizontal axis and the contrast on the vertical axis. Now, assume that the PSF of the target endoscope and the PSF of the sample-scale field (PSF) of the sample-scale field have the characteristics shown in Figure 2.

[0022] Various low-pass filters (LPFs) can be used as filters to blur images. As a method for calculating the frequency response characteristics (hereinafter referred to as corrected PSF) of such an LPF, a technique is sometimes employed in which the corrected PSF is obtained from the PSF of the target endoscope, and this corrected PSF is approximated to the PSF of the target endoscope / PSF of the target endoscope (hereinafter referred to as the ideal frequency response characteristics). By applying a filter with a corrected PSF that matches the ideal frequency response characteristics to a high-resolution image obtained with the target endoscope, it is possible to reliably generate a low-resolution image that would be obtained with the target endoscope. However, as shown in Figure 2, the corrected PSF does not perfectly match the ideal frequency response characteristics; therefore, an approximation is performed such that the characteristics roughly match at, for example, half the Nyquist frequency.

[0023] Alternatively, a Gaussian filter may be used as the low-pass filter (LPF) for blurring. A Gaussian filter uses weighting based on a Gaussian function to smoothly change pixel values, enabling a natural blurring effect. By applying the filter to the high-resolution image obtained with the target endoscope, a low-resolution image that would likely be obtained with the target endoscope is generated.

[0024] Thus, conventionally, low-resolution student images are generated by applying corrected PSF to high-resolution images (teacher images), and the teacher image is given a uniform blur corresponding to the corrected PSF when the student images are generated. In other words, conventionally, the blur processing is not adapted to the scene.

[0025] Therefore, in this embodiment, it is possible to generate student images with different blur differences depending on the image characteristics of the target pixel and its surrounding area in the training image. This makes it possible to construct an inference model that achieves scene-adaptive high image quality. In this embodiment, in order to generate student images with different blur differences adaptively to the scene, different correction PSF filters are switched according to the image characteristics, and good student images adapted to the scene can be obtained across the entire bandwidth.

[0026] (Configuration) The image processing device 10 shown in Figure 1 includes a reduction processing unit 1, an image characteristic calculation unit 2, an LPF selection unit 3, a target frequency response characteristic setting unit 4, a filter generation unit 5, and an LPF processing unit 6. Each part of the image processing device 10 may be composed of a processor using a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), an NPU (Neural Processing Unit), etc., and may operate according to a program stored in a memory (not shown) to control each part, or may some or all of the functions be realized by hardware electronic circuits.

[0027] The image processing device 10 receives a high-resolution image obtained by a high-resolution imaging device as a training image PL, and generates a student image with a quality equivalent to the quality of a low-resolution image obtained by a target imaging device through image processing of the high-resolution image. By using the high-resolution image as the training image and providing a dataset of training images and student images to a neural network for training, it becomes possible to construct an inference model for image quality enhancement processing that converts low-resolution images into high-resolution images.

[0028] The training image PL is supplied to the reduction processing unit 1. The reduction processing unit 1 performs a reduction process on the input training image PL. For example, the reduction processing unit 1 performs the reduction process by known interpolation or decimation processes. The reduction processing by the reduction processing unit 1 generates a training image PT. The generated training image PT is supplied to the image characteristic calculation unit 2 and the LPF processing unit 6.

[0029] By reducing the size of the training image through image reduction processing, the number of pixels relative to the number of line pairs decreases, thus improving the resolution of the training image and enhancing its frequency characteristics. However, if a sufficiently high-quality image can be obtained using a high-resolution imaging device, image reduction processing by the reduction processing unit 1 is not always necessary, and the reduction processing unit 1 can be omitted. In this case, the training image PT is input as the learning image PL, and the input training image PT is supplied to the image characteristic calculation unit 2 and the LPF processing unit 6.

[0030] The image characteristic calculation unit 2 calculates local image characteristics of the training image PT. For example, the image characteristic calculation unit 2 divides the entire area of ​​the training image PT according to predetermined conditions and calculates image characteristics for the divided area (hereinafter referred to as the area of ​​interest). For example, the image characteristic calculation unit 2 uses the pixel of interest and its surrounding area as the area of ​​interest and calculates at least one image characteristic selected from the group consisting of, for example, brightness, color, frequency band distribution, frequency characteristic distribution, edge, contrast distribution, or subject distance. Note that luminance, brightness, etc. may be used as a measure of brightness, and hue, saturation, etc. may be used as a measure of color.

[0031] The image characteristic calculation unit 2 calculates image characteristics for each region of interest corresponding to a pixel of interest across the entire area of ​​the training image PT, and outputs the calculated image characteristic information to the LPF selection unit 3. Note that the image characteristic calculation unit 2 does not necessarily need to calculate image characteristics for all regions of interest corresponding to pixels of interest across the entire area of ​​the training image PT; it may calculate image characteristics for multiple regions of interest within the training image PT.

[0032] The target frequency response characteristic setting unit 4 sets the target frequency response characteristics of the filter required when blurring the teacher image PT to create the student image PS in the filter generation unit 5. The target frequency response characteristics are a setting that determines how much the high-resolution image will be blurred, and are set empirically based on the resolution of the high-resolution image and the resolution of the low-resolution endoscopic image obtained by the target endoscope.

[0033] In this embodiment, the filter generation unit 5 generates a group of low-pass filters (LPFs) consisting of multiple types of low-pass filters (LPFs) that approximate the target frequency response characteristics. The multiple filters generated by the filter generation unit 5 have different frequency response characteristics (corrected PSFs) from each other, and are capable of applying different blurring differences to the training image PT.

[0034] The LPF selection unit 3 outputs selection information to the LPF processing unit 6 for selecting a filter from a group of low-pass filters according to the image characteristics for each area of ​​interest in the teacher image PT. The LPF processing unit 6 generates each pixel that makes up the student image PS. Based on the selection information provided by the LPF selection unit 3, the LPF processing unit 6 selects one of the multiple filters generated by the filter generation unit 5 and applies a blurring effect using the selected filter to the area of ​​interest in the teacher image PT, thereby generating pixels in the student image PS corresponding to the pixels of interest in the teacher image PT (hereinafter referred to as target pixels of interest). The student image PS is generated by the LPF processing unit 6 generating target pixels of interest across the entire area of ​​the student image PS.

[0035] The selection information generated by the LPF selection unit 3 may indicate which stage the image is in, by dividing the image characteristics into multiple stages according to their characteristic values. For example, if the image characteristics are divided into three stages, the LPF selection unit 3 may select a filter according to which stage each area of ​​interest is in. If the filters generated by the filter generation unit 5 provide three stages of blur with different strengths, the LPF selection unit 3 selects a filter that provides blur according to the stage of the image characteristics based on the selection information and uses it for processing in the LPF processing unit 6.

[0036] Thus, in this embodiment, filtering can be performed on a single training image PT using multiple filters, and an appropriate filter is adopted for each area of ​​interest according to its image characteristics to perform blurring. For example, filtering may be omitted for bright spots and surrounding pixels. The image characteristic calculation unit 2 may also perform the division on a pixel-by-pixel basis. In this case, the pixel of interest becomes the area of ​​interest.

[0037] (Operation) Next, the operation of the first embodiment will be described with reference to Figure 3. Figure 3 is a flowchart for explaining the operation of the first embodiment.

[0038] First, let's look at an example of using contrast as an image characteristic for selecting a filter.

[0039] (Example of contrast) Figure 4 is an explanatory diagram to illustrate the change in contrast characteristic values ​​based on differences in images. Figure 5 is a graph showing an example of a corrected PSF for the LPF based on differences in contrast, with the frequency band on the horizontal axis and contrast on the vertical axis.

[0040] In S1 of Figure 3, the target frequency response characteristic setting unit 4 sets the target frequency response characteristics of the filter required when blurring the teacher image PT to create the student image PS in the filter generation unit 5. The filter generation unit 5 generates multiple filters according to the setting in the target frequency response characteristic setting unit 4 (S2).

[0041] The reduction processing unit 1 acquires the training image PL (S3) and performs reduction processing. The reduction processing yields the training image PT (S5). The training image PT obtained by the reduction processing unit 1 is also provided to the image characteristic calculation unit 2. The image characteristic calculation unit 2 calculates the image characteristics for the region of interest of the training image PT (S6).

[0042] When contrast is used as the image characteristic for selecting a filter, the image characteristic calculation unit 2 calculates the contrast of the area of ​​interest. Figure 4 shows images PT1, PT2, and PT3 as examples of training images PT. Images PT1, PT2, and PT3 are images obtained by imaging the inside of the body with a high-resolution endoscope, and show the pattern of blood vessels. The graphs shown below each of the images PT1 to PT3 show histograms of the brightness values ​​of a predetermined area of ​​interest in each of the images PT1, PT2, and PT3, with the brightness value on the horizontal axis and the number of pixels on the vertical axis.

[0043] An image with a histogram distribution that spreads out overall has many high pixel values and low pixel values, so it is considered to have high contrast. Conversely, an image with high and low pixel values in a narrow range has relatively few distributions of large luminance differences, so it is considered to have low contrast. That is, the images PT1 to PT3 in FIG. 4 are examples of an image with medium contrast, an image with high contrast, and an image with low contrast, respectively. For example, the image characteristic calculation unit 2 can determine the contrast of an image using the histogram distribution of luminance values and the number of pixels as an index.

[0044] The calculation result of the image characteristics of the image characteristic calculation unit 2 is given to the LPF selection unit 3, and the LPF selection unit 3 selects a filter (LPF) based on the image characteristics (S7).

[0045] FIG. 5 shows characteristics A1 to A3 of three filters LPF_A1 to LPF_A3 as examples of filters generated by the filter generation unit 5. Characteristic A1 has the smallest degree of blurring (weak blurring), characteristic A2 has a medium degree of blurring, and characteristic A3 has the largest degree of blurring (strong blurring). In this case, the LPF selection unit 3 selects, for example, the filter with characteristic A1 for a high-contrast attention area, the filter with characteristic A2 for a medium-contrast attention area, and the filter with characteristic A3 for a low-contrast attention area. Note that the number of filters that can be generated by the filter generation unit 5 is not limited to three, and two or four or more filters with different blurring intensities can be generated.

[0046] That is, the LPF processing unit 6 selects and applies the LPF_A3 with characteristic A3, which has a relatively low frequency response characteristic and gives a strong blurring, to the low-contrast attention area of the image PT3. Then, by applying the filter LPF_A3 to the attention area of the image PT3 by the LPF processing unit 6, a target attention pixel of a student image with relatively strong blurring is generated (S9). The student image generated in this way is sufficiently blurred, and by performing high-quality processing using the inference model obtained by learning this student image and the teacher image, a high-contrast image with less blurring is obtained.

[0047] Further, by applying the filter LPF_A2 with characteristic A2 to the attention area of medium contrast in the image PT1 by the LPF processing unit 6, target attention pixels of a student image with medium blur are generated. Also, the LPF processing unit 6 applies the filter LPF_A1 with characteristic A1 to the attention area of the image PT2. Since the filter LPF_A1 has a relatively high frequency response characteristic and strong blur cannot be obtained even if it is applied, target attention pixels of a student image with a relatively small amount of added blur are generated.

[0048] By applying the inference model obtained by learning using the student image and teacher image thus generated to the low-quality image, for low-contrast image parts and scenes, the effect of improving blur is large and the contrast is sufficiently improved, and for high-contrast image parts, the effect of improving blur is reduced to suppress an increase in noise.

[0049] It is determined whether all student pixels have been generated (S10). If not, the process returns to S6 to perform processing on the next attention area, and when all student pixels have been generated, the process ends.

[0050] Although an example of adopting a filter that adds stronger blur as the contrast of the attention area is lower and adds weaker blur as the contrast of the attention area is higher has been described, a case of adopting a filter that adds stronger blur as the contrast of the attention area is higher and adds weaker blur as the contrast of the attention area is lower is also conceivable. In this case, by applying the inference model obtained by learning using the generated student image and teacher image to the low-quality image, an image with more emphasized contrast can be obtained.

[0051] (Example of frequency band) Next, an example of the case of adopting the local frequency band of the teacher image as an image characteristic for selecting a filter will be described.

[0052] FIG. 6 is a graph showing the frequency band of the attention area with the frequency band on the horizontal axis and the contrast on the vertical axis. Also, FIG. 5 is a graph showing an example of the corrected PSF of the LPF based on the difference in the frequency band with the frequency band on the horizontal axis and the contrast on the vertical axis.

[0053] When a frequency band is used as the image characteristic for selecting a filter, the image characteristic calculation unit 2 converts the image of the region of interest to the frequency space and calculates the frequency band of the region of interest. The frequency space can be defined by line pairs / pixel, and the higher the spatial frequency, the finer the image pattern. Figure 3 shows examples of the frequency band of the region of interest, with a dashed line indicating a low frequency band, a solid line indicating a medium frequency band, and a dashed-dotted line indicating a high frequency band.

[0054] Figure 7 is an explanatory diagram showing the characteristics A11 to A13 of three filters LPF_A11 to LPF_A13 as examples of filters generated by the filter generation unit 5. Characteristic A11 provides the smallest degree of blurring, characteristic A12 provides a medium degree of blurring, and characteristic A13 provides the largest degree of blurring. The frequency band calculation result from the image characteristic calculation unit 2 is provided to the LPF selection unit 3, and the LPF selection unit 3 selects a filter based on the calculation result of the frequency band of the region of interest. For example, if the low-frequency band distribution (centroid, center, peak, etc.) is in the low-frequency band, the LPF selection unit 3 selects filter LPF_A11 with characteristic A11; if it is in the medium-frequency band, it selects filter LPF_A12 with characteristic A12; and if it is in the high-frequency band, it selects filter LPF_A13 with characteristic A13.

[0055] The LPF processing unit 6 uses the filter selected by the LPF selection unit 3 to find student pixels and generate a student image. For example, the LPF processing unit 6 selects and applies LPF_A11, which has relatively high frequency response characteristics and a small degree of blurring, to the low-frequency band of interest. The LPF processing unit 6 also selects and applies LPF_A12, which has relatively moderate frequency response characteristics and a moderate degree of blurring, to the mid-frequency band of interest. The LPF processing unit 6 also selects and applies LPF_A13, which has relatively low frequency response characteristics and a large degree of blurring, to the high-frequency band of interest.

[0056] When the inference model constructed using the student and teacher images thus created is applied to low-resolution images, the blur reduction effect increases with higher frequency bands, resulting in images with enhanced contrast. Even in this case, the filter selected can be changed as appropriate depending on the frequency band; the filter selection should be determined according to the desired image characteristics created by the image enhancement process.

[0057] (Example of edges) Next, we will explain an example of using local edges of the training image as an image characteristic for selecting a filter.

[0058] When edges are used as an image characteristic for selecting a filter, the image characteristic calculation unit 2 detects edges from the image of the area of ​​interest and calculates the strength and fineness of the edges in the area of ​​interest. Alternatively, the image characteristic calculation unit 2 may detect edges based on the difference between the pixel value of the area of ​​interest and the pixel value of a pixel adjacent to the area of ​​interest.

[0059] For example, if the image characteristics relate to edge fineness, and the edge fineness is divided into three stages: fine, normal, and flat, the LPF selection unit 3 may select a filter according to which stage of edge fineness each area of ​​interest is in. If the filters generated by the filter generation unit 5 provide three stages of blur with different strengths, the LPF processing unit 6 selects a filter that provides blur according to the stage of edge fineness based on the selection information. By performing filtering on all images of the teacher image PT, the student image PS is generated.

[0060] For example, when generating student images, a filter that provides strong blur is selected for areas of interest with fine edge structures, and a filter that provides weak blur is selected for areas of interest with flat edges. When an inference model is built using these student and training images, it becomes possible to enhance the contrast of areas of interest with fine edge structures and to improve image quality with weaker contrast for areas of interest with flat edges. Furthermore, for example, in areas of interest with many edges, the contrast can be improved to emphasize the edges, while in areas of interest with few edges, image quality improvement with weaker image enhancement can be expected. Areas with many edges may require more thorough observation than areas with flat edges, and improving the contrast of areas with many edges enables proper observation.

[0061] (Example of brightness) Next, we will explain an example of using the local brightness of the training image as an image characteristic for selecting a filter.

[0062] When brightness is used as an image characteristic for selecting a filter, the image characteristic calculation unit 2 calculates the brightness of the image of the area of ​​interest. For example, the image characteristic calculation unit 2 may use the sum or average value of the luminance values ​​(pixel values) of the area of ​​interest in the training image PT as the brightness of the image.

[0063] The LPF selection unit 3 determines which of the filters with different blur strengths to select according to the brightness value. The LPF processing unit 6 uses the filter selected by the LPF selection unit 3 to find student pixels and generate a student image.

[0064] For example, the LPF processing unit 6 selects and applies a filter with characteristics that provide stronger blurring to brighter areas of focus. Conversely, the LPF processing unit 6 selects and applies a filter with characteristics that provide weaker blurring to darker areas of focus.

[0065] This generates student images in which bright areas are strongly blurred and dark areas are not blurred much. When an inference model created using such student and teacher images is applied to the target endoscope, the contrast can be strongly improved in bright areas (where noise is usually not noticeable), and the contrast can be weakly improved in dark areas (where noise is usually noticeable).

[0066] (Example of color) Next, we will explain an example of using local colors from the training image as image characteristics for selecting a filter.

[0067] When color is used as an image characteristic for selecting a filter, the image characteristic calculation unit 2 calculates the color of the image of the region of interest. For example, the image characteristic calculation unit 2 may use the hue and saturation values ​​of the region of interest of the training image PT, such as the average value, as the color of the image.

[0068] In this case, the filter generation unit 5 generates filters corresponding to each color, each with a different blur intensity. The LPF selection unit 3 decides which of the filters with different blur intensities to select according to the color. The LPF processing unit 6 uses the filter selected by the LPF selection unit 3 to find student pixels and generate a student image.

[0069] For example, the LPF processing unit 6, when the area of ​​interest is not red, selects and applies a filter corresponding to that color, and when it is red, selects and applies a filter that corresponds to red and has characteristics that provide a stronger blur.

[0070] This generates student images in which the red areas are strongly blurred, while the non-red areas are not very blurred. When an inference model created using such student and teacher images is applied to the target endoscope, an image with higher contrast in the red areas can be obtained.

[0071] (Example of object point distance of the subject) Next, we will explain an example of using the object point distance of the subject at a local point in the training image as an image characteristic for selecting a filter. Endoscopes often use single-focus lenses. In this case, the region with the highest contrast is the region where the distance to the subject (hereinafter referred to as the object point distance) is a predetermined distance (hereinafter referred to as the best distance), and this region also has many high-frequency components. Conversely, the larger the difference between the object point distance and the best distance, the lower the contrast becomes, the more low-frequency components there are, and the blurred image occurs. In other words, the local image characteristics can be determined by the local object point distance of the image.

[0072] Figure 8 is an explanatory diagram showing an example of an endoscopic image. In the example in Figure 8, the dark area in the center indicates the deep region of the organ. In region R1 of the image in Figure 8, the blood vessels indicated by lines are clearly visible, and this region R1 has the best object distance and is a high-contrast region. In contrast, for example, region R2 enclosed in a rectangular frame contains many deep regions, and the object distance is relatively farther than the best distance, resulting in insufficient contrast.

[0073] When determining the object distance as an image characteristic for selecting a filter, the image characteristic calculation unit 2 calculates the object distance of the subject in the area of ​​interest. The image characteristic calculation unit 2 can calculate the object distance using various known methods, such as a method using a 3D endoscope, a method using a distance sensor, or a method using phase difference. When using a 3D endoscope, the training image PT is also an image obtained by capturing data with this 3D endoscope.

[0074] Figure 9 is an explanatory diagram showing the characteristics A21 to A23 of three filters LPF_A21 to LPF_A23 as examples of filters generated by the filter generation unit 5. Characteristic A21 provides the smallest degree of blurring, characteristic A22 provides a moderate degree of blurring, and characteristic A23 provides the largest degree of blurring. For example, the filter generation unit 5 may use a simulation to determine the corrected PSF when converting the PSF of a target endoscope with a best distance of 15 mm to the PSF of a target endoscope with a best distance of 7 mm, and use that as the filter with characteristic A21. Similarly, the filter generation unit 5 may use a simulation to determine the corrected PSF when converting the PSF of a target endoscope with a best distance of 15 mm to the PSF of a target endoscope with a best distance of 6 mm, and use that as the filter with characteristic A22. Furthermore, the filter generation unit 5 may use a simulation to determine the corrected PSF when converting the PSF of a target endoscope with a best distance of 15 mm to the PSF of a target endoscope with a best distance of 5 mm, and use that as the filter with characteristic A23.

[0075] The result of the object point distance calculation by the image characteristic calculation unit 2 is provided to the LPF selection unit 3, and the LPF selection unit 3 selects a filter based on the calculated object point distance of the area of ​​interest. For example, if the object point distance of the image of interest is 14-16 mm, the LPF selection unit 3 selects the filter LPF_A21 with characteristic A21; if the object point distance of the image of interest is 12-14 mm or 16-18 mm, it selects the filter LPF_A22 with characteristic A22; and if the object point distance of the image of interest is 12 mm or less or 18 mm or more, it selects the filter LPF_A23 with characteristic A23.

[0076] The LPF processing unit 6 uses the filter selected by the LPF selection unit 3 to find student pixels and generate a student image. For example, if the object distance of the target image is close to the best distance, such as 14-16 mm, the LPF processing unit 6 selects and applies filter LPF_A1, which provides a small degree of blurring. If the object distance of the target image is 12-14 mm or 16-18 mm, the LPF processing unit 6 applies filter LPF_A22, which provides a moderate degree of blurring. If the object distance of the target image is 12 mm or less or 18 mm or more, the LPF processing unit 6 applies filter LPF_A23, which provides the greatest degree of blurring. As a result, the closer the object distance of the target region in the teacher image is to the best distance, the less blurring is generated in the student image, and the further away the object distance is from the best distance, the more strongly blurred the student image is generated.

[0077] By applying the inference model constructed using the student and teacher images thus created to the target image, the blur reduction effect is small when the object distance is close to the best distance, and the blur reduction effect becomes larger as the object distance moves further from the best distance. This results in an image with good contrast regardless of the object distance, enabling high image quality similar to dynamic focus processing applied to the entire screen.

[0078] In this embodiment, the image characteristics of the region of interest in the training image are determined, and student images are generated by applying filters with varying degrees of blur based on these image characteristics. This makes it possible to generate student images with a scene-adaptive difference in the degree of blur compared to the training image. As a result, when an inference model (artificial intelligence) constructed using such student and training images is applied to low-resolution images, scene-adaptive image enhancement can be achieved.

[0079] (Second Embodiment) Figure 10 is a block diagram showing a second embodiment of the present invention. In Figure 10, the same reference numerals are used for components that are the same as those in Figure 1, and their descriptions are omitted. In the first embodiment described above, multiple filters were generated in advance, and one filter was selected according to the image characteristics to generate the student image. In contrast, this embodiment generates a filter according to the image characteristics each time.

[0080] In Figure 10, the image processing device 20 differs from the image processing device 10 in Figure 1 in that the LPF selection unit 3 is omitted and a filter generation unit 21 is used instead of the filter generation unit 5. The image characteristic calculation unit 2 provides the filter generation unit 21 with image characteristic information obtained for the region of interest of the training image PT. The filter generation unit 21 is provided with the target frequency response characteristics from the target frequency response characteristic setting unit 4. The filter generation unit 21 generates a filter that has characteristics approximating the target frequency response characteristics and corresponds to the image characteristics.

[0081] Figure 11 is an explanatory diagram for illustrating the target frequency response characteristics obtained by the target frequency response characteristic setting unit 4, and Figures 12 and 13 are graphs for illustrating filter generation, with the frequency bandwidth on the horizontal axis and the contrast on the vertical axis.

[0082] Figure 11 shows an example of determining the target frequency response characteristics using a chart CH for SFR (Spatial Frequency Response) measurement. The filter generation unit 21 simulates obtaining a high-resolution image obtained by taking an image with the target endoscope using the prepared SFR chart CH. Similarly, the filter generation unit 21 simulates obtaining a low-resolution image obtained by taking an image with the target endoscope using the prepared SFR chart CH. The filter generation unit 21 determines the response spatial frequency characteristics SFR_High and SFR_Low for each spatial frequency by performing Fourier transform processing on the obtained high-resolution and low-resolution images.

[0083] The central graph in Figure 11 shows these spatial frequency characteristics, SFR_High and SFR_Low. The target frequency response characteristic setting unit 4 determines the target frequency response characteristic shown in the right column of Figure 11 using SFR_Low / SFR_High.

[0084] The filter generation unit 21 generates a filter having characteristics that approximate the target frequency response characteristics according to the image characteristics obtained for the region of interest of the training image PT. Figure 12 shows the target frequency response characteristics and three characteristics A31 to A33 that approximate these target frequency response characteristics. Characteristic A31 approximates the high-frequency band characteristics of the target frequency response characteristics, characteristic A32 approximates the mid-frequency band characteristics of the target frequency response characteristics, and characteristic A33 approximates the low-frequency band characteristics of the target frequency response characteristics. The filter generation unit 21 can generate filters having characteristics such as characteristics A31 to A33, and determines the characteristics of the filter to be generated according to the image characteristics obtained for the region of interest. Note that the filters that the filter generation unit 21 can generate are not limited to three, and it can generate filters having characteristics that approximate each frequency band of the target frequency response characteristics.

[0085] The filter generation unit 21 may generate multiple filters with different blur intensities through trial and error or empirically. For example, as shown in Figure 13, the filter generation unit 21 may generate filters with different blur strengths, A41 to A43, based on a Gaussian function.

[0086] Thus, the same effects as in the first embodiment can be obtained in this embodiment as well. Furthermore, in this embodiment, a filter is generated according to the image characteristics, making it possible to perform blurring that is more in line with the image characteristics.

[0087] Figure 14 is a block diagram illustrating an AI image enhancement system that utilizes teacher images and student images generated in each of the above embodiments. The AI ​​image enhancement system in Figure 14 includes a learning device that generates an inference model using the teacher images and student images generated in each of the above embodiments, and an image processing device that performs image enhancement processing using the inference model (artificial intelligence) constructed by this learning device.

[0088] The deep learning unit 31 is given the teacher image PT generated by each of the above embodiments and the student image PS generated based on this teacher image, and performs deep learning using these teacher image PT and student image PS. The deep learning unit 31 can be configured, for example, as a neural network. The deep learning unit 31 obtains the parameters of the neural network through deep learning and provides the obtained parameter information as model information to the inference model unit 32, which is an artificial intelligence.

[0089] A neural network consists of an input layer, hidden layers, and an output layer, each composed of multiple nodes. Each node is connected to the nodes in the preceding and succeeding layers, and each connection is assigned a parameter called a weight coefficient. Learning is the process of updating the parameters to minimize the learning loss between high-resolution and low-resolution images. For example, a convolutional neural network (CNN) may be used as the neural network. An inference model constructed in this way will, for example, take a low-resolution input image as input, and through inference processing obtain and output a high-resolution image.

[0090] The inference model unit 32 is composed of a neural network similar to that of the deep learning unit 31. It receives model information from the deep learning unit 31 and constructs an inference model by setting the parameters obtained through deep learning into the neural network.

[0091] The low-resolution endoscope 33 is an endoscope that obtains endoscopic images of the inside of the body for diagnosis, treatment, etc., and outputs relatively low-resolution endoscopic images (low-resolution images) to the inference model unit 32. The inference model unit 32 converts the low-resolution images into high-resolution images through inference processing.

[0092] The inference model obtained by the deep learning unit 31 enables scene-adaptive high-resolution image processing, and the inference model unit 32 enables scene-adaptive high-resolution image processing.

[0093] The present invention is not limited to the embodiments described above, and the components can be modified and implemented in practice without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining the multiple components disclosed in the embodiments described above. For example, some components of all the components shown in the embodiments may be deleted. Moreover, components from different embodiments may be appropriately combined.

[0094] Furthermore, many of the controls and functions described here, primarily those illustrated in flowcharts, can be configured by programs, and these controls and functions can be realized by a computer reading and executing these programs. These programs can be recorded or stored, in whole or in part, as computer program products on portable media such as flexible disks, CD-ROMs, and non-volatile memory, or on storage media such as hard disks and volatile memory. They can be distributed or provided at the time of product shipment or via portable media or communication lines. Users can easily realize the image processing method, machine learning method, image processing apparatus, image processing program, non-volatile recording medium containing the image processing program, and computer program product containing the image processing program by downloading and installing the program on a computer via a communication network, or by installing it from a recording medium to a computer.

Claims

1. An image processing method for generating student images for machine learning from a teacher image, wherein the image characteristic calculation unit calculates a first image characteristic which is the image characteristic of a first region of the teacher image, calculates a second image characteristic which is the image characteristic of a second region different from the first region of the teacher image, the low-pass filter selection unit selects a first low-pass filter corresponding to the first image characteristic from a group of low-pass filters consisting of multiple types of low-pass filters, and selects a second low-pass filter corresponding to the second image characteristic, and the low-pass filter processing unit performs image processing on the first region of the teacher image using the first low-pass filter, and performs image processing on the second region of the teacher image using the second low-pass filter.

2. The image processing method according to claim 1, comprising dividing the entire area of ​​the training image according to predetermined conditions, calculating image characteristics for each divided area, and selecting and applying a low-pass filter corresponding to the image characteristics of the divided area from the low-pass filter group.

3. The image processing method according to claim 2, wherein the predetermined conditions are for each pixel, and the low-pass filter selection unit selects a corresponding low-pass filter from the low-pass filter group for each pixel constituting the training image.

4. The image processing method according to claim 1, wherein the image characteristic is at least one selected from the group consisting of brightness, color, frequency band distribution, frequency response distribution, edge, contrast distribution, or subject distance.

5. The image processing method according to claim 4, wherein, when the image characteristic is the brightness, the low-pass filter selection unit selects a corresponding low-pass filter from the low-pass filter group based on the average value of the pixel values ​​in the region.

6. The image processing method according to claim 4, wherein, when the image characteristic is an edge, the low-pass filter selection unit selects a corresponding low-pass filter from the low-pass filter group based on the difference between the pixel value of the region and the pixel value adjacent to the region.

7. The image processing method according to claim 4, wherein, when the image characteristic is contrast, the low-pass filter selection unit selects a corresponding low-pass filter from the low-pass filter group based on the training image histogram distribution of the region.

8. The image processing method according to claim 4, wherein, when the image characteristic is subject distance, the training image is an image captured by a 3D endoscope.

9. A machine learning method to obtain an inference model for image quality enhancement processing that converts low-resolution input images into high-resolution images by providing a neural network with the training image described in claim 1 and the student image generated by the image processing method described in claim 1 for training.

10. An image processing apparatus comprising: an artificial intelligence trained by the machine learning method described in claim 9; and an image input unit that inputs received image information to the artificial intelligence.

11. An image processing method comprising: at least one processor calculating a first image characteristic which is an image characteristic of a first region of a training image; calculating a second image characteristic which is an image characteristic of a second region of the training image different from the first region; selecting a first low-pass filter corresponding to the first image characteristic from a group of low-pass filters consisting of multiple types of low-pass filters; selecting a second low-pass filter corresponding to the second image characteristic; performing image processing on the first region of the training image using the first low-pass filter; and performing image processing on the second region of the training image using the second low-pass filter; thereby generating a student image for machine learning from the training image.

12. An image processing program that causes at least one processor to calculate a first image characteristic, which is an image characteristic of a first region of a training image; calculate a second image characteristic, which is an image characteristic of a second region of the training image different from the first region; select a first low-pass filter corresponding to the first image characteristic from a group of low-pass filters consisting of multiple types of low-pass filters; select a second low-pass filter corresponding to the second image characteristic; perform image processing on the first region of the training image using the first low-pass filter; perform image processing on the second region of the training image using the second low-pass filter; and generate student images for machine learning from the training image.

13. A non-volatile recording medium on which an image processing program is recorded, wherein the image processing program causes at least one processor to calculate a first image characteristic, which is an image characteristic of a first region of a training image; calculate a second image characteristic, which is an image characteristic of a second region of the training image different from the first region; select a first low-pass filter corresponding to the first image characteristic from a group of low-pass filters consisting of a plurality of types of low-pass filters; select a second low-pass filter corresponding to the second image characteristic; perform image processing on the first region of the training image using the first low-pass filter; perform image processing on the second region of the training image using the second low-pass filter; and generate student images for machine learning from the training image.

14. A computer program product that records an image processing program that causes at least one computer to calculate a first image characteristic, which is the image characteristic of a first region of a training image; calculate a second image characteristic, which is the image characteristic of a second region of the training image that is different from the first region; select a first low-pass filter corresponding to the first image characteristic from a group of low-pass filters consisting of multiple types of low-pass filters; select a second low-pass filter corresponding to the second image characteristic; perform image processing on the first region of the training image using the first low-pass filter; perform image processing on the second region of the training image using the second low-pass filter; and generate student images for machine learning from the training image.