Machine learning image processing method, machine learning image generation device, machine learning method, machine learning image generation program, and endoscope device

By generating low-quality student images from high-quality teacher images with matched frequency characteristics, the method addresses the challenge of insufficient inference performance due to differing imaging conditions, enabling high-quality image processing without product design information.

WO2026004136A1PCT designated stage Publication Date: 2026-01-02OLYMPUS MEDICAL SYST CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/023648
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing methods for generating training images for AI image quality improvement systems face challenges when the imaging conditions of low-resolution devices differ from the actual conditions, leading to insufficient inference performance without product design information.

Method used

An image processing method that generates low-quality student images from high-quality teacher images by adjusting image size and quality based on frequency characteristics, ensuring the Nyquist frequency of the student images matches that of sample images captured by a target low-quality device.

Benefits of technology

Enables the construction of an inference model with sufficient performance even when product design information is unavailable, allowing for high-quality image processing of low-quality images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024023648_02012026_PF_FP_ABST
    Figure JP2024023648_02012026_PF_FP_ABST
Patent Text Reader

Abstract

This machine learning image processing method comprises: generating, from a first teacher image, a first student image having a lower image quality than the first teacher image; calculating a first frequency characteristic of the first student image; calculating a sample frequency characteristic from a sample image obtained by capturing an image with a low-quality endoscope, said sample image being set as a low-quality image sample; comparing the first frequency characteristic with the sample frequency characteristic; and if the Nyquist frequency of the first frequency characteristic is lower than the Nyquist frequency of the sample frequency characteristic, reducing the image size of the first teacher image to generate a second teacher image, and lowering the image quality of the second teacher image to generate a second student image.
Need to check novelty before this filing date? Find Prior Art

Description

Machine learning image processing method, machine learning image generation device, machine learning method, machine learning image generation program, and endoscope device

[0001] The present invention relates to an image processing method for machine learning suitable for machine learning in AI for high-image-quality processing, an image generation device for machine learning, a machine learning method, a program for generating images for machine learning, and an endoscope device.

[0002] In recent years, systems that utilize AI (artificial intelligence) to achieve high-quality image processing have been developed. For example, there is an AI image quality improvement system that uses deep learning to learn a dataset of low-quality student images and high-quality teacher images of the same subject with the same configuration, thereby obtaining a high-quality image from an input low-quality image. In the medical field, AI image quality improvement systems are also sometimes used to improve the quality of low-quality endoscopic images.

[0003] It is difficult to obtain a low-quality student image and a high-quality teacher image of the same subject with the same configuration through endoscopic imaging. Therefore, a method is sometimes adopted in which a low-quality image (student image) is generated by performing degradation processing on a high-quality image (teacher image) obtained by imaging with a high-quality endoscope capable of high-quality imaging. The degradation processing is performed based on information about the optical system and imaging element of the high-quality endoscope and information about the optical system and imaging element of a low-quality endoscope that captures low-resolution images during actual use.

[0004] For example, Japanese Patent Application Laid-Open Publication No. 2018-195069 (hereinafter referred to as Patent Document 1) discloses a technique for adding the influence of the imaging element to the generated low-resolution training image by convolving a PSF (Point Spread Function) with the high-resolution training image when generating a low-resolution training image from a high-resolution training image.

[0005] Japanese Patent Application Publication No. 2018-195069

[0006] However, with the technology of Patent Document 1, if the imaging conditions of the low-resolution imaging device to be reproduced differ from the imaging conditions of the imaging device actually used, the inference model obtained by learning high-resolution training images and generated low-resolution training images cannot achieve sufficient inference performance. That is, Patent Document 1 has a problem in that it is not possible to obtain training images for building an effective inference model unless the imaging conditions (F-number, wavelength, magnification, pixel size, and aperture ratio) of the high-resolution imaging device and the imaging conditions (F-number, wavelength, magnification, pixel size, and aperture ratio) of the low-resolution imaging device to be reproduced are understood. The present invention aims to provide an image processing method for machine learning, a machine learning image generation device, a machine learning method, a program for machine learning image generation, and an endoscope device that can generate training images that enable sufficient inference performance even when product design information of the imaging device is not available.

[0007] An image processing method for machine learning according to one aspect of the present invention generates a first student image from a first teacher image, the first student image having a lower image quality than the first teacher image, calculates a first frequency characteristic of the first student image, calculates a sample frequency characteristic from a sample image that is a sample of a low-image-quality image obtained by capturing an image using a low-image-quality endoscope, compares the first frequency characteristic with the sample frequency characteristic, and if the Nyquist frequency of the first frequency characteristic is smaller than the Nyquist frequency of the sample frequency characteristic, reduces the image size of the first teacher image to generate a second teacher image, and reduces the image quality of the second teacher image to generate a second student image.

[0008] An image generation device for machine learning according to one aspect of the present invention includes an image receiving unit that receives a first teacher image and a sample image that is a sample of a low-quality image obtained by capturing an image using a low-quality endoscope; a student image generation unit that generates a first student image from the first teacher image, the first student image having a lower quality than the first teacher image; a frequency characteristic calculation unit that calculates a first frequency characteristic of the first student image and a sample frequency characteristic of the sample image; a comparison unit that compares the first frequency characteristic with the sample frequency characteristic; and a teacher image generation unit that reduces the image size of the first teacher image to generate a second teacher image if the Nyquist frequency of the first frequency characteristic is smaller than the Nyquist frequency of the sample frequency characteristic, and the student image generation unit reduces the quality of the second teacher image to generate the second student image.

[0009] Another aspect of the present invention provides an image generation device for machine learning, which includes a processor that receives a first teacher image, receives a sample image that is a sample of a low-quality image obtained by capturing an image using a low-quality endoscope, generates a first student image from the first teacher image, which has lower quality than the first teacher image, calculates a first frequency characteristic of the first student image, calculates a sample frequency characteristic of the sample image, compares the first frequency characteristic with the sample frequency characteristic, and if the Nyquist frequency of the first frequency characteristic is smaller than the Nyquist frequency of the sample frequency characteristic, reduces the image size of the first teacher image to generate a second teacher image, and reduces the image quality of the second teacher image to generate a second student image.

[0010] A machine learning method according to one aspect of the present invention performs machine learning using the first teacher image as a teacher image and the first student image, the Nyquist frequency of which is greater than the Nyquist frequency of the sample frequency characteristic, as a student image.

[0011] A machine learning method according to another aspect of the present invention performs machine learning using the second teacher image as a teacher image and the second student image, the Nyquist frequency of which is greater than the Nyquist frequency of the sample frequency characteristic, as a student image.

[0012] A program for generating images for machine learning according to one aspect of the present invention causes an image receiving unit to receive a first teacher image and a sample image that is a sample of a low-quality image obtained by capturing an image using a low-quality endoscope, causes a student image generation unit to generate a first student image from the first teacher image, the first student image having lower quality than the first teacher image, causes a frequency characteristic calculation unit to calculate a first frequency characteristic of the first student image and a sample frequency characteristic of the sample image, causes a comparison unit to compare the first frequency characteristic with the sample frequency characteristic, and if the Nyquist frequency of the first frequency characteristic is smaller than the Nyquist frequency of the sample frequency characteristic, causes the teacher image generation unit to reduce the image size of the first teacher image to generate a second teacher image, and causes the student image generation unit to generate a second student image by reducing the quality of the second teacher image.

[0013] An endoscopic device according to one aspect of the present invention includes a receiving unit that receives endoscopic images, an artificial intelligence for high-definition processing that has learned using a machine learning method, and an output unit that outputs the results of high-definition processing of the endoscopic image by the artificial intelligence for high-definition processing to a display.

[0014] According to the present invention, even when there is no product design information of the imaging device, it is possible to generate a training image that enables sufficient inference performance to be obtained.

[0015] FIG. 1 is a block diagram showing an image generation device for machine learning according to a first embodiment of the present invention. FIG. 2 is a graph showing an example of the frequency response characteristics of a Gaussian filter set by the LPF generation unit 3, with the horizontal axis representing the frequency of the image and the vertical axis representing the contrast. FIG. 3 is an explanatory diagram for explaining the frequency characteristics of a student image and a sample image, using a graph with the horizontal axis representing line pairs / pixel and the vertical axis representing the contrast. FIG. 4 is a flowchart for explaining the operation of an embodiment. FIG. 5 is an explanatory diagram for explaining the operation of an embodiment. FIG. 6 is a block diagram showing a second embodiment. FIG. 7 is a flowchart for explaining the operation of the second embodiment. FIG. 8 is a block diagram showing an example in which AI for image quality improvement processing is incorporated into an endoscope device.

[0016] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0017] First Embodiment Fig. 1 is a block diagram showing an image generation device for machine learning according to a first embodiment of the present invention. The image generation device for machine learning in Fig. 1 implements an image processing method for machine learning according to an embodiment. This method generates low-quality (low-resolution) student images by applying appropriate blurring to high-quality teacher images, and regenerates the student images by adjusting the size of the teacher images based on a comparison between the generated student images and an image (hereinafter referred to as a sample image) captured by an imaging device (hereinafter referred to as an assumed imaging device) assumed to be used for capturing an image to be inferred, thereby enabling the generation of training images (student images) that enable sufficient inference performance to be obtained.

[0018] In this embodiment, an example will be described in which a high-quality endoscope is used as an imaging device for obtaining high-quality images (hereinafter referred to as a high-quality imaging device), and a low-quality endoscope (hereinafter referred to as an assumed endoscope) with relatively lower image quality than the high-quality endoscope is used as an assumed imaging device for obtaining sample images, but the imaging device for obtaining high-quality images and sample images is not limited to an endoscope, and various imaging devices can be used. Furthermore, the image to be inferred is not limited to an in-vivo image.

[0019] The machine-learning image generating device 10 shown in FIG. 1 includes a teacher image generating unit 1, a student image generating unit 2, and a learning image validity evaluating unit 4. Each unit of the machine-learning image generating device 10 may be configured by a processor using a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), an NPU (Neural Processing Unit), or the like, may operate according to a program stored in a memory (not shown) to control each unit, or may realize some or all of its functions using hardware electronic circuits. Furthermore, all of the components of the machine-learning image generating device 10 shown in FIG. 1 do not need to be housed in the same single housing.

[0020] The machine learning image generation device 10 receives high-quality images captured by a high-quality imaging device as training images, and processes the high-quality images to generate student images with image quality equivalent to that of low-resolution images captured by a target imaging device. Using the high-quality images as teacher images, the device enables the construction of an inference model for high-quality image processing by learning a model using a dataset of teacher images and student images.

[0021] The learning images are supplied to the teacher image generation unit 1. For example, high-quality endoscopic images captured by a high-quality endoscope are used as the learning images. The teacher image generation unit 1 can perform a reduction process on the input learning images. The reduction ratio of the reduction process by the teacher image generation unit 1 is determined by teacher size setting information from the learning image validity evaluation unit 4, which will be described later. The reduction process by the teacher image generation unit 1 can be achieved by known pixel interpolation processing or pixel thinning processing.

[0022] The teacher image generated by the teacher image generation unit 1 is supplied to the student image generation unit 2. The student image generation unit 2 includes an LPF generation unit 3. The student image generation unit 2 receives the teacher image via an image receiving unit (not shown) and generates a student image by processing the received teacher image. The student image generation unit 2 may employ, for example, various low-pass filters (LPFs) that blur images. For example, the student image generation unit 2 may employ a Gaussian filter. A Gaussian filter performs weighting based on a Gaussian function to smoothly change pixel values ​​and achieve a natural blurring effect. The frequency response characteristics (filter characteristics) of the LPFs constituting the student image generation unit 2 are determined by the LPF generation unit 3. The LPF generation unit 3 receives frequency response characteristic information from an external source, and the LPF generation unit 3 determines the filter characteristics of the student image generation unit 2 based on the received frequency response characteristic information. The frequency response characteristic information is information that determines how blurry a high-quality image should be, and is information that is empirically determined based on the resolution of the high-quality image and the resolution of the low-resolution endoscopic image obtained by capturing an image using a target endoscope.

[0023] FIG. 2 is a graph showing an example of the frequency response characteristics of a Gaussian filter set by the LPF generation unit 3, with the horizontal axis representing the image frequency and the vertical axis representing the contrast. The example in FIG. 2 shows frequency response characteristics δ1 to δ3. All of the frequency response characteristics δ1 to δ3 exhibit characteristics that decrease contrast as the frequency increases. In other words, applying the filter characteristics represented by the frequency response characteristics δ1 to δ3 to a high-quality image reduces the contrast in finer image portions, resulting in blurring of the image. FIG. 2 shows that the degree of blur increases in the order of frequency response characteristics δ1, δ2, and δ3.

[0024] The student image generation unit 2 generates a student image by reducing the image quality (resolution) of the teacher image by applying the filter characteristics set by the LPF generation unit 3 to the teacher image. The student image from the student image generation unit 2 is supplied to a learning image validity evaluation unit 4. The learning image validity evaluation unit 4 includes a control unit 41, a frequency characteristic calculation unit 42, and a teacher size setting information generation unit 43. The control unit 41 comprehensively controls each unit of the learning image validity evaluation unit 4.

[0025] The learning image validity evaluation unit 4 also receives sample images captured by the assumed endoscope via an image receiving unit (not shown). The frequency characteristic calculation unit 42 of the learning image validity evaluation unit 4 is controlled by the control unit 41 to calculate the frequency characteristics of the student image and the frequency characteristics of the sample image. For example, the frequency characteristic calculation unit 42 may obtain the frequency characteristics of the student image and the sample image by performing FFT (Fast Fourier Transform) on each of the student image and the sample image.

[0026] The control unit 41, which serves as a comparison unit, determines the validity of the student image by comparing the frequency characteristics of the student image from the student image generation unit 2 with the frequency characteristics of the sample image. That is, the control unit 41 determines whether the frequency characteristics of the student image encompass the frequency characteristics of the sample image. Incidentally, "encompassing" means that the contrast of the student image is greater at each frequency than the contrast of the sample image, but this does not have to be strict; it is sufficient if the frequency-contrast curves are approximately the same or the contrast of the student image is generally greater at each frequency than the contrast of the sample image. In particular, as will be described later, it is sufficient if the maximum frequency of the student image (Nyquist frequency) is equal to or greater than the maximum frequency of the sample image (Nyquist frequency).

[0027] If the control unit 41 determines that the frequency characteristics of the student image include the frequency characteristics of the sample image (hereinafter referred to as a match determination), then if an inference model is constructed using the student image as a learning image, the inference performance of the inference model for images obtained by the assumed imaging device will be sufficiently high, and the control unit 41 evaluates the student image as being valid. Conversely, if the control unit 41 determines that the frequency characteristics of the student image do not include the frequency characteristics of the sample image (hereinafter referred to as a mismatch determination), then if an inference model is constructed using the student image as a learning image, the inference performance of the inference model for images obtained by the assumed imaging device will not be high, and the control unit 41 evaluates the student image as being invalid.

[0028] FIG. 3 is an explanatory diagram illustrating the frequency characteristics of a student image and a sample image using a graph with line pairs / pixel (frequency intensity) on the horizontal axis and contrast (contrast intensity) on the vertical axis. Note that line pairs / pixel corresponds to the number of black and white line pairs per pixel, i.e., the frequency (resolution) of the image. The upper left column of FIG. 3 shows multiple student images obtained by the student image generation unit 2. The frequency characteristic calculation unit 42 performs an FFT transformation on these student images to obtain the frequency characteristics of the student images. The upper right column of FIG. 3 shows the average or representative value of the frequency characteristics of the multiple student images. The lower left column of FIG. 3 shows multiple sample images. The frequency characteristic calculation unit 42 performs an FFT transformation on these sample images to obtain the frequency characteristics of the sample images. The lower right column of FIG. 3 shows the average or representative value of the frequency characteristics of the multiple sample images.

[0029] A two-dimensional FFT transform of an image can obtain the frequency characteristics of each position in the image. For example, three-dimensional frequency characteristics can be obtained by taking the horizontal frequency of the image in the X direction, the vertical frequency of the image in the Y direction, and the contrast of the image in the Z direction. For the sake of simplicity, the example in Figure 3 shows the frequency characteristics of a specific direction of the image on a two-dimensional plane.

[0030] The control unit 41 evaluates the student image as valid if the learning image includes the frequency characteristics of the sample image, i.e., the contrast vs. line pair / pixel characteristics. The control unit 41 evaluates the student image as invalid if the learning image does not include the frequency characteristics of the sample image. However, the amount of calculation required to determine the contrast vs. line pair / pixel characteristics for the student image and the sample image is relatively large. Furthermore, even if the contrast changes slightly, the impact on visibility relative to resolution is considered to be relatively small. Therefore, the control unit 41 may determine whether the frequency characteristics are included using the Nyquist frequency of a specific frequency characteristic, for example, the Nyquist frequency.

[0031] In the following description, the highest frequency (line pair / pixel) at which the contrast becomes 0 is referred to as the Nyquist frequency. In the example of FIG. 3, the Nyquist frequency of the student image is N1, and the Nyquist frequency of the sample image is N2. Note that the control unit 41 may compare frequency characteristics by treating the highest frequency at which the contrast is equal to or less than a predetermined value near 0 as the Nyquist frequency.

[0032] For example, the control unit 41 compares the Nyquist frequencies of the student image and the sample image, and if the Nyquist frequency of the student image is higher than the Nyquist frequency of the sample image, evaluates the student image as valid. Furthermore, if the Nyquist frequency of the student image is lower than the Nyquist frequency of the sample image, the control unit 41 evaluates the student image as invalid. If the control unit 41 evaluates the validity as invalid, it controls the teacher size setting information generation unit 43 to generate teacher size setting information based on the difference between the Nyquist frequency of the student image and the Nyquist frequency of the sample image. The control unit 41 outputs the teacher size setting information generated by the teacher size setting information generation unit 43 to the teacher image generation unit 1, causing the teacher image generation unit 1 to create teacher data again.

[0033] In this embodiment, the teacher image generating unit 1 repeats generating teacher images and the student image generating unit 2 repeats generating student images until the learning image validity evaluating unit 4 determines whether the images match.

[0034] The training image generating unit 1 reduces the training image at a reduction ratio based on the training image size setting information from the training image validity evaluating unit 4 to generate a training image.

[0035] When the teacher image is reduced in size by shrinking the teacher image, the number of pixels relative to the number of line pairs becomes smaller, so the resolution of the teacher image becomes finer and the Nyquist frequency becomes higher. Therefore, in this case, the resolution of the student image from the student image generation unit 2 also becomes finer and the Nyquist frequency becomes higher.

[0036] Therefore, by generating teacher size setting information according to the difference in Nyquist frequency between the student image and the sample image in the training image validity evaluation unit 4, it is possible to ultimately make the Nyquist frequencies of the student image and the sample image match or approximately match. When a match is determined, the training image validity evaluation unit 4 outputs the student image as a training image.

[0037] Next, the operation of the embodiment configured as above will be described with reference to Figures 4 and 5. Figure 4 is a flowchart for explaining the operation of the embodiment, and Figure 5 is an explanatory diagram for explaining the operation of the embodiment.

[0038] 4, the LP characteristics of the student image generation unit 2 are set in advance. That is, frequency response characteristic information for setting, in the student image generation unit 2, frequency response characteristics empirically determined based on the resolution of a high-quality image obtained by a high-quality imaging device and a sample image obtained by a target imaging device is provided to the LPF generation unit 3. The LPF generation unit 3 determines the filter characteristics of the student image generation unit 2 based on the frequency response characteristic information.

[0039] High-quality images obtained by a high-quality imaging device are supplied to a teacher image generation unit 1 as learning images. The teacher image generation unit 1 reduces the input learning images at an initial reduction rate to generate teacher images, and outputs the generated teacher images to a student image generation unit 2. Note that the teacher image generation unit 1 may adopt a reduction rate of 1 as the initial value and use the input learning images as teacher images as they are. The student image generation unit 2 filters the received teacher images to generate low-resolution student images (S2). The student image generation unit 2 outputs the generated student images to a learning image validity evaluation unit 4.

[0040] The learning image validity evaluation unit 4 evaluates the validity of the student image, and if it evaluates that the student image is invalid, it provides teacher size setting information to the teacher image generation unit 1 to regenerate a teacher image. That is, the teacher image generation unit 1 repeats generating teacher images depending on the evaluation of the learning image validity evaluation unit 4, and the student image generation unit 2 generates a student image each time a teacher image is output from the teacher image generation unit 1. In the following description, the teacher images output from the teacher image generation unit 1 will be referred to in order as the first teacher image, the second teacher image, ..., and the student images output from the student image generation unit 2 will be referred to in order as the first student image, the second student image, .... That is, the teacher image generation unit 1 first generates and outputs the first teacher image, and the student image generation unit 2 generates the first student image from the first teacher image.

[0041] The training image validity evaluation unit 4 also receives the sample image and calculates a Nyquist frequency (i) based on the frequency characteristics of the first student image (hereinafter referred to as "first frequency characteristics"), which indicate the relationship between frequency intensity and contrast intensity, obtained by FFT-transforming the first student image (S3). The training image validity evaluation unit 4 also calculates a Nyquist frequency (ii) based on the frequency characteristics of the sample image (hereinafter referred to as "sample frequency characteristics"), which indicate the relationship between frequency intensity and contrast intensity, obtained by FFT-transforming the sample image (S4). The training image validity evaluation unit 4 compares the Nyquist frequency of the first frequency characteristics with the Nyquist frequency of the sample frequency characteristics for the first student image and the sample image (S5). The training image validity evaluation unit 4 evaluates the validity of the first student image by determining whether the difference between the Nyquist frequencies is within a predetermined range (S6). The training image validity evaluation unit 4 branches the process depending on the validity in S6. If the result is a match (the difference between the Nyquist frequencies is within a predetermined range), the training image validity evaluation unit 4 determines that the images are valid and proceeds to S7, outputting the input first student image as a training image. If the result is a mismatch (the difference between the Nyquist frequencies is outside the predetermined range), the training image validity evaluation unit 4 determines that the images are not valid and generates teacher size setting information for changing the size of the teacher image and outputs this information to the teacher image generation unit 1, thereby changing the size of the teacher image (S8). Thereafter, the processes of S2 to S6 are repeated until the student image is evaluated as valid.

[0042] FIG. 5 illustrates teacher size setting information based on a comparison of Nyquist frequencies.

[0043] The example in Figure 5 shows the frequency characteristics of example A, where the validity evaluation of the first student image is valid, and example B, where the validity evaluation is invalid, a comparison between the Nyquist frequency of the sample image and the Nyquist frequency of the first student image, and whether or not to change the size of the learning images (teacher image and student image) and the mechanism behind it.

[0044] Example A shows an example in which, among the frequency characteristics of the first student image (solid line) and the frequency characteristics of the sample image (dashed line), the Nyquist frequency of the sample image is the same as or smaller than the Nyquist frequency of the student image. In this case, the training image validity evaluation unit 4 determines that the first student image is valid. Therefore, the training image validity evaluation unit 4 outputs the input first student image as a training image without generating teacher size setting information to be supplied to the teacher image generation unit 1.

[0045] In Example B, as shown in the frequency characteristics of the first student image (solid line) and the frequency characteristics of the sample image (dashed line), the Nyquist frequency of the first student image is lower than the Nyquist frequency of the sample image. In this case, the training image validity assessment unit 4 determines that the first student image is not valid. In order to bring the Nyquist frequency of the first student image closer to the Nyquist frequency of the sample image, the training image validity assessment unit 4 generates teacher size setting information for reducing the teacher image and provides it to the teacher image generation unit 1.

[0046] In this case, the teacher image generation unit 1 generates a teacher image (second teacher image) by reducing the size of the first teacher image using thinning processing, pixel interpolation, or the like, based on the teacher size setting information. In the frequency space defined by line pairs / pixel, the Nyquist frequency of the second teacher image is higher than the Nyquist frequency of the first strong image. As a result, the Nyquist frequency of the student image (second student image) generated by the student image generation unit 2 is also higher than the Nyquist frequency of the first student image, and the second student image contains higher frequency components. In this way, the Nyquist frequency of the second student image approaches the Nyquist frequency of the sample image.

[0047] The learning image validity evaluation unit 4 compares the Nyquist frequency of the second student image with the Nyquist frequency of the sample image, and generates teacher size setting information for controlling the reduction process of the teacher image generation unit 1 until both values ​​fall within a predetermined range.

[0048] For example, if the second student image is evaluated as not being valid based on the second frequency characteristics of the second student image and the sample frequency characteristics of the sample image, the second teacher image is reduced in size to generate a third teacher image, and this third teacher image is reduced in image quality to generate the third student image.

[0049] In this way, the student image output from the training image validity evaluation unit 4 has the same resolution as the sample image. As a result, the inference model constructed by learning the inference model using the finally generated teacher image and student image can obtain sufficient inference performance for images captured by the assumed imaging device, and can convert low-resolution images into high-quality images.

[0050] As described above, in this embodiment, when a high-quality teacher image is subjected to appropriate blurring to generate a low-quality student image, the size of the teacher image is reduced so that the Nyquist frequency of the frequency characteristics of the sample image captured by the assumed imaging device is equal to or smaller than the Nyquist frequency of the frequency characteristics of the student image. This allows the image quality of the student image used for learning to match the image quality of the sample image. By constructing an inference model through learning using such teacher images and student images, the inference model can demonstrate sufficient inference performance for low-quality images obtained by the assumed imaging device, thereby obtaining high-quality images. In other words, even if there is no high-quality imaging device for obtaining the teacher image and product design information for the assumed imaging device, it is possible to construct an inference model with high inference performance.

[0051] Second Embodiment Fig. 6 is a block diagram showing a second embodiment. In Fig. 6, the same components as those in Fig. 1 are assigned the same reference numerals, and their description will be omitted. This embodiment is applied to an AI image quality improvement system that constructs an inference model using training images (teacher images and student images) generated by the machine learning image generation device of the first embodiment.

[0052] The embodiment of Figure 6 differs from Figure 1 in that it adds a deep learning unit 5, an inference model unit 6, and a low-image-quality endoscope 7. The deep learning unit 5 is provided with a student image when the learning image validity evaluation unit 4 evaluates the student image as valid and a teacher image corresponding to the student image, and performs deep learning using the teacher image and the student image. The deep learning unit 5 can be configured, for example, by a neural network. The deep learning unit 5 obtains neural network parameters through deep learning and provides information on the obtained parameters to the inference model unit 6 as model information.

[0053] A neural network is composed of an input layer, an intermediate layer (hidden layer), and an output layer, each of which is made up of multiple nodes. Each node is connected to nodes in the previous and next layers, and each connection is assigned a parameter called a weight coefficient. Learning is a process of updating parameters to minimize the learning loss between high-resolution images and low-resolution images. A convolutional neural network (CNN), for example, may be used as the neural network. An inference model configured in this way receives a low-resolution input image, and then performs inference processing to obtain and output a high-resolution image.

[0054] The inference model unit 6 is composed of a neural network similar to that of the deep learning unit 5, and constructs an inference model by receiving model information from the deep learning unit 5 and setting the parameters obtained by deep learning into the neural network.

[0055] The low-quality endoscope 7 is an endoscope that obtains endoscopic images of the inside of the body for diagnosis, treatment, etc., and outputs endoscopic images of relatively low quality (low-quality images) to the inference model unit 6, the low-quality images having the same quality as the sample images given to the learning image validity evaluation unit 4. The inference model unit 6 converts the low-quality images into high-quality images by inference processing.

[0056] Next, the operation of the embodiment configured as above will be described with reference to Fig. 7. Fig. 7 is a flowchart for explaining the operation of the second embodiment. In Fig. 7, the same steps as in Fig. 4 are assigned the same reference numerals and the description thereof will be omitted.

[0057] In this embodiment, high-quality teacher images and low-quality student images adjusted to a predetermined size are obtained using a procedure similar to that of the flow in Fig. 4. In S11 of Fig. 7, the deep learning unit 5 generates a trained inference model through deep learning using these teacher images and student images. This inference model is adopted by the inference model unit 6. The inference model unit 6 performs inference on low-quality endoscopic images captured by the low-quality endoscope 7 to obtain high-quality images. Because the image quality of the student images used for learning by the deep learning unit 5 is equivalent to the image quality of the low-quality images obtained by the low-quality endoscope 7, the inference model unit 6 can demonstrate relatively high inference performance and obtain high-quality images.

[0058] In this way, in this embodiment, an inference model is constructed using training images (teacher images and student images) generated by the machine learning image generation device of the first embodiment, making it possible to improve the image quality of images obtained by endoscopes that are actually used.

[0059] FIG. 8 is a block diagram showing an example in which AI for image quality improvement processing is incorporated into an endoscope device.

[0060] The AI ​​110 for high-quality image processing, which has been trained using the above-mentioned machine learning method, can also be used as an endoscopic device 100 together with a receiving unit 120 that receives endoscopic images and an output unit 130 that outputs the results of high-quality image processing of the endoscopic images to a display 200.

[0061] Furthermore, the AI ​​110 for high-quality image processing may be installed in an endoscope processor installed in the examination room, or may reside on the cloud.

[0062] The present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some of the components shown in the embodiments may be omitted. Furthermore, components from different embodiments may be appropriately combined.

[0063] Furthermore, among the technologies described herein, many of the controls and functions, mainly those described in the flowcharts, can be set by a program, and the above-described controls and functions can be realized by a computer reading and executing the program. The program can be recorded or stored, in whole or in part, as a computer program product on a portable medium such as a flexible disk, CD-ROM, or nonvolatile memory, or on a storage medium such as a hard disk or volatile memory, and can be distributed or provided at the time of product shipment, via a portable medium, or via a communication line. A user can easily realize the machine learning image processing method, machine learning image generation device, machine learning method, machine learning image generation program, and endoscope device of the present embodiment by downloading the program via a communication network and installing it on a computer, or by installing it on a computer from a recording medium.

Claims

1. An image processing method for machine learning, comprising: generating a first student image from a first teacher image, the first student image having lower image quality than the first teacher image; calculating a first frequency characteristic of the first student image; calculating a sample frequency characteristic from a sample image that is a sample of a low-image-quality image obtained by capturing an image using a low-image-quality endoscope; comparing the first frequency characteristic with the sample frequency characteristic; and, if the Nyquist frequency of the first frequency characteristic is smaller than the Nyquist frequency of the sample frequency characteristic, reducing the image size of the first teacher image to generate a second teacher image; and reducing the image quality of the second teacher image to generate a second student image.

2. The image processing method for machine learning described in claim 1, wherein the first frequency characteristic is a representative value or average value of the frequency characteristics of a plurality of first student images, and the sample frequency characteristic is a representative value or average value of the frequency characteristics of a plurality of sample images.

3. The image processing method for machine learning described in claim 1, further comprising: calculating a second frequency characteristic from the second student image; comparing the second frequency characteristic with the sample frequency characteristic; and, if the Nyquist frequency of the second frequency characteristic is smaller than the Nyquist frequency of the sample frequency characteristic, reducing the image size of the second teacher image to generate a third teacher image; and reducing the image quality of the third teacher image to generate a third student image.

4. An image generation device for machine learning, comprising: an image receiving unit that receives a first teacher image and a sample image that is a sample of a low-quality image obtained by capturing an image using a low-quality endoscope; a student image generation unit that generates a first student image of lower quality than the first teacher image from the first teacher image; a frequency characteristic calculation unit that calculates a first frequency characteristic of the first student image and a sample frequency characteristic of the sample image; a comparison unit that compares the first frequency characteristic with the sample frequency characteristic; and a teacher image generation unit that reduces the image size of the first teacher image to generate a second teacher image if the Nyquist frequency of the first frequency characteristic is smaller than the Nyquist frequency of the sample frequency characteristic, wherein the student image generation unit reduces the quality of the second teacher image to generate the second student image.

5. The image generation device for machine learning described in claim 4, wherein the frequency characteristic calculation unit calculates a second frequency characteristic from the second student image, the comparison unit compares the second frequency characteristic with the sample frequency characteristic, the teacher image generation unit reduces the image size of the second teacher image to generate a third teacher image when the Nyquist frequency of the second frequency characteristic is smaller than the Nyquist frequency of the sample frequency characteristic, and the student image generation unit reduces the image quality of the third teacher image to generate a third student image.

6. An image generation device for machine learning, comprising a processor that receives a first teacher image, receives a sample image that is a sample of a low-quality image obtained by capturing an image using a low-quality endoscope, generates a first student image from the first teacher image, the first student image having lower quality than the first teacher image, calculates a first frequency characteristic of the first student image, calculates a sample frequency characteristic of the sample image, compares the first frequency characteristic with the sample frequency characteristic, and, if the Nyquist frequency of the first frequency characteristic is smaller than the Nyquist frequency of the sample frequency characteristic, reduces the image size of the first teacher image to generate a second teacher image, and reduces the quality of the second teacher image to generate a second student image.

7. The image generation device for machine learning described in claim 6, wherein the processor: calculates a second frequency characteristic from the second student image; compares the second frequency characteristic with the sample frequency characteristic; and, if the Nyquist frequency of the second frequency characteristic is smaller than the Nyquist frequency of the sample frequency characteristic, reduces the image size of the second teacher image to generate a third teacher image; and reduces the image quality of the third teacher image to generate a third student image.

8. A machine learning method according to claim 1, wherein machine learning is performed using the first teacher image as a teacher image and the first student image, the Nyquist frequency of which is greater than the Nyquist frequency of which is the sample frequency characteristic, as a student image.

9. A machine learning method according to claim 2, wherein machine learning is performed using the second teacher image as a teacher image and the second student image, the Nyquist frequency of which is greater than the Nyquist frequency of the sample frequency characteristic, as a student image.

10. A program for generating images for machine learning, comprising: an image receiving unit receiving a first teacher image and a sample image that is a sample of a low-quality image obtained by capturing an image using a low-quality endoscope; a student image generating unit generating a first student image from the first teacher image, the first student image having lower quality than the first teacher image; a frequency characteristic calculation unit calculating a first frequency characteristic of the first student image and a sample frequency characteristic of the sample image; a comparison unit comparing the first frequency characteristic with the sample frequency characteristic; and, if the Nyquist frequency of the first frequency characteristic is smaller than the Nyquist frequency of the sample frequency characteristic, causing the teacher image generating unit to reduce the image size of the first teacher image to generate a second teacher image; and causing the student image generating unit to generate a second student image by reducing the quality of the second teacher image.

11. The image generation program for machine learning described in claim 10, wherein the frequency characteristic calculation unit calculates a second frequency characteristic from the second student image, the comparison unit compares the second frequency characteristic with the sample frequency characteristic, the teacher image generation unit reduces the image size of the second teacher image to generate a third teacher image if the Nyquist frequency of the second frequency characteristic is smaller than the Nyquist frequency of the sample frequency characteristic, and the student image generation unit reduces the image quality of the third teacher image to generate a third student image.

12. An endoscopic device comprising: a receiving unit that receives endoscopic images; an artificial intelligence for image quality improvement processing that has learned using the machine learning method described in claim 6; and an output unit that outputs the results of image quality improvement processing of the endoscopic images by the artificial intelligence for image quality improvement processing to a display.

Citation Information

Patent Citations

  • Image-processing method

    JP1998271323A

  • Learning device, medical information processing device, learning data generation method, learning method, and program

    JP2023088665A

  • Information processing system, endoscope system, trained model, information storage medium, and information processing method

    WO2021090469A1