Machine learning image processing method, machine learning image generation device, machine learning method, machine learning image generation program, and endoscope device

By calculating and adjusting frequency characteristics to match Nyquist frequencies, the method generates training images that ensure sufficient inference performance for low-quality images, addressing the challenge of mismatched imaging conditions in AI systems.

WO2026004143A1PCT designated stage Publication Date: 2026-01-02OLYMPUS MEDICAL SYST CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/023681
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing AI image quality improvement systems face challenges in generating low-quality training images that achieve sufficient inference performance when the imaging conditions of the low-resolution imaging device differ from those of the high-resolution imaging device, as they require precise matching of imaging conditions which is often unavailable.

Method used

An image processing method that calculates frequency characteristics of teacher and sample images, adjusts image size or quality using low-pass filters based on predetermined frequency response characteristics, ensuring the Nyquist frequency of the generated student images matches that of the target low-quality endoscope, thereby generating training images suitable for constructing an inference model.

Benefits of technology

Enables the creation of training images that support sufficient inference performance even without product design information of the imaging device, allowing for effective image quality improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024023681_02012026_PF_FP_ABST
    Figure JP2024023681_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a machine learning image processing method, wherein: a first frequency characteristic of a first teacher image is calculated; a sample frequency characteristic of a sample image, which serves as a sample of a low-quality image obtained by imaging with a low-image-quality endoscope, is calculated; a ratio of the sample frequency characteristic and the first frequency characteristic is calculated as a first frequency response characteristic; the first frequency response characteristic is compared with a predetermined second frequency response characteristic for student image generation; when a Nyquist frequency of the first frequency response characteristic is N2 and a Nyquist frequency of the second frequency response characteristic is N1, if N1 < N2, the image size of the first teacher image is reduced to generate a second teacher image; and if N1 ≥ N2, the first teacher image is reduced in image quality using a low-pass filter set on the basis of the second frequency response characteristic to generate a first student image.
Need to check novelty before this filing date? Find Prior Art

Description

Machine learning image processing method, machine learning image generation device, machine learning method, machine learning image generation program, and endoscope device

[0001] The present invention relates to an image processing method for machine learning suitable for machine learning in AI for high-image-quality processing, an image generation device for machine learning, a machine learning method, an image generation program for machine learning, and an endoscope device.

[0002] In recent years, systems that utilize AI (artificial intelligence) to achieve high-quality image processing have been developed. For example, there is an AI image quality improvement system that uses deep learning to learn a dataset of low-quality student images and high-quality teacher images of the same subject with the same configuration, thereby obtaining a high-quality image from an input low-quality image. In the medical field, AI image quality improvement systems are also sometimes used to improve the quality of endoscopic images.

[0003] It is difficult to obtain a low-quality student image and a high-quality teacher image of the same subject with the same configuration through endoscopic imaging. Therefore, a method is sometimes adopted in which a low-quality image (student image) is generated by performing degradation processing on a high-quality image (teacher image) obtained by imaging with a high-quality endoscope capable of high-quality imaging. The degradation processing is performed based on information about the optical system and imaging element of the high-quality endoscope and information about the optical system and imaging element of a low-quality endoscope that captures low-resolution images during actual use.

[0004] For example, Japanese Patent Application Laid-Open Publication No. 2018-195069 (hereinafter referred to as Patent Document 1) discloses a technique for adding the influence of the imaging element to the generated low-resolution training image by convolving a PSF (Point Spread Function) with the high-resolution training image when generating a low-resolution training image from a high-resolution training image.

[0005] Japanese Patent Application Publication No. 2018-195069

[0006] However, with the technology of Patent Document 1, if the imaging conditions of the low-resolution imaging device to be reproduced differ from the imaging conditions of the imaging device actually used, the inference model obtained by learning high-resolution training images and generated low-resolution training images cannot achieve sufficient inference performance. That is, Patent Document 1 has a problem in that unless the imaging conditions (F-number, wavelength, magnification, pixel size, aperture ratio) of the high-resolution imaging device and the imaging conditions (F-number, wavelength, magnification, pixel size, aperture ratio) of the low-resolution imaging device to be reproduced are understood and low-resolution training images are accurately generated from the high-resolution training images, it may be impossible to obtain training images for building an effective inference model. The present invention aims to provide an image processing method for machine learning, a machine learning image generation device, a machine learning method, a machine learning image generation program, and an endoscope device that can generate training images that enable sufficient inference performance even when product design information of the imaging device is not available.

[0007] An image processing method for machine learning according to one aspect of the present invention calculates a first frequency characteristic of a first teacher image, calculates a sample frequency characteristic of a sample image to be used as a sample of a low-quality image obtained by capturing an image using a low-quality endoscope, calculates a ratio of the sample frequency characteristic to the first frequency characteristic as a first frequency response characteristic, compares the first frequency response characteristic with a predetermined second frequency response characteristic for generating a student image, and, assuming that the Nyquist frequency of the first frequency response characteristic is N2 and the Nyquist frequency of the second frequency response characteristic is N1, if N1 < N2, reduces the image size of the first teacher image to generate a second teacher image, and if N1 ≥ N2, reduces the image quality of the first teacher image using a low-pass filter set based on the second frequency response characteristic to generate a first student image.

[0008] An image generation device for machine learning according to one aspect of the present invention has a processor, wherein the processor receives a first teacher image, receives a sample image to be used as a sample of a low-quality image obtained by capturing an image using a low-quality endoscope, calculates sample frequency characteristics of the sample image, calculates a first frequency response characteristic from the first frequency characteristic and the sample frequency characteristic, compares the first frequency response characteristic with a predetermined second frequency response characteristic for generating a student image, and, assuming that the Nyquist frequency of the first frequency response characteristic is N2 and the Nyquist frequency of the second frequency response characteristic is N1, if N1 < N2, reduces the image size of the first teacher image to generate a second teacher image, and if N1 ≧ N2, reduces the image quality of the first teacher image using a low-pass filter set based on the second frequency response characteristic to generate a first student image.

[0009] A machine learning method according to one aspect of the present invention performs machine learning using the first teacher image as a teacher image and the first student image as a student image.

[0010] In a machine learning method according to another aspect of the present invention, machine learning is performed using the second teacher image as a teacher image and the second student image as a student image.

[0011] A machine learning image generation program according to one aspect of the present invention causes an image receiving unit to receive a first teacher image and a sample image that is a sample of a low-quality image obtained by capturing an image using a low-quality endoscope, causes a frequency characteristic calculation unit to calculate a first frequency characteristic of the first teacher image and a sample frequency characteristic of the sample image, causes a frequency response characteristic calculation unit to calculate a first frequency response characteristic from the first frequency characteristic and the sample frequency characteristic, and causes a comparison unit to compare the first frequency response characteristic with a predetermined second frequency response characteristic for generating a student image, and when N1 < N2, causes the teacher image generation unit to reduce the image size of the first teacher image to generate a second teacher image, and when N1 ≥ N2, causes the student image generation unit to reduce the image quality of the first teacher image using a low-pass filter set based on the second frequency response characteristic to generate a first student image.

[0012] An endoscopic device according to one aspect of the present invention includes a receiving unit that receives endoscopic images, an artificial intelligence for high-definition processing that has learned using a machine learning method, and an output unit that outputs the results of high-definition processing of the endoscopic image by the artificial intelligence for high-definition processing to a display.

[0013] According to the present invention, even when there is no product design information of the imaging device, it is possible to generate a training image that enables sufficient inference performance to be obtained.

[0014] FIG. 1 is a block diagram showing an image generation device for machine learning according to a first embodiment of the present invention. FIG. 2 is a graph showing an example of the frequency response characteristics of a Gaussian filter set by an LPF generation unit 22, with the horizontal axis representing the frequency of an image and the vertical axis representing the contrast. FIG. 3 is an explanatory diagram for explaining a second frequency response characteristic when generating a student image and a first frequency response characteristic that depends on a teacher image, using a graph with the horizontal axis representing line pairs / pixel (frequency intensity) and the vertical axis representing the contrast (contrast intensity). FIG. 4 is a flowchart for explaining the operation of an embodiment. FIG. 5 is an explanatory diagram for explaining the operation of an embodiment. FIG. 6 is a block diagram showing a second embodiment. FIG. 7 is a flowchart for explaining the operation of the second embodiment. FIG. 8 is a block diagram showing an example in which AI for image quality improvement processing is incorporated into an endoscope device.

[0015] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0016] First Embodiment Fig. 1 is a block diagram showing an image generation device for machine learning according to a first embodiment of the present invention. The image generation device for machine learning in Fig. 1 implements an image processing method for machine learning according to an embodiment. This method generates low-quality (low-resolution) student images by applying appropriate blurring to high-quality (high-resolution) teacher images. The method determines the characteristics required for blurring based on an image (hereinafter referred to as a sample image) captured by an imaging device (hereinafter referred to as an assumed imaging device) assumed to capture an image to be inferred, and the teacher image. Furthermore, the size of the teacher image is adjusted based on a comparison between these characteristics and a predetermined second frequency response characteristic when actually generating the student image. This enables the generation of training images (student images) that have image quality equivalent to that of the sample image and enable sufficient inference performance.

[0017] In this embodiment, an example will be described in which a high-quality endoscope is used as an imaging device for obtaining high-quality (high-resolution) images (hereinafter referred to as a high-quality imaging device), and a low-quality endoscope (hereinafter referred to as an assumed endoscope) with relatively lower image quality than the high-quality endoscope is used as an assumed imaging device for obtaining sample images, but the imaging device for obtaining high-quality images and sample images is not limited to an endoscope, and various imaging devices can be used. Furthermore, the image to be inferred is not limited to an in-vivo image.

[0018] The machine-learning image generating device 10 shown in FIG. 1 includes a teacher image generating unit 1, a student image generating unit 2, and a learning image validity evaluating unit 4. Each unit of the machine-learning image generating device 10 may be configured by a processor using a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), an NPU (Neural Processing Unit), or the like, may operate according to a program stored in a memory (not shown) to control each unit, or may realize some or all of its functions using hardware electronic circuits. Furthermore, all of the components of the machine-learning image generating device 10 shown in FIG. 1 do not need to be housed in the same single housing.

[0019] The machine learning image generation device 10 receives high-resolution images captured by a high-quality imaging device as training images, and processes the high-resolution images to generate student images with image quality equivalent to that of low-resolution images captured by a target imaging device. Using super-resolution images as teacher images, the device enables the construction of an inference model for image quality improvement processing by learning a model using a dataset of teacher images and student images.

[0020] The learning images are supplied to the teacher image generation unit 1. For example, high-resolution endoscopic images captured by a high-resolution endoscope are used as the learning images. The teacher image generation unit 1 can perform a reduction process on the input learning images. The reduction ratio of the reduction process by the teacher image generation unit 1 is determined by teacher size setting information from the learning image validity evaluation unit 4, which will be described later. The reduction process by the teacher image generation unit 1 can be achieved by known interpolation or thinning processes.

[0021] The teacher image generated by the teacher image generation unit 1 is supplied to the student image generation unit 2 and the learning image validity evaluation unit 4. The student image generation unit 2 includes an image receiving unit 21 and an LPF generation unit 22. The student image generation unit 2 receives the teacher image via the image receiving unit 21. The image receiving unit 21 may be configured to receive the teacher image when information indicating the validity of the teacher image is provided from the learning image validity evaluation unit 4 (described later). The student image generation unit 2 generates a student image by performing image degradation processing on the received teacher image. The student image generation unit 2 may employ, for example, various low-pass filters (LPFs) that blur images. For example, the student image generation unit 2 may employ a Gaussian filter. A Gaussian filter performs weighting based on a Gaussian function to smoothly change pixel values ​​and achieve a natural blurring effect. The frequency response characteristics (filter characteristics) of the LPFs constituting the student image generation unit 2 are determined by the LPF generation unit 22. Frequency response characteristic information is provided to the LPF generation unit 22 from the outside, and based on the received frequency response characteristic information, the LPF generation unit 22 determines the filter characteristics of the student image generation unit 2. The frequency response characteristic information is information that determines how blurred a high-resolution image should be, and is empirically provided information based on the resolution of the high-resolution image and the resolution of a low-resolution endoscopic image obtained by capturing an image using a hypothetical endoscope.

[0022] FIG. 2 is a graph showing an example of the frequency response characteristics of a Gaussian filter set by the LPF generation unit 22, with the horizontal axis representing the image frequency and the vertical axis representing the contrast. The example in FIG. 2 shows frequency response characteristic curves δ1 to δ3. Each of the frequency response characteristics δ1 to δ3 exhibits a characteristic that decreases contrast as the frequency increases. In other words, applying the filter characteristics represented by the frequency response characteristics δ1 to δ3 to a high-resolution image reduces contrast in finer image portions, resulting in blurring of the image. FIG. 2 shows that the degree of blur increases in the order of frequency response characteristics δ1, δ2, and δ3.

[0023] The student image generation unit 2 generates a student image by reducing the image quality (resolution) of the teacher image by applying the filter characteristics set by the LPF generation unit 22 to the teacher image. In the following description, the filter characteristics set by the LPF generation unit 22 based on the frequency response characteristic information will be referred to as the frequency response characteristics when generating the student image. The student image from the student image generation unit 2 is output as a learning image together with the teacher image.

[0024] The learning image validity evaluation unit 4 includes a control unit 41, a frequency characteristic calculation unit 42, and a teacher size setting information generation unit 43. The control unit 41 comprehensively controls each unit of the learning image validity evaluation unit 4, and provides the image receiving unit 21 of the student image generation unit 2 with information indicating whether the teacher image is valid or not.

[0025] The training image validity evaluation unit 4 receives the teacher image from the teacher image generation unit 1 and the sample image captured by the assumed endoscope via an image receiving unit (not shown). The frequency characteristic calculation unit 42 of the training image validity evaluation unit 4 is controlled by the control unit 41 to calculate the frequency characteristics of the teacher image and the frequency characteristics of the sample image. For example, the frequency characteristic calculation unit 42 may obtain the frequency characteristics of the teacher image and the sample image by performing FFT (Fast Fourier Transform) on the teacher image and the sample image, respectively.

[0026] The control unit 41, which functions as a frequency response characteristic calculation unit, calculates the ratio between the frequency characteristics of the sample image calculated by the frequency characteristic calculation unit 42 and the frequency characteristics of the teacher image to calculate a first frequency response characteristic dependent on the teacher image. If the frequency characteristics of the teacher image are S1 and the frequency characteristics of the sample image are S2, the first frequency response characteristic dependent on the teacher image is calculated as S2 / S1. If the Nyquist frequency N1 of the second frequency response characteristic during student image generation in the student image generation unit 2 is compared with the Nyquist frequency N2 of the first frequency response characteristic dependent on the teacher image, and N1 ≥ N2, a teacher image with frequency characteristic S1 can be provided to the student image generation unit 2 for image processing to produce a student image with a wider frequency characteristic than that of the sample image. However, if the Nyquist frequency N1 of the second frequency response characteristic during student image generation is compared with the Nyquist frequency N2 of the first frequency response characteristic dependent on the teacher image, and N1 < N2, the frequency characteristic of the student image will be narrower than that of the sample image when the teacher image generation unit 2 is provided with frequency characteristic S1 for image processing.

[0027] Therefore, in this embodiment, the frequency characteristics of the teacher image are changed so that the Nyquist frequency N2 of the first frequency response characteristic that depends on the teacher image satisfies N1≧N2 when compared with the Nyquist frequency N1 of the second frequency response characteristic when generating the student image.

[0028] The control unit 41 is provided with predetermined second frequency response characteristic information, and determines frequency response characteristics during student image generation that are identical to the frequency response characteristics determined by the LPF generation unit 22. The frequency response characteristics during student image generation are specified characteristics based on the frequency response characteristic information. The control unit 41, which serves as a comparison unit, compares the second frequency response characteristics during student image generation with first frequency response characteristics that depend on the teacher image to determine the validity of the teacher image, i.e., the validity of the student image generated based on the teacher image. The control unit 41 makes a determination by comparing the second frequency response characteristics during student image generation with the first frequency response characteristic that depends on the teacher image. The control unit 41 determines that, when an inference model is constructed using a data set of teacher images and student images generated based on the teacher images as training images, the inference performance of the inference model for images obtained by the assumed imaging device is sufficiently high, and evaluates the teacher image and student image in this case as valid. Conversely, when the Nyquist frequency N1 of the second frequency response characteristic when generating the student image is compared to the Nyquist frequency N2 of the first frequency response characteristic that depends on the teacher image, the control unit 41 determines that, when an inference model is constructed using a data set of teacher images and student images generated based on the teacher images as training images, the inference performance of the inference model for images obtained by the assumed imaging device is not high, and evaluates the teacher image and student image in this case as invalid. Note that in this embodiment, the validity of the student image and the validity of the teacher image are mutually consistent, and if one is valid, the other is also valid.

[0029] FIG. 3 is an explanatory diagram illustrating a predetermined second frequency response characteristic and a first frequency response characteristic dependent on a teacher image during student image generation, using a graph with line pairs / pixel (frequency intensity) on the horizontal axis and contrast (contrast intensity) on the vertical axis. Note that line pairs / pixel corresponds to the number of black and white line pairs per pixel, i.e., the frequency (resolution) of the image. The upper left column of FIG. 3 shows multiple teacher images obtained by the teacher image generation unit 1. The frequency characteristic calculation unit 42 performs an FFT transformation on these teacher images to obtain the frequency characteristic S1 of the teacher images. The upper center column of FIG. 3 shows the average frequency characteristic S1 of the multiple teacher images. The lower left column of FIG. 3 shows multiple sample images. The frequency characteristic calculation unit 42 performs an FFT transformation on these sample images to obtain the frequency characteristic S2 of the sample images. The lower center column of FIG. 3 shows the average frequency characteristic S2 of the multiple sample images.

[0030] The frequency characteristics of an image can be obtained by performing a two-dimensional FFT transform on the image. For example, three-dimensional frequency characteristics can be obtained by taking the horizontal frequency of the image in the X direction, the vertical frequency of the image in the Y direction, and the contrast of the image in the Z direction. For the sake of simplicity, the example in Figure 3 shows the frequency characteristics of a specific direction of the image on a two-dimensional plane.

[0031] The control unit 41 calculates a first frequency response characteristic that depends on the teacher image by calculating the ratio between the frequency characteristic S2 of the sample image calculated by the frequency characteristic calculation unit 42 and the frequency characteristic S1 of the teacher image. The lower right column of Fig. 3 shows the first frequency response characteristic that depends on the teacher image. The upper right column of Fig. 3 shows the predetermined second frequency response characteristic used when generating the student image.

[0032] The control unit 41 evaluates that the teacher image and the student image generated from this teacher image have validity when the Nyquist frequency N1 of the predetermined second frequency response characteristic when generating the student image is compared with the Nyquist frequency N2 of the first frequency response characteristic that depends on the teacher image, if N1 ≧ N2. The control unit 41 evaluates that the teacher image and the student image generated from this teacher image have invalidity when the Nyquist frequency N1 of the predetermined second frequency response characteristic when generating the student image is compared with the Nyquist frequency N2 of the first frequency response characteristic that depends on the teacher image, if N1 < N2.

[0033] For example, the control unit 41 compares the Nyquist frequencies of the predetermined second frequency response characteristics used in generating the student images with the first frequency response characteristics dependent on the teacher image, and evaluates the teacher image and student image as valid if the Nyquist frequency N1 of the second frequency response characteristics used in generating the student images is compared with the Nyquist frequency N2 of the first frequency response characteristics dependent on the teacher image, and N1 ≥ N2. Furthermore, the control unit 41 evaluates the teacher image and student image as invalid if the Nyquist frequency N1 of the predetermined second frequency response characteristics used in generating the student images is compared with the Nyquist frequency N2 of the first frequency response characteristics dependent on the teacher image, and N1 < N2. If the control unit 41 evaluates the teacher image and student image as invalid, it controls the teacher size setting information generation unit 43 to generate teacher size setting information based on the difference between the Nyquist frequency N1 of the predetermined second frequency response characteristics used in generating the student images and the Nyquist frequency N2 of the first frequency response characteristics dependent on the teacher image. The control unit 41 outputs the teacher size setting information generated by the teacher size setting information generating unit 43 to the teacher image generating unit 1, and causes the teacher image generating unit 1 to generate teacher data again.

[0034] In this embodiment, the generation of teacher images by the teacher image generation unit 1 and the evaluation of validity by the learning image validity evaluation unit 4 are repeated until a judgment is made by the learning image validity evaluation unit 4. When the control unit 41 judges that the teacher image is valid, it gives permission to the student image generation unit 2 to receive the teacher image, and causes the student image generation unit 2 to generate student images based on the teacher image.

[0035] The training image generating unit 1 reduces the training image at a reduction ratio based on the training image size setting information from the training image validity evaluating unit 4 to generate a training image.

[0036] When the teacher image is reduced in size by shrinking the image, the number of pixels relative to the number of line pairs becomes smaller, so the resolution of the teacher image becomes finer and the value of the frequency characteristic S1 of the teacher image becomes larger, so the first frequency response characteristic that depends on the teacher image becomes smaller and the Nyquist frequency N2 also becomes smaller.

[0037] Therefore, by generating teacher size setting information in the training image validity evaluation unit 4 according to the difference between the Nyquist frequency N1 of the predetermined second frequency response characteristic when generating the student image and the Nyquist frequency N2 of the first frequency response characteristic that depends on the teacher image, the Nyquist frequency N1 of the second frequency response characteristic when generating the student image will ultimately satisfy the relationship N1 ≧ N2, where N1 is the Nyquist frequency N2 of the first frequency response characteristic that depends on the teacher image. By providing the teacher image in this case to the student image generation unit 2 and imparting filter characteristics based on the frequency response characteristics when generating the student image to the teacher image, the student image will include the frequency characteristics of the sample image.

[0038] In Figure 3, a method for determining the first frequency response characteristic is described in which a plurality of biological endoscope images are used as the teacher image and the sample image. However, the first frequency response characteristic can also be determined by using an SFR chart image for determining the frequency response characteristic for each of the teacher image and the sample image.

[0039] Next, the operation of the embodiment configured as above will be described with reference to Figures 4 and 5. Figure 4 is a flowchart for explaining the operation of the embodiment, and Figure 5 is an explanatory diagram for explaining the operation of the embodiment.

[0040] High-resolution images obtained by a high-quality imaging device are supplied as learning images to a teacher image generation unit 1. The teacher image generation unit 1 reduces the input learning images at an initial reduction rate to generate teacher images, and outputs the generated teacher images to the learning image validity evaluation unit 4 and the student image generation unit 2. Note that the teacher image generation unit 1 may adopt a reduction rate of 1 as the initial value, and use the input learning images as teacher images as they are.

[0041] 4, the frequency characteristic calculation unit 42 of the training image validity evaluation unit 4 calculates a frequency characteristic S1 of the teacher image, which indicates the relationship between frequency intensity and contrast intensity obtained by FFT-transforming the teacher image. A sample image is also provided to the frequency characteristic calculation unit 42, and the frequency characteristic calculation unit 42 calculates a frequency characteristic S2 of the sample image (hereinafter referred to as the sample frequency characteristic), which indicates the relationship between frequency intensity and contrast intensity obtained by FFT-transforming the sample image (S12). The control unit 41 calculates a ratio between the frequency characteristic S2 and the frequency characteristic S1 to obtain a first frequency response characteristic (a first frequency response characteristic dependent on the teacher image) (S13).

[0042] The control unit 41 also calculates a second frequency response characteristic (a predetermined frequency response characteristic at the time of generating the student image) to be convolved when generating the student image from the teacher image (D). The control unit 41 controls the teacher size setting information generation unit 43 to calculate the amount of resizing of the teacher image and generate teacher size setting information. That is, the teacher size setting information generation unit 43 compares the Nyquist frequencies of the first and second frequency response characteristics and determines whether resizing of the teacher image is necessary (S15).

[0043] The training image validity evaluation unit 4 evaluates the validity of the teacher image based on the magnitude of the Nyquist frequencies. In S15, the training image validity evaluation unit 4 branches the process depending on whether the teacher image is valid, and proceeds to S18 if the teacher image is valid. The LPF generation unit 22 of the student image generation unit 2 sets a second frequency response characteristic based on the frequency response characteristic information as a low-pass filter characteristic in the student image generation unit 2 (S17). The student image generation unit 2 performs image processing on the teacher image supplied from the teacher image generation unit 1 using the second frequency response characteristic to generate a low-quality student image (S18).

[0044] On the other hand, the training image validity evaluation unit 4 determines that the teacher image is not valid, and changes the size of the teacher image by generating teacher size setting information for changing the size of the teacher image and outputting the information to the teacher image generation unit 1 (S16). Thereafter, the processes of S11 to S16 are repeated until the teacher image is evaluated as valid.

[0045] That is, the teacher image generation unit 1 repeats generating teacher images depending on the evaluation by the training image validity evaluation unit 4. In the following description, the teacher images output from the teacher image generation unit 1 will be referred to in order as the first teacher image, the second teacher image, ..., and the frequency characteristics of the first teacher image, the second teacher image, ... will be referred to as the first frequency characteristic, the second frequency characteristic, .... Furthermore, the first frequency response characteristic dependent on the teacher image obtained by the ratio between the second frequency characteristic and the sample frequency characteristic will be referred to as the third frequency response characteristic.

[0046] FIG. 5 illustrates teacher size setting information based on a comparison of Nyquist frequencies.

[0047] The example in Figure 5 shows Example A, in which the validity evaluation of the first teacher image (student image) is valid, and Example B, in which the validity evaluation is invalid, and shows the validity evaluation results, a comparison between the Nyquist frequencies of the frequency response characteristics, and whether or not to change the size of the learning images (teacher image and student image) and the mechanism behind it.

[0048] Example A shows an example in which the Nyquist frequencies N2, N1 of the first frequency response characteristic (solid line) dependent on the teacher image and the predetermined second frequency response characteristic (dashed line) during student image generation satisfy N1≧N2. In this case, the training image validity evaluation unit 4 determines that the teacher image is valid. Therefore, the training image validity evaluation unit 4 instructs the student image generation unit 2 to perform filtering on the teacher image from the teacher image generation unit 1 without generating teacher size setting information to be supplied to the teacher image generation unit 1. The student image generation unit 2 filters the teacher image based on the second frequency response characteristic to generate a student image.

[0049] Example B shows an example in which the Nyquist frequencies N2, N1 of the first frequency response characteristic (solid line) dependent on the teacher image and the predetermined second frequency response characteristic (dashed line) during student image generation are N1 < N2. In this case, the training image validity evaluation unit 4 determines that the first student image is not valid. The training image validity evaluation unit 4 needs to reduce the first frequency response characteristic dependent on the teacher image in order to bring the Nyquist frequency N2 of the first frequency response characteristic dependent on the teacher image closer to the Nyquist frequency N1 of the frequency response characteristic during student image generation. Therefore, the training image validity evaluation unit 4 generates teacher size setting information for reducing the teacher image and provides it to the teacher image generation unit 1.

[0050] The teacher image generation unit 1 generates a teacher image (hereinafter referred to as a second teacher image) by reducing the size of the first teacher image using processes such as thinning and pixel interpolation based on the teacher size setting information. In the frequency space defined by line pairs / pixel, the Nyquist frequency (second frequency characteristic) of the second teacher image is greater than the Nyquist frequency (first frequency characteristic) of the first teacher image. As a result, the first frequency response characteristic dependent on the teacher image is reduced. In this way, the Nyquist frequency N2 of the first frequency response characteristic dependent on the teacher image can be made closer to the Nyquist frequency N1 of the frequency response characteristic used to generate the student image.

[0051] The learning image validity evaluation unit 4 compares the Nyquist frequency N2 of the first frequency response characteristic that depends on the teacher image with the Nyquist frequency N1 of the predetermined second frequency response characteristic when generating the student image, and generates teacher size setting information for controlling the enlargement or reduction processing of the teacher image generation unit 1 until N1≧N2.

[0052] For example, if the second teacher image is evaluated as not being valid based on the Nyquist frequency of the first frequency response characteristic (third frequency response characteristic) that depends on the teacher image based on the second teacher image and the Nyquist frequency of the frequency response characteristic when the student image is generated, the second teacher image is reduced to generate a third teacher image, and the validity of this third teacher image is evaluated.

[0053] In this way, when the training image validity evaluation unit 4 evaluates the teacher image as valid, the teacher image evaluated as valid is supplied to the student image generation unit 2, which generates a student image. This student image has the same resolution as the sample image. As a result, the inference model constructed by learning the inference model using the finally generated teacher image and student image can obtain sufficient inference performance for images captured by the assumed imaging device, and can convert low-resolution images into high-resolution images.

[0054] In this embodiment, when a low-quality student image is generated by applying appropriate blurring to a high-quality teacher image, a frequency response characteristic is calculated based on the frequency characteristics of a sample image captured by an assumed imaging device and the frequency characteristics of the teacher image, the validity of the teacher image is determined based on the frequency response characteristic and a predetermined frequency response characteristic for generating the student image, and an inference model is constructed by learning using the teacher image determined to be valid and the student image generated based on it. This inference model can demonstrate sufficient inference performance for low-quality images obtained by the assumed imaging device, and high-quality images can be obtained. In other words, even if there is no high-quality imaging device for obtaining the teacher image and no product design information for the assumed imaging device, it is possible to construct an inference model with high inference performance.

[0055] Second Embodiment Fig. 6 is a block diagram showing a second embodiment. In Fig. 6, the same components as those in Fig. 1 are assigned the same reference numerals, and their description will be omitted. This embodiment is applied to an AI image quality improvement system that constructs an inference model using training images (teacher images and student images) generated by the machine learning image generation device of the first embodiment.

[0056] The embodiment of Figure 6 differs from Figure 1 in that it adds a deep learning unit 5, an inference model unit 6, and a low-image-quality endoscope 7. The deep learning unit 5 is provided with a teacher image when the training image validity evaluation unit 4 evaluates the teacher image as valid and a student image generated based on the teacher image, and performs deep learning using the teacher image and the student image. The deep learning unit 5 can be configured, for example, by a neural network. The deep learning unit 5 obtains neural network parameters through deep learning and provides information on the obtained parameters to the inference model unit 6 as model information.

[0057] A neural network is composed of an input layer, an intermediate layer (hidden layer), and an output layer, each of which is made up of multiple nodes. Each node is connected to nodes in the previous and next layers, and each connection is assigned a parameter called a weight coefficient. Learning is a process of updating parameters to minimize the learning loss between the super-resolution image and the low-resolution image. A convolutional neural network (CNN), for example, may be used as the neural network. An inference model configured in this way receives a low-resolution input image, and then performs inference processing to obtain and output a high-resolution image.

[0058] The inference model unit 6 is composed of a neural network similar to that of the deep learning unit 5, and constructs an inference model by receiving model information from the deep learning unit 5 and setting the parameters obtained by deep learning into the neural network.

[0059] The low-quality endoscope 7 is an endoscope that obtains endoscopic images of the inside of the body for diagnosis, treatment, etc., and outputs endoscopic images of relatively low quality (low-quality images) to the inference model unit 6, the low-quality images having the same quality as the sample images given to the learning image validity evaluation unit 4. The inference model unit 6 converts the low-quality images into high-quality images by inference processing.

[0060] Next, the operation of the embodiment configured as above will be described with reference to Fig. 7. Fig. 7 is a flowchart for explaining the operation of the second embodiment. In Fig. 7, the same steps as in Fig. 4 are assigned the same reference numerals and the description thereof will be omitted.

[0061] In this embodiment, high-quality teacher images and low-quality student images adjusted to a predetermined size are obtained using a procedure similar to that of the flow in Fig. 4. In S21 of Fig. 7, the deep learning unit 5 generates a trained inference model through deep learning using these teacher images and student images. This inference model is adopted by the inference model unit 6. The inference model unit 6 performs inference on low-quality endoscopic images captured by the low-quality endoscope 7 to obtain high-quality images. Because the image quality of the student images used for learning by the deep learning unit 5 is equivalent to the image quality of the low-quality images obtained by the low-quality endoscope 7, the inference model unit 6 can demonstrate relatively high inference performance and obtain high-quality images.

[0062] In this way, in this embodiment, an inference model is constructed using training images (teacher images and student images) generated by the machine learning image generation device of the first embodiment, making it possible to improve the image quality of images obtained by endoscopes that are actually used.

[0063] FIG. 8 is a block diagram showing an example in which AI for image quality improvement processing is incorporated into an endoscope device.

[0064] The AI ​​110 for high-quality image processing, which has been trained using the above-mentioned machine learning method, can also be used as an endoscopic device 100 together with a receiving unit 120 that receives endoscopic images and an output unit 130 that outputs the results of high-quality image processing of the endoscopic images to a display 200.

[0065] Furthermore, the AI ​​110 for high-quality image processing may be installed in an endoscope processor installed in the examination room, or may reside on the cloud.

[0066] The present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some of the components shown in the embodiments may be omitted. Furthermore, components from different embodiments may be appropriately combined.

[0067] Furthermore, among the technologies described herein, many of the controls and functions, mainly those described in the flowcharts, can be set by a program, and the above-described controls and functions can be realized by a computer reading and executing the program. The program can be recorded or stored, in whole or in part, as a computer program product on a portable medium such as a flexible disk, CD-ROM, or nonvolatile memory, or on a storage medium such as a hard disk or volatile memory, and can be distributed or provided at the time of product shipment, via a portable medium, or via a communication line. A user can easily realize the machine learning image processing method, machine learning image generation device, machine learning method, machine learning image generation program, and endoscope device of the present embodiment by downloading the program via a communication network and installing it on a computer, or by installing it on a computer from a recording medium.

Claims

1. An image processing method for machine learning, comprising the steps of: calculating a first frequency characteristic of a first teacher image; calculating a sample frequency characteristic of a sample image to be used as a sample of a low-quality image obtained by capturing an image using a low-quality endoscope; calculating a ratio of the sample frequency characteristic to the first frequency characteristic as a first frequency response characteristic; comparing the first frequency response characteristic with a predetermined second frequency response characteristic for generating a student image; and, assuming that the Nyquist frequency of the first frequency response characteristic is N2 and the Nyquist frequency of the second frequency response characteristic is N1, if N1 < N2, reducing the image size of the first teacher image to generate a second teacher image; and if N1 ≥ N2, reducing the image quality of the first teacher image using a low-pass filter set based on the second frequency response characteristic to generate a first student image.

2. The image processing method for machine learning described in claim 1, which includes: calculating a second frequency characteristic from a second teacher image; calculating a ratio between the sample frequency characteristic and the second frequency characteristic as a third frequency response characteristic; comparing the third frequency response characteristic with a predetermined second frequency response characteristic for generating a student image; and, assuming that the Nyquist frequency of the third frequency response characteristic is N3, if N1 < N3, reducing the image size of the second teacher image to generate a third teacher image; and if N1 ≥ N3, reducing the image quality of the second teacher image using a low-pass filter set based on the second frequency response characteristic to generate a second student image.

3. An image generation device for machine learning, comprising: an image receiving unit that receives a first teacher image and a sample image that is a sample of a low-quality image obtained by capturing an image using a low-quality endoscope; a frequency characteristic calculation unit that calculates a first frequency characteristic of the first teacher image and a sample frequency characteristic of the sample image; a frequency response characteristic calculation unit that calculates a first frequency response characteristic from the first frequency characteristic and the sample frequency characteristic; a comparison unit that compares the first frequency response characteristic with a predetermined second frequency response characteristic for generating a student image; a teacher image generation unit that generates a second teacher image by reducing the image size of the first teacher image if N1 < N2, where N2 is the Nyquist frequency of the first frequency response characteristic and N1 is the Nyquist frequency of the second frequency response characteristic; and a student image generation unit that generates a first student image by reducing the image quality of the first teacher image using a low-pass filter set based on the second frequency response characteristic if N1 ≧ N2.

4. The machine learning image generation device according to claim 3, wherein the frequency characteristic calculation unit calculates a second frequency characteristic from a second teacher image, the frequency response characteristic calculation unit calculates a ratio between the sample frequency characteristic and the second frequency characteristic as a third frequency response characteristic, the comparison unit compares the third frequency response characteristic with a predetermined second frequency response characteristic for generating a student image, and when N1<N3, the teacher image generation unit reduces the image size of the second teacher image to generate a third teacher image, and when N1≧N3, the student image generation unit reduces the image quality of the second teacher image using a low-pass filter set based on the second frequency response characteristic to generate a second student image.

5. An image generation device for machine learning, comprising a processor that receives a first teacher image, receives a sample image that is a sample of a low-quality image obtained by capturing an image using a low-quality endoscope, calculates sample frequency characteristics of the sample image, calculates a first frequency response characteristic from the first frequency characteristic and the sample frequency characteristic, compares the first frequency response characteristic with a predetermined second frequency response characteristic for generating a student image, and, where N2 is the Nyquist frequency of the first frequency response characteristic and N1 is the Nyquist frequency of the second frequency response characteristic, if N1 < N2, reduces the image size of the first teacher image to generate a second teacher image, and if N1 ≥ N2, reduces the image quality of the first teacher image using a low-pass filter set based on the second frequency response characteristic to generate a first student image.

6. The image generation device for machine learning described in claim 5, wherein the processor: calculates a second frequency characteristic from a second teacher image; calculates a ratio between the sample frequency characteristic and the second frequency characteristic as a third frequency response characteristic; compares the third frequency response characteristic with a predetermined second frequency response characteristic for generating a student image; and, where N1 is the Nyquist frequency of the third frequency response characteristic, if N1 < N3, generates a third teacher image by reducing the image size of the second teacher image; and if N1 ≥ N3, reduces the image quality of the second teacher image using a low-pass filter set based on the second frequency response characteristic to generate a second student image.

7. The machine learning method according to claim 1, wherein machine learning is performed using the first teacher image as a teacher image and the first student image as a student image.

8. The machine learning method according to claim 4, wherein machine learning is performed using the second teacher image as a teacher image and the second student image as a student image.

9. A machine learning image generation program, comprising: an image receiving unit receiving a first teacher image and a sample image used as a sample of a low-quality image obtained by capturing an image using a low-quality endoscope; a frequency characteristic calculation unit calculating a first frequency characteristic of the first teacher image and a sample frequency characteristic of the sample image; a frequency response characteristic calculation unit calculating a first frequency response characteristic from the first frequency characteristic and the sample frequency characteristic; a comparison unit comparing the first frequency response characteristic with a predetermined second frequency response characteristic for generating a student image; where N2 is a Nyquist frequency of the first frequency response characteristic and N1 is a Nyquist frequency of the second frequency response characteristic, if N1 < N2, then the teacher image generation unit reducing the image size of the first teacher image to generate a second teacher image; and if N1 ≧ N2, then the student image generation unit reducing the image quality of the first teacher image using a low-pass filter set based on the second frequency response characteristic to generate a first student image.

10. The machine learning image generation program of claim 9, wherein the frequency characteristic calculation unit calculates a second frequency characteristic from a second teacher image, the frequency response characteristic calculation unit calculates a ratio between the sample frequency characteristic and the second frequency characteristic as a third frequency response characteristic, the comparison unit compares the third frequency response characteristic with a predetermined second frequency response characteristic for generating a student image, and, when N1 < N3, the teacher image generation unit reduces the image size of the second teacher image to generate a third teacher image, and, when N1 ≥ N3, the student image generation unit reduces the image quality of the second teacher image using a low-pass filter set based on the second frequency response characteristic to generate a second student image.

11. An endoscopic device comprising: a receiving unit that receives endoscopic images; an artificial intelligence for image quality improvement processing that has learned using the machine learning method described in claim 5; and an output unit that outputs the results of image quality improvement processing of the endoscopic images by the artificial intelligence for image quality improvement processing to a display.

Citation Information

Patent Citations

  • Image-processing method

    JP1998271323A

  • Learning device, medical information processing device, learning data generation method, learning method, and program

    JP2023088665A

  • Information processing system, endoscope system, trained model, information storage medium, and information processing method

    WO2021090469A1