Training data generation method for machine learning, machine learning method, machine learning program, and image processing device
Patent Information
- Application Number
- JP2025527170
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2023-06-15
- Filing Date
- 2023-06-15
- Publication Date
- 2026-01-22
AI Technical Summary
Existing machine learning techniques for generating high-quality inference images fail to accurately account for differences in subject distance between imaging devices, leading to incorrect simulation images and trained models, resulting in low-quality inference image generation.
A method that generates simulation images as training data by correcting the image quality of captured images based on a conversion table, considering the correlation between different imaging devices, to accurately reflect the subject distance of both the first and second imaging devices, enabling the creation of a correct trained model for high-quality image inference.
This approach allows for the accurate generation of high-quality inference images by correctly simulating the image quality of a second captured image from a first captured image, ensuring a correct trained model is generated, thereby improving the accuracy of image quality enhancement processing.
Abstract
Description
Machine learning training data generation method, machine learning method, machine learning program, and image processing device
[0001] The present invention relates to a method for generating training data for machine learning, a machine learning method, a machine learning program, and an image processing device.
[0002] Recently, a super-resolution technology has been known that performs image quality improvement processing on a processing target image generated by an imaging device that generates low-quality captured images, thereby generating a high-quality inference image that appears as if it were generated by an imaging device that generates high-quality captured images (see, for example, Patent Document 1). Hereinafter, the imaging device that generates the high-quality captured image will be referred to as the first imaging device, and the imaging device that generates the processing target image will be referred to as the second imaging device. In the technology described in Patent Document 1, image quality improvement processing is performed on the processing target image using a trained model generated by machine learning. Hereinafter, the teacher data (correct answer image) and training data used to generate the trained model are as follows: The teacher data is a captured image generated by the first imaging device (hereinafter referred to as the first captured image). On the other hand, the training data is a simulation image in which blur is added to the first captured image.
[0003] Japanese Patent Application Laid-Open No. 2018-195069
[0004] The technology described in Patent Literature 1 generates simulation images as training data from first captured images as teacher data without taking into account the difference between the subject distance of the first captured image captured by the first imaging device and the subject distance of the image captured by the second imaging device. That is, the technology described in Patent Literature 1 generates simulation images by adding blur to the first captured image, assuming that the subject distance is a specific subject distance. However, if the first captured image is captured at a subject distance different from the specific subject distance, the simulation image cannot be generated correctly. As a result, a correct trained model cannot be generated, and a high-quality inference image cannot be generated with high accuracy.
[0005] The present invention has been made in consideration of the above, and aims to provide a method for generating training data for machine learning, a machine learning method, a machine learning program, and an image processing device that enable the accurate generation of high-quality inference images.
[0006] In order to solve the above-mentioned problems and achieve the object, the method for generating training data for machine learning of the present invention acquires a first captured image generated by a first imaging device and a first subject distance of the first captured image, and generates a simulation image that imitates a second captured image captured at a second subject distance by a second imaging device that generates an captured image of lower image quality than the first captured image, as training data for machine learning corresponding to the first captured image that serves as teacher data, by correcting the image quality of the first captured image based on a conversion table that determines the amount of correction based on the correlation between the first subject distance and the first imaging device, and the second subject distance and the second imaging device.
[0007] A machine learning method according to the present invention acquires teacher data and training data, and generates a learned model by machine learning based on the teacher data and the training data, wherein the teacher data is a first captured image captured by a first imaging device at a first subject distance, and the training data is a simulation image simulating a second captured image captured at a second subject distance by a second imaging device that generates an captured image of lower image quality than the first captured image, and the simulation image is one in which the image quality of the first captured image has been corrected based on a conversion table that determines a correction amount based on the correlation between the first subject distance and the first imaging device, and the second subject distance and the second imaging device.
[0008] The machine learning program of the present invention is a machine learning program that uses as training data a first captured image captured by a first imaging device at a first subject distance and a simulation image that imitates a second captured image captured at a second subject distance by a second imaging device that generates an captured image of lower image quality than the first captured image, and the training data is an image in which the image quality of the first captured image has been corrected based on a conversion table that determines the amount of correction based on the correlation between the first subject distance and the first imaging device, and the second subject distance and the second imaging device.
[0009] The image processing device of the present invention includes a distance information acquisition unit that acquires a subject distance indicating the subject distance of a processing target image on which image quality improvement processing is to be performed, a trained model selection unit that selects a trained model from a plurality of trained models that corresponds to the subject distance acquired by the distance information acquisition unit, and an image processing unit that acquires the processing target image and performs the image quality improvement processing on the processing target image using the trained model selected by the trained model selection unit to generate a high-image quality inference image.
[0010] The machine learning training data generation method, machine learning method, machine learning program, and image processing device according to the present invention make it possible to generate high-quality inference images with high accuracy.
[0011] FIG. 1 is a block diagram showing a configuration of a machine learning training data generation device according to an embodiment. FIG. 2 is a flowchart showing a machine learning training data generation method. FIG. 3 is a diagram explaining the generation process (step S1C). FIG. 4 is a diagram explaining the generation process (step S1C). FIG. 5 is a block diagram showing a configuration of a machine learning device. FIG. 6 is a flowchart showing a machine learning method. FIG. 7 is a diagram explaining the learning process (step S2B). FIG. 8 is a block diagram showing a configuration of a second endoscope system. FIG. 9 is a flowchart showing an image processing method. FIG. 10 is a diagram explaining a first modification of the embodiment. FIG. 11 is a diagram explaining a second modification of the embodiment. FIG. 12 is a diagram explaining a second modification of the embodiment. FIG. 13 is a diagram explaining a second modification of the embodiment. FIG. 14 is a diagram explaining a second modification of the embodiment. FIG. 15 is a diagram explaining a third modification of the embodiment. FIG. 16 is a diagram explaining a fourth modification of the embodiment. FIG. 17 is a diagram explaining a fourth modification of the embodiment.
[0012] Hereinafter, a mode for carrying out the present invention (hereinafter referred to as an embodiment) will be described with reference to the drawings. Note that the present invention is not limited to the embodiment described below. Furthermore, in the description of the drawings, the same parts are given the same reference numerals.
[0013] Below, we will explain in order the machine learning training data generation device 100 and method for generating training data for machine learning, the machine learning device 200 and method for performing machine learning using the training data to generate a trained model, and the second endoscopic system 300 and method for performing image quality improvement processing using the trained model to generate a high-quality inference image.
[0014] [Configuration of Machine Learning Training Data Generation Device] First, the configuration of a machine learning training data generation device 100 that generates training data for machine learning will be described. FIG. 1 is a block diagram showing the configuration of the machine learning training data generation device 100 according to an embodiment. The machine learning training data generation device 100 is an information processing device such as a PC (Personal Computer) or a server, and generates simulation images that serve as training data required when generating a trained model used in image quality improvement processing (super-resolution). Here, the machine learning training data generation device 100 generates the training data from a first captured image generated by a first endoscope system 400.
[0015] Before describing the configuration of the machine learning training data generation device 100, the configuration of a first endoscope system 400 will be described. The first endoscope system 400 is a system used in the medical field for observing the inside of a subject (inside a living organism). As shown in FIG. 1 , the first endoscope system 400 includes a first endoscope 410 and a first image processing device 420.
[0016] The first endoscope 410 corresponds to a first imaging device according to the present invention. This first endoscope 410 is configured, for example, as a flexible endoscope having an imaging unit 411 ( FIG. 1 ) that is partially inserted into a living body and captures an image of a subject inside the living body. The imaging unit 411 has an imaging element such as a CCD (Charge Coupled Device) or a CMOS (Complementary Metal Oxide Semiconductor) that receives light from the subject image and converts it into an electrical signal. Hereinafter, a captured image generated by capturing the subject image using the imaging unit 411 will be referred to as a first captured image.
[0017] The first image processing device 420 includes a controller such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array), and controls the overall operation of the first endoscope system 400. As shown in FIG. 1 , the first image processing device 420 includes an image processing unit 421 and an external interface 422.
[0018] The image processing unit 421 performs predetermined image processing on the first captured image. The first captured image after the image processing is then displayed on a display device (not shown). The first captured image is also output to the outside via the external interface 422.
[0019] As shown in FIG. 1 , the machine learning training data generation device 100 includes an external interface 110, a storage unit 120, and a generation processing unit 130. The storage unit 120 stores a first captured image acquired via the external interface 110. Note that FIG. 1 illustrates the machine learning training data generation device 100 as being configured to acquire the first captured image directly from the first endoscope system 400, but this is not limiting. The machine learning training data generation device 100 may also be configured to acquire the first captured image output from the first endoscope system 400 and stored on a server or the like from the server via the external interface 110. The storage unit 120 also stores various programs executed by the generation processing unit 130, information necessary for processing by the generation processing unit 130, and the like.
[0020] The generation processing unit 130 includes a controller such as a CPU or MPU, or an integrated circuit such as an ASIC or FPGA, and generates a simulation image by executing a generation process described below. Note that detailed functions of the generation processing unit 130 will be described in the "Method for generating training data for machine learning" section described below.
[0021] [Machine Learning Training Data Generation Method] Next, a machine learning training data generation method executed by the above-described machine learning training data generation device 100 will be described. FIG. 2 is a flowchart illustrating the machine learning training data generation method. First, the generation processing unit 130 acquires a first captured image stored in the storage unit 120 (step S1A) and acquires first object distance information indicating the first object distance of the first captured image when captured by the first endoscope 410 (step S1B). Furthermore, the generation processing unit 130 may acquire information about the imaging device that captured the first captured image (hereinafter also referred to as endoscope information). Examples of the imaging device information include at least one of the model and model number of the imaging device, and the optimal object distance at which the endoscope that captured the image is in focus. Examples of imaging devices include an endoscope, an eye catheter, and a digital camera. The imaging device information can be acquired from at least one of an imaging device such as an endoscope, an endoscope processor, and an image. After obtaining the imaging device information, the method for generating training data for machine learning can be executed again using the same image, or when executing the method for generating training data for machine learning on an image or series of images extracted from the same video, the acquisition of imaging device information can be omitted by using the imaging device information that has already been obtained.
[0022] Here, it is assumed that the imaging unit 411 is configured with a stereo camera. In this case, in step S1B, the generation processing unit 130 acquires first object distance information or first object distance information and endoscope information as described below. Specifically, the generation processing unit 130 calculates (acquires) first object distance information indicating the first object distance based on the principle of triangulation by using the relative shift amount in images of the same object in each captured image (first captured images) simultaneously captured from different viewpoints by the stereo camera. Here, the generation processing unit 130 calculates first object distance information indicating the first object distance to an object captured in the center of the first captured image or the first object distance to a specific object captured in the first captured image.
[0023] Also, assume that the imaging element constituting the imaging unit 411 is configured with an imaging element including phase difference pixels. In this case, in step S1B, the generation processing unit 130 acquires first object distance information or first object distance information and endoscope information as described below. Specifically, the generation processing unit 130 calculates (acquires) first object distance information indicating the first object distance based on pixel information corresponding to phase difference pixels in the first captured image. Here, the generation processing unit 130 calculates first object distance information indicating the first object distance to an object captured in the center of the first captured image or the first object distance to a specific object captured in the first captured image.
[0024] After step S1B, the generation processing unit 130 executes a generation process as follows to generate a simulation image (step S1C). Figures 3 and 4 are diagrams illustrating the generation process (step S1C). Specifically, Figure 4 is a diagram showing a corrected point spread function (PSF) stored in the storage unit 120.
[0025] In this embodiment, the generation processing unit 130 generates a simulation image corresponding to an image captured at a predetermined object distance using a second endoscope having a predetermined optimal object distance different from the first endoscope, based on the first object distance information and endoscope information acquired in step S1B. The simulation image serves as training data for machine learning. Here, the first object distance is, for example, 3 mm. Note that the above-mentioned value indicating the first object distance is merely an example, and other values may be used. For convenience of explanation, the above-mentioned value will be used below.
[0026] 3, the generation processing unit 130 generates an optical image of the first endoscope 410 by projecting onto an image plane a first captured image CI1 captured by the first endoscope 410 at an optimal subject distance without blurring for the first endoscope 410 (step S1C1). Specifically, the generation processing unit 130 generates an optical image of the first endoscope 410 by enlarging the length and width of the first captured image CI1 by a predetermined magnification.
[0027] After step S1C1, the generation processing unit 130 reads from the storage unit 120 a corrected PSF (corrected PSF(1) ( FIG. 4 )) corresponding to the first object distance of the first captured image CI1, as shown in FIG. 3 (step S1C2). Here, the storage unit 120 stores image quality correction information (corrected PSF(1)) as a corrected PSF corresponding to the first object distance (3 mm) and the second object distance (2 mm) of the second endoscope 310 (see FIG. 8 ), as shown in FIG. 4 . When generating a blurred image captured with the second imaging device at an object distance of 2 mm as training data, simply correcting the first captured image at the optimal object distance to the object distance of 2 mm will not result in an appropriate image as training data. Therefore, a conversion table is used in which the necessary correction PSF is calculated in advance by correlating the degree of blur of the first imaging device with the degree of blur of the second imaging device. 4, for example, it can be seen that to obtain a blurred image with a subject distance of 2 mm using the second imaging device, the image captured with the first imaging device at the optimum subject distance must be blurred to a subject distance of 3 mm. In other words, it can be seen that the blur must be 3 mm, not 2 mm. When multiple conversion tables for different combinations of imaging devices are stored, the corresponding conversion table may be identified from the imaging device information described above.
[0028] The second endoscope 310 generates a captured image of lower image quality than the first captured image. In this embodiment, the first captured image captured by the first endoscope 410 at a first object distance (3 mm) is used as teacher data (correct image), and the simulation image is training data corresponding to the teacher data, and is a simulation image that imitates a second captured image captured by the second endoscope 310 at a second object distance (2 mm). Note that the above-mentioned value indicating the second object distance is merely an example, and other values may be used. For convenience of explanation, the above-mentioned value will be used below.
[0029] The corrected PSF(1) is calculated as follows. MTF1, which is the amount of blur (MTF (Modulation Transfer Function)) on the image plane of the optical system constituting the first endoscope 410 at a first object distance (3 mm), is calculated by optical simulation using the first object distance (3 mm) and first lens configuration information related to the optical system. MTF2, which is the amount of blur (MTF) on the image plane of the optical system constituting the second endoscope 310 at a second object distance (2 mm), is calculated by optical simulation using the second object distance (2 mm) and second lens configuration information related to the optical system. The corrected PSF(1) is image quality correction information for correcting the amount of blur so that MTF1 becomes MTF2, and is configured using, for example, a two-dimensional filter.
[0030] After step S1C2, the generation processing unit 130 convolves the optical image generated in step S1C2 with the corrected PSF(1) to correct blur (image quality) of the optical image (first captured image) (step S1C3), as shown in Fig. 3. This generates an optical image of the second endoscope 310 (step S1C4).
[0031] After step S1C4, the generation processing unit 130 samples (thins out pixels) the optical image generated in step S1C4, as shown in FIG. 3, to generate a simulation image SI1.
[0032] Then, the generation processing unit 130 stores in the storage unit 120 a pair of the first captured image CI1 (teacher data) used to generate the simulation image SI1 and the simulation image SI1 (training data).
[0033] [Configuration of Machine Learning Device] Next, a description will be given of the configuration of a machine learning device 200 that performs machine learning using the simulation image SI1 generated by the machine learning training data generation device 100 to generate a trained model. FIG. 5 is a block diagram showing the configuration of the machine learning device 200. The machine learning device 200 is an information processing device such as a PC or a server, and performs machine learning using teacher data and training data to generate a trained model. As shown in FIG. 5, the machine learning device 200 includes an external interface 210, a storage unit 220, and a learning processing unit 230.
[0034] The storage unit 220 stores multiple pairs of teacher data (first captured image CI1) and training data (simulation image SI1) acquired via the external interface 210. While FIG. 5 illustrates a configuration in which the machine learning device 200 acquires multiple pairs of teacher data (first captured image CI1) and training data (simulation image SI1) directly from the machine learning training data generation device 100, this is not limiting. The machine learning device 200 may also be configured to acquire multiple pairs of teacher data (first captured image CI1) and training data (simulation image SI1) output from the machine learning training data generation device 100 and stored on a server or the like from the server via the external interface 210. The storage unit 220 also stores various programs executed by the learning processing unit 230, information necessary for processing by the learning processing unit 230, and the like.
[0035] The learning processing unit 230 is configured to include a controller such as a CPU or MPU, or an integrated circuit such as an ASIC or FPGA, and generates a trained model by executing a learning process described below. Note that detailed functions of the learning processing unit 230 will be described in the "machine learning method" section described below.
[0036] [Machine Learning Method] Next, a description will be given of the machine learning method executed by the above-described machine learning device 200. Fig. 6 is a flowchart showing the machine learning method. First, the learning processing unit 230 acquires multiple sets of teacher data (first captured images CI1) and training data (simulation images SI1) stored in the storage unit 220 (step S2A).
[0037] After step S2A, the learning processing unit 230 generates a trained model by executing a learning process as shown below (step S2B). FIG. 7 is a diagram illustrating the learning process (step S2B). As shown in FIG. 7, the learning processing unit 230 repeatedly executes a learning process on the learning model using multiple sets of teacher data (first captured images CI1) and training data (simulation images SI1), and generates the learned model after learning as a trained model MD. The learning model used in the learning process is, for example, a convolutional neural network (CNN). Then, the learning processing unit 230 calculates weight values and bias values for each layer of the CNN, generates these values as a trained model MD, and stores the trained model MD in the storage unit 220.
[0038] The neural network used in the learning process (step S1B) is not limited to a CNN, and other neural networks may be used. Furthermore, various well-known learning algorithms may be used as machine learning algorithms in the neural network. For example, a supervised learning algorithm using backpropagation may be used.
[0039] [Configuration of Second Endoscope System] Next, the configuration of a second endoscope system 300 that generates a high-quality inference image by performing image quality improvement processing using the trained model MD generated by the machine learning device 200 will be described. Fig. 8 is a block diagram showing the configuration of the second endoscope system 300. The second endoscope system 300 is used in the medical field and is a system for observing the inside of a subject (inside a living organism). As shown in Fig. 8, the second endoscope system 300 includes a second endoscope 310 and a second image processing device 320.
[0040] The second endoscope 310 corresponds to a second imaging device according to the present invention. This second endoscope 310 is configured, for example, by a flexible endoscope having an imaging unit 311 ( FIG. 8 ) that is partially inserted into a living body and captures an image of a subject inside the living body. The imaging unit 311 differs from the imaging unit 411 only in that the imaging unit 311 generates a captured image of lower image quality than the imaging unit 411. Hereinafter, the captured image generated by the imaging unit 311 will be referred to as a second captured image.
[0041] The second image processing device 320 includes a controller such as a CPU or an MPU, or an integrated circuit such as an ASIC or an FPGA, and controls the overall operation of the second endoscope system 300. As shown in FIG. 8 , the second image processing device 320 includes an external interface 321, a storage unit 322, and an image processing unit 323.
[0042] The storage unit 322 stores the trained model MD acquired via the external interface 321. Note that, although FIG. 8 illustrates a configuration in which the second endoscope system 300 acquires the trained model MD directly from the machine learning device 200, the configuration is not limited to this. The second endoscope system 300 may be configured to acquire the trained model MD output from the machine learning device 200 and stored in a server or the like from the server via the external interface 321. The storage unit 322 also stores various programs executed by the image processing unit 323, information necessary for processing by the image processing unit 323, and the like.
[0043] The image processing unit 323 performs predetermined image processing on the second captured image. The second captured image after the image processing is then displayed on a display device (not shown). The detailed functions of the image processing unit 323 will be described later in the "Image Processing Method" section.
[0044] [Image Processing Method] Next, the image processing method executed by the second image processing device 320 described above will be described. FIG. 9 is a flowchart illustrating the image processing method. First, the image processing unit 323 acquires a second captured image (step S3A) and reads the trained model MD from the storage unit 322 (step S3B). Then, after step S3B, the image processing unit 323 performs image quality improvement processing on the second captured image (processing target image) acquired in step S3A using the trained model MD to generate a high-quality inference image (step S3C). That is, by performing image quality improvement processing on the second captured image captured by the second endoscope 310, a high-quality inference image such as the first captured image CI1 captured at the first subject distance (3 mm) by the first endoscope 410, which generates a captured image with higher image quality than the second endoscope 310, is generated.
[0045] The present embodiment described above provides the following advantages. In the generation process (step S1C) according to the present embodiment, a simulation image simulating a second captured image captured at a second object distance by a second endoscope 310 that generates a captured image of lower image quality than the first captured image CI1 is generated as training data for machine learning corresponding to the first captured image CI1 (teacher data). The simulation image SI1 is generated by correcting the image quality of the first captured image CI1 based on the first object distance. That is, the simulation image SI1 is generated taking into account the first object distance of the first captured image CI1 when captured by the first endoscope 410. Therefore, according to the present embodiment, the simulation image SI1 can be generated correctly, and as a result, a correct trained model MD can be generated, and a high-image-quality inference image can be generated with high accuracy.
[0046] Other Embodiments Although the embodiments for carrying out the present invention have been described above, the present invention should not be limited to the above-described embodiments. The following modifications 1 to 4 may also be adopted in the above-described embodiments.
[0047] (Modification 1) Fig. 10 is a diagram illustrating Modification 1 of the embodiment. Specifically, Fig. 10 corresponds to Fig. 3 and illustrates the generation process (step S1C) according to Modification 1. In the above-described embodiment, the corrected PSF (corrected PSF(1)) used in the blur correction process (step S1C3) is stored in the storage unit 120, but this is not limiting. For example, as in Modification 1 shown in Fig. 10, the corrected PSF (corrected PSF(1)) may be calculated before the blur correction process (step S1C3).
[0048] Specifically, in the generation process (step S1C) according to Modification 1, step S1C5 is employed instead of step S1C2, as shown in Fig. 10. In step S1C5, the generation processing unit 130 calculates MTF1, which is the amount of blur (MTF) on the image plane of the optical system constituting the first endoscope 410 at the first object distance (3 mm) D1, by optical simulation based on a first object distance (approximately 3 mm) D1 and first lens configuration information I1 related to the optical system constituting the first endoscope 410, as shown in Fig. 10. The generation processing unit 130 also calculates MTF2, which is the amount of blur (MTF) on the image plane of the optical system constituting the second endoscope 310 at the second object distance (2 mm) D2, by optical simulation based on a second object distance (2 mm) D2 and second lens configuration information I2 related to the optical system constituting the second endoscope 310. Then, the generation processing unit 130 calculates a corrected PSF (corrected PSF(1)), which is image quality correction information for correcting the amount of blur, so that MTF1 becomes MTF2. The calculated corrected PSF (corrected PSF(1)) is used in the blur correction process (step S1C3).
[0049] Even when the configuration of the present modified example 1 described above is adopted, the same effects as those of the above-described embodiment are achieved.
[0050] (Modification 2) FIGS. 11 to 14 are diagrams illustrating Modification 2 of the embodiment. Specifically, FIG. 11 corresponds to FIG. 4 and is a diagram illustrating a corrected PSF stored in the storage unit 120 according to Modification 2. FIG. 12 corresponds to FIG. 7 and is a diagram illustrating the learning process (step S1B) according to Modification 2. FIG. 13 corresponds to FIG. 8 and is a block diagram illustrating the configuration of a second endoscope system 300A according to Modification 2. FIG. 14 corresponds to FIG. 9 and is a flowchart illustrating an image processing method according to Modification 2. In the above-described embodiment, the machine learning training data generation device 100 generates only the simulation image SI1 as training data corresponding to the first captured image CI1 (teacher data) in which the first subject distance is 3 mm. However, this is not limited to this. For example, it may generate a simulation image as training data corresponding to a first captured image in which the first subject distance is a value other than 3 mm.
[0051] The generation processing unit 130 of this variant example 2 generates a simulation image SI1 that serves as training data corresponding to a first captured image CI1 in which the first subject distance based on the first subject distance information is a first specific distance, a simulation image SI2 that serves as training data corresponding to a first captured image CI2 in which the first subject distance is a second specific distance, and a simulation image SI3 that serves as training data corresponding to a first captured image CI3 in which the first subject distance is a third specific distance.
[0052] Here, the first specific distance is 3 mm, as in the above-described embodiment. The second specific distance is, for example, 7.5 mm. The third specific distance is, for example, 12.5 mm. Note that the above values indicating the first to third specific distances are merely examples, and other values may also be used. For the sake of convenience, the above values will be used below.
[0053] That is, in step S1A, the generation processing unit 130 of this variant example 2 acquires a first captured image CI1 captured at a first subject distance (3 mm), a first captured image CI2 captured at a first subject distance (7.5 mm), and a first captured image CI3 captured at a first subject distance (12.5 mm).
[0054] Furthermore, in step S1B, the generation processing unit 130 according to this second variant acquires first subject distance information indicating the first subject distance (3 mm), first subject distance information indicating the first subject distance (7.5 mm), and first subject distance information indicating the first subject distance (12.5 mm).
[0055] 11 , the storage unit 120 according to the present modified example 2 stores image quality correction information (corrected PSF(1)) as a corrected PSF corresponding to the first object distance (3 mm) of the first endoscope 410 and the second object distance (2 mm) of the second endoscope 310. The storage unit 120 also stores image quality correction information (corrected PSF(2)) as a corrected PSF corresponding to the first object distance (7.5 mm) of the first endoscope 410 and the second object distance (5 mm) of the second endoscope 310. The storage unit 120 also stores image quality correction information (corrected PSF(3)) as a corrected PSF corresponding to the first object distance (12.5 mm) of the first endoscope 410 and the second object distance (9 mm) of the second endoscope 310.
[0056] The generation processing unit 130 according to the second modification then executes step S1C using the first captured image CI1 captured at the first object distance (3 mm) and the corrected PSF(1) to generate a simulation image SI1. This simulation image SI1 is training data corresponding to the first captured image CI1 captured by the first endoscope 410 at the first object distance (3 mm), and is a simulation image simulating a second captured image captured by the second endoscope 310 at the second object distance (2 mm).
[0057] Furthermore, the generation processing unit 130 according to the second modification executes step S1C using the first captured image CI2 captured at the first object distance (7.5 mm) and the corrected PSF(2) to generate a simulation image SI2. This simulation image SI2 is training data corresponding to the first captured image CI2 captured by the first endoscope 410 at the first object distance (7.5 mm), and is a simulation image simulating a second captured image captured by the second endoscope 310 at the second object distance (5 mm). Note that the above-described values indicating the second object distance are merely examples, and other values may be used. For ease of explanation, the above-described values will be used below.
[0058] Furthermore, the generation processing unit 130 according to the second modification executes step S1C using the first captured image CI3 captured at the first object distance (12.5 mm) and the corrected PSF(3) to generate a simulation image SI3. This simulation image SI3 is training data corresponding to the first captured image CI3 captured by the first endoscope 410 at the first object distance (12.5 mm), and is a simulation image simulating a second captured image captured by the second endoscope 310 at the second object distance (9 mm). Note that the above-described values indicating the second object distance are merely examples, and other values may be used. For ease of explanation, the above-described values will be used below.
[0059] The generation processing unit 130 then stores a pair of the first captured image CI1 (teacher data) used to generate the simulation image SI1 and the simulation image SI1 (training data) in the storage unit 120. The generation processing unit 130 also stores a pair of the first captured image CI2 (teacher data) used to generate the simulation image SI2 and the simulation image SI2 (training data) in the storage unit 120. The generation processing unit 130 also stores a pair of the first captured image CI3 (teacher data) used to generate the simulation image SI3 and the simulation image SI3 (training data) in the storage unit 120.
[0060] Furthermore, in Modification 2, the processing of the machine learning device 200 (learning processing unit 230) also differs from that of the above-described embodiment. Specifically, as shown in Fig. 12 , the learning processing unit 230 according to Modification 2 repeatedly executes a learning process (step S2B (step S2B1)) for a learning model using a plurality of sets of teacher data (first captured images CI1) and training data (simulation images SI1), and generates the learning model after the learning as a learned model MD1.
[0061] Furthermore, as shown in FIG. 12, the learning processing unit 230 of this variant example 2 repeatedly performs a learning process (step S2B (step S2B2)) on the learning model using multiple sets of teacher data (first captured image CI2) and training data (simulation image SI2), and generates the learning model after learning as a learned model MD2.
[0062] Furthermore, as shown in Figure 12, the learning processing unit 230 of this variant example 2 repeatedly performs a learning process (step S2B (step S2B3)) on the learning model using multiple sets of teacher data (first captured image CI3) and training data (simulation image SI3), and generates the learning model after learning as a learned model MD3.
[0063] That is, in step S2A, the learning processing unit 230 of this variant example 2 acquires multiple sets of teacher data (first captured image CI1) and training data (simulation image SI1), multiple sets of teacher data (first captured image CI2) and training data (simulation image SI2), and multiple sets of teacher data (first captured image CI3) and training data (simulation image SI3).
[0064] Then, the learning processing unit 230 according to the present modification 2 associates the generated trained model MD1 with the second subject distance (2 mm) and stores it in the storage unit 220. The learning processing unit 230 also associates the generated trained model MD2 with the second subject distance (5 mm) and stores it in the storage unit 220. Furthermore, the learning processing unit 230 associates the generated trained model MD3 with the second subject distance (9 mm) and stores it in the storage unit 220.
[0065] In addition, in this Modification 2, the configuration of the second endoscopic system 300A is also different from the second endoscopic system 300 described in the above-mentioned embodiment. Specifically, in the second endoscopic system 300A according to this Modification 2, the configuration of the second image processing device 320 is changed compared to the second endoscopic system 300, as shown in Fig. 13 . Hereinafter, the second image processing device 320 according to this Modification 2 will be referred to as the second image processing device 320A.
[0066] The second image processing device 320A corresponds to the image processing device according to the present invention. As shown in Fig. 13 , the second image processing device 320A has the functions of a distance information acquisition unit 324 and a trained model selection unit 325 added to the functions of the second image processing device 320.
[0067] Below, we will explain the image processing method according to this modification 2, and also explain the functions of the distance information acquisition unit 324 and the trained model selection unit 325. As shown in Figure 14, the image processing method according to this modification 2 is different from the image processing method described in the above embodiment in that steps S3D and S3E are added instead of step S3B.
[0068] Step S3D is executed after step S3A. Specifically, in step S3D, the distance information acquisition unit 324 acquires second subject distance information indicating the second subject distance of the second captured image acquired in step S3A.
[0069] Here, it is assumed that the imaging unit 311 is configured with a stereo camera. In this case, the distance information acquisition unit 324 acquires second subject distance information in step S3D as described below. Specifically, the distance information acquisition unit 324 calculates (acquires) second subject distance information indicating the second subject distance based on the principle of triangulation by using the relative shift amount in images of the same subject in each captured image (second captured images) simultaneously captured from different viewpoints by the stereo camera. Here, the distance information acquisition unit 324 calculates the second subject distance information indicating the second subject distance to a subject captured in the center of the second captured image or the second subject distance to a specific subject captured in the second captured image.
[0070] Also, assume that the imaging element constituting the imaging unit 311 is configured with an imaging element including phase difference pixels. In this case, the distance information acquisition unit 324 acquires second subject distance information in step S3D as described below. Specifically, the distance information acquisition unit 324 calculates (acquires) first subject distance information indicating the second subject distance based on pixel information corresponding to the phase difference pixels in the second captured image. Here, the distance information acquisition unit 324 calculates second subject distance information indicating the second subject distance to a subject appearing in the center of the second captured image or the second subject distance to a specific subject appearing in the second captured image.
[0071] After step S3D, the trained model selection unit 325 selects, from the trained models MD1 to MD3 stored in the storage unit 322, a trained model corresponding to the second object distance based on the second object distance information acquired in step S3D (step S3E). Then, in the image quality improvement process of step S3C, the trained model selected in step S3E is used. That is, by performing image quality improvement processing on the second captured image captured by the second endoscope 310, a high-image-quality inference image such as the first captured image CI1 captured at the first object distance and the optimal object distance by the first endoscope 410, which generates an image with higher image quality than the second endoscope 310, is generated. In this case, for example, in machine learning, the image to be improved in image quality to be applied to a learning model trained to obtain an image with an object distance of 2 mm as the optimal object distance does not necessarily have to have an object distance of 2 mm. In this case, even an image with an object distance of 1 mm to 3 mm can be image quality converted to obtain the optimal object distance. In other words, even a blurred image whose subject distance differs by, for example, ±1.5 mm from the learning image used during machine learning can be enhanced in image quality using the same machine learning model. Furthermore, by performing image quality enhancement processing on a second captured image captured by the second endoscope 310 at a second subject distance (approximately 5 mm), a high-quality inference image such as the first captured image CI2 captured at the optimal subject distance by the first endoscope 410, which generates a captured image with higher image quality than the second endoscope 310, is generated. Furthermore, by performing image quality enhancement processing on a second captured image captured by the second endoscope 310 at a second subject distance (approximately 9 mm), a high-quality inference image such as the first captured image CI3 captured at the optimal subject distance by the first endoscope 410, which generates a captured image with higher image quality than the second endoscope 310, is generated.
[0072] Even when the configuration of the second modification described above is adopted, the same effects as those of the above-described embodiment are achieved. Furthermore, in the second modification, simulation images SI1 to SI3 are generated for each of the first captured images CI1 to CI3, which have different first subject distances, based on the first subject distance. Furthermore, trained models MD1 to MD3 are generated by machine learning for each of the pair of the first captured image CI1 (teacher data) and the simulation image SI1 (training data), the pair of the first captured image CI2 (teacher data) and the simulation image SI2 (training data), and the pair of the first captured image CI3 (teacher data) and the simulation image SI3 (training data). Then, among the trained models MD1 to MD3, a trained model corresponding to the second subject distance of the second captured image captured by the second endoscope 310 is used to perform image quality improvement processing on the second captured image, thereby generating a high-quality inference image. Therefore, even if the second endoscope 310 generates a second captured image at various second subject distances, a high-quality inference image can be generated with high accuracy from the second captured image using a trained model corresponding to the second subject distance.
[0073] (Variation 3) Fig. 15 is a diagram illustrating Variation 3 of the embodiment. Specifically, Fig. 15 is a diagram corresponding to Fig. 12 and illustrates the learning process (step S1B) according to Variation 3. In Variation 2 described above, the trained models MD1 to MD3 are generated by machine learning from multiple sets of teacher data (first captured image CI1) and training data (simulation image SI1), multiple sets of teacher data (first captured image CI2) and training data (simulation image SI2), and multiple sets of teacher data (first captured image CI3) and training data (simulation image SI3). However, the present invention is not limited to this.
[0074] 15 , in step S2B, the learning processing unit 230 according to the third modification repeatedly executes a learning process for the same learning model using multiple sets of teacher data (first captured image CI1) and training data (simulation image SI1), multiple sets of teacher data (first captured image CI2) and training data (simulation image SI2), and multiple sets of teacher data (first captured image CI3) and training data (simulation image SI3), and generates the learning model after learning as a learned model MD4. The learned model MD4 corresponds to a third learned model according to the present invention.
[0075] The image processing method according to Modification 3 is similar to the embodiment described above, except that the trained model MD is changed to trained model MD4. That is, in Modification 3, as in Modification 2 described above, image quality improvement processing is performed on a second captured image captured by a second endoscope 310, thereby generating a high-quality inference image similar to the first captured image CI1 captured at a first object distance and an optimal object distance by a first endoscope 410 that generates captured images of higher image quality than the second endoscope 310. In this case, for example, during machine learning, the target image for image quality improvement to be applied to a learning model trained to set an image with an object distance of 2 mm to the optimal object distance does not necessarily have to have an object distance of 2 mm. In this case, even an image with an object distance of 1 mm to 3 mm can be converted to the optimal object distance. In other words, even a blurred image with an object distance that differs by, for example, ±1.5 mm from the learning image used during machine learning can be enhanced in image quality using the same machine learning model. Furthermore, by performing image quality improvement processing on the second captured image captured by the second endoscope 310 at the second object distance (approximately 5 mm), a high-quality inference image is generated, such as the first captured image CI2 captured at the optimum object distance by the first endoscope 410, which generates a captured image with higher image quality than the second endoscope 310. Furthermore, by performing image quality improvement processing on the second captured image captured by the second endoscope 310 at the second object distance (approximately 9 mm), a high-quality inference image is generated, such as the first captured image CI3 captured at the optimum object distance by the first endoscope 410, which generates a captured image with higher image quality than the second endoscope 310.
[0076] Even when the configuration of Modification 3 described above is adopted, the same effects as those of the above-described embodiment and Modification 2 are achieved. Furthermore, in Modification 3, a single trained model MD4 is generated by machine learning using multiple sets of teacher data (first captured image CI1) and training data (simulation image SI1), multiple sets of teacher data (first captured image CI2) and training data (simulation image SI2), and multiple sets of teacher data (first captured image CI3) and training data (simulation image SI3). Therefore, when the second endoscope 310 generates second captured images at various second subject distances, a high-quality inference image can be generated with high accuracy from the second captured images using the trained model MD4, even without acquiring the second subject distances.
[0077] 16 and 17 are diagrams illustrating a fourth modification of the embodiment. Specifically, FIG. 16 corresponds to FIG. 3 and illustrates a generation process (step S1C) according to the fourth modification. In addition to generating the simulation image SI1, the generation processing unit 130 according to the fourth modification generates the following simulation image SI4. The simulation image SI4 is a simulation image that uses the simulation image SI1 as training data and is teacher data corresponding to the training data, simulating an image captured at a predetermined subject distance by a third endoscope (not shown) that generates an image with lower image quality than the first captured image CI1 but higher image quality than the second captured image generated by the second endoscope 310.
[0078] Specifically, in step S1C, the generation processing unit 130 according to this fourth modification generates a simulation image SI1 in the same manner as in the above-described embodiment, and also generates a simulation image SI4 as described below. First, as shown in Fig. 16, the generation processing unit 130 according to this fourth modification generates an optical image of the first endoscope 410 by projecting a first captured image CI1 captured by the first endoscope 410 at a first subject distance (3 mm) onto an image plane (step S1C6). Specifically, the generation processing unit 130 generates an optical image of the first endoscope 410 by enlarging the first captured image CI1 vertically and horizontally by a predetermined magnification.
[0079] After step S1C6, the generation processing unit 130 according to the fourth modification reads the corrected PSF (corrected PSF(4)) corresponding to the first object distance of the first captured image CI1 from the storage unit 120 (step S1C7), as shown in Fig. 16. Here, the storage unit 120 stores image quality correction information (corrected PSF(4)) as the corrected PSF corresponding to the first object distance (3 mm) and the predetermined object distance of the third endoscope (not shown).
[0080] Here, the corrected PSF(1) is calculated as follows. MTF1, which is the amount of blur (MTF) on the image plane of the optical system constituting the first endoscope 410 at a first object distance (3 mm), is calculated by optical simulation using the first object distance (3 mm) and first lens configuration information related to the optical system. MTF3, which is the amount of blur (MTF) on the image plane of the optical system constituting the third endoscope at a predetermined object distance, is calculated by optical simulation using the predetermined object distance and third lens configuration information related to the optical system. The corrected PSF(4) is image quality correction information for correcting the amount of blur so that MTF1 becomes MTF3, and is configured using, for example, a two-dimensional filter.
[0081] After step S1C7, the generation processing unit 130 according to the fourth modification performs convolution of the optical image generated in step S1C6 with the corrected PSF(4) to correct blur (image quality) of the optical image (first captured image) (step S1C7), as shown in Fig. 16. This generates a third endoscopic optical image (step S1C8).
[0082] After step S1C8, the generation processing unit 130 according to the fourth modification samples (thins out pixels) the optical image generated in step S1C8, as shown in FIG. 16, to generate a simulation image SI4.
[0083] Then, the generation processing unit 130 according to the fourth modification stores the simulation image SI4 (teacher data) and the simulation image SI1 (training data) as a pair in the storage unit 120.
[0084] In addition, in this modification 4, the processing of the machine learning device 200 (learning processing unit 230) also differs from that of the above-described embodiment. Specifically, as shown in FIG. 17 , the learning processing unit 230 in this modification 4 repeatedly executes a learning process (step S2B) on a learning model using multiple sets of teacher data (simulation images SI4) and training data (simulation images SI1), and generates the learning model after learning as a trained model MD5. The learning processing unit 230 then stores the generated trained model MD5 in the storage unit 220.
[0085] That is, in step S2A, the learning processing unit 230 according to the fourth modification acquires a plurality of sets of teacher data (simulation image SI4) and training data (simulation image SI1).
[0086] The image processing method according to Modification 4 is similar to the above-described embodiment, except that the trained model MD is changed to trained model MD5. That is, in Modification 4, by performing image quality improvement processing on a second captured image captured by the second endoscope 310 at a second subject distance (approximately 2 mm), a high-quality inference image is generated that resembles an image captured at a predetermined subject distance by a third endoscope that generates a captured image of higher image quality than the second endoscope 310.
[0087] Even when the configuration of the fourth modification described above is adopted, the same effects as those of the above-described embodiment are achieved. Furthermore, in the fourth modification, the simulation image SI1 is used as training data, and a simulation image SI4 is generated as teacher data corresponding to the training data, simulating an image captured at a predetermined subject distance by a third endoscope (not shown) that generates images with lower image quality than the first captured image CI1 but higher image quality than the second captured image generated by the second endoscope 310. Furthermore, in the fourth modification, a trained model MD5 is generated by machine learning using multiple sets of teacher data (simulation image SI4) and training data (simulation image SI1). Therefore, by performing image quality improvement processing on the second captured image captured by the second endoscope 310 using the trained model MD5, a high-quality inference image can be generated that resembles an image captured at a predetermined subject distance by the third endoscope that generates images with higher image quality than the second endoscope 310. In other words, since there is no need to use the captured image captured by the target endoscope, such as the third endoscope, as training data for machine learning, it is possible to generate a high-quality inference image similar to the captured image captured by the target endoscope without retaking the actual image of the target endoscope.
[0088] (Variation 5) In the above-described embodiment, the first to third imaging devices according to the present invention are configured using endoscopes, but this is not limited thereto, and any other imaging device may be adopted as long as it is an imaging device that generates an image by capturing an image.
[0089] (Variation 6) In the above-described embodiment, the PSF of the optical image (two-dimensional image) is corrected when generating the simulation image SI1, but this is not limiting. For example, the simulation image SI1 may be generated by transforming the first captured image CI1 into a frequency space and correcting the OTF (Optical Transfer Function) in the frequency space.
[0090] 100 Machine learning training data generation device 110 External interface 120 Storage unit 130 Generation processing unit 200 Machine learning device 210 External interface 220 Storage unit 230 Learning processing unit 300, 300A Second endoscope system 310 Second endoscope 311 Imaging unit 320, 320A Second image processing device 321 External interface 322 Storage unit 323 Image processing unit 324 Distance information acquisition unit 325 Trained model selection unit 400 First endoscope system 410 First endoscope 411 Imaging unit 420 First image processing device 421 Image processing unit 422 External interface CI1 to CI3 First captured image D1 First subject distance D2 Second subject distance I1 First lens configuration information I2 Second lens configuration information MD, MD1 to MD5 Trained model SI1 to SI4 Simulation images
Claims
1. acquiring a first captured image generated by a first imaging device and a first subject distance of the first captured image; A method for generating training data for machine learning, in which a simulation image simulating a second captured image captured at a second subject distance by a second imaging device that generates an image of lower quality than the first captured image is generated as training data for machine learning corresponding to the first captured image that serves as teacher data, is generated by correcting the image quality of the first captured image based on a conversion table that determines the amount of correction based on the correlation between the first subject distance and the first imaging device, and the second subject distance and the second imaging device.
2. acquiring first imaging device information that is information about the first imaging device associated with the first captured image; The method for generating training data for machine learning according to claim 1 , wherein the conversion table is selected based on the first image capture device information.
3. The method for generating machine learning training data according to claim 2 , wherein first lens configuration information relating to an optical system that constitutes the first imaging device is acquired as the first imaging device information.
4. The first object distance is The method for generating training data for machine learning according to claim 1 , wherein the subject distance is the subject distance to the center of the first captured image, or the subject distance to a specified subject that appears in the first captured image.
5. The method for generating training data for machine learning according to claim 1, further comprising generating a simulation image that imitates a third captured image captured by the second imaging device at a third subject distance based on the conversion table.
6. A method for detecting a subject distance by a first imaging device, comprising: receiving a first captured image captured at a first subject distance; a simulation image simulating a second captured image captured at a second object distance by a second imaging device that generates a captured image of lower image quality than the first captured image, the simulation image being generated by correcting the image quality of the first captured image using the first captured image based on a conversion table that defines a correction amount based on a correlation between the first object distance and the first imaging device, and the second object distance and the second imaging device; A machine learning method for performing learning processing by setting a learning data set in which the first captured image is used as teacher data and the simulation image is used as training data.
7. The machine learning method of claim 6 further comprises performing machine learning using a simulation image as training data that imitates a third captured image captured by the second imaging device at a third subject distance based on the conversion table.
8. receiving a first captured image captured by a first imaging device at a first subject distance; a simulation image that simulates a second captured image captured at a second subject distance by a second imaging device that generates a captured image of lower image quality than the first captured image, generating the simulation image, in which image quality of the first captured image is corrected, using the first captured image based on a conversion table that defines a correction amount based on a correlation between the first object distance and the first imaging device, and the second object distance and the second imaging device; a machine learning program that sets a learning data set using the first captured image as teacher data and the simulation image as training data, and performs learning processing.
9. A distance information acquisition unit that acquires a subject distance of a processing target image that is captured by a first imaging device and that is to undergo image quality improvement processing; a trained model selection unit that selects a trained model from a plurality of trained models that have been trained using teacher data and training data captured by a second image capture device, and in which the teacher data and the training data used in the training process have different subject distances, based on a conversion table that defines a correction amount based on a correlation between a first subject distance and the first image capture device, and a second subject distance and the second image capture device; an image processing unit that acquires the processing target image, and performs the image quality improvement process on the processing target image using the trained model selected by the trained model selection unit to generate a high-quality inference image.