Medical image processing device, operating method of medical image processing device, medical image processing program, and recording medium

The medical image processing device enhances endoscopic image analysis by using distance information to select appropriate estimation modes and weighted averaging, addressing shooting distance variations and improving lesion detection and classification accuracy.

JP7805187B2Active Publication Date: 2026-01-23FUJIFILM CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022014091
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-01
Publication Date
2026-01-23
Estimated Expiration
2042-02-01

AI Technical Summary

Technical Problem

The estimation performance of the state of an observation region from an endoscopic image is degraded due to variations in shooting distance, affecting the accuracy of lesion detection and classification.

Method used

A medical image processing device that utilizes distance information from the endoscope to select appropriate estimation modes and performs weighted averaging of estimation results using learning models tailored for different shooting distances, incorporating normal and special light images to enhance accuracy.

Benefits of technology

Accurately estimates the state of an observation region regardless of shooting distance, improving lesion detection and classification in endoscopic images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007805187000001
    Figure 0007805187000001
  • Figure 0007805187000002
    Figure 0007805187000002
  • Figure 0007805187000003
    Figure 0007805187000003
Patent Text Reader

Abstract

To provide a medical image processing device, an operation method for a medical image processing device, a medical image processing program, and a recording medium that are capable of accurately estimating a state of an observation area of an observation image captured by an endoscope.SOLUTION: In a medical image processing device with a processor 22, an image acquisition unit 110 of the processor 22 acquires an observation image 100 of an observation area in a human body captured by an endoscope. A distance information acquisition unit 112 of the processor 22 acquires (estimates) distance information for a distance between the endoscope and the observation area from the observation image 100. A state estimation unit 114 of the processor 22 estimates a state of the observation area based on the observation image 100 and the distance information.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a medical image processing device, an operating method for a medical image processing device, a medical image processing program, and a recording medium, and in particular to a technology for accurately estimating the state of an observation area inside a body from an observation image of the observation area captured by an endoscope. [Background technology]

[0002] In the past, to prevent overlooking lesions during endoscopic examinations, lesions have been automatically detected and reported using AI (Artificial Intelligence). There are also discrimination AIs that can determine whether a detected lesion is truly a lesion.

[0003] Patent Document 1 proposes an endoscope processor that has a high ability to detect lesions.

[0004] The endoscopic processor described in Patent Document 1 generates a first processed image and a second processed image by applying different image processing to an image captured by an endoscope, and uses a learning model to which the generated first processed image and second processed image are input, and outputs the state of the disease, which is the estimated result of the learning model.

[0005] Furthermore, Patent Document 1 describes that a learning model is generated for each disease to be diagnosed, or for each part to be diagnosed. A user selects and instructs which disease or part to be diagnosed, and the endoscope processor outputs the state of the disease using the learning model corresponding to the user's selection instruction. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Publication No. 2020-156903 Summary of the Invention [Problem to be solved by the invention]

[0007] When estimating the state of an observation region (e.g., the presence or absence of a lesion) from an observation image of the observation region captured by an endoscope, the estimation performance of the state of the observation region may be degraded depending on the shooting distance of the observation image. This is thought to be because the appearance of the observation image differs depending on the shooting distance.

[0008] The present invention has been made in consideration of the above circumstances, and aims to provide a medical image processing device, an operating method for a medical image processing device, a medical image processing program, and a recording medium that can accurately estimate the state of the observation area of ​​an observation image captured by an endoscope. [Means for solving the problem]

[0009] In order to achieve the above object, the invention of a first aspect is a medical image processing device equipped with a processor, in which the processor acquires an observation image of an observation area inside the body photographed by an endoscope, acquires distance information regarding the distance from the endoscope to the observation area, and estimates the state of the observation area based on the observation image and the distance information.

[0010] According to the first aspect of the present invention, when estimating the state of an observation image from an observation area inside the body captured by an endoscope, distance information regarding the distance from the endoscope to the observation area is also used, thereby making it possible to accurately estimate the state of the observation area regardless of the shooting distance of the observation image.

[0011] In the medical image processing device according to the second aspect of the present invention, it is preferable that the processor selects one of a plurality of estimation modes for estimating the state of the observation area based on the acquired distance information, and estimates the state of the observation area using the selected estimation mode.

[0012] In the medical image processing device according to the third aspect of the present invention, it is preferable that the processor performs a weighted average of the estimation results of the state of the observation area, which are respectively estimated using a plurality of estimation modes for estimating the state of the observation area, according to the acquired distance information, to obtain the final output.

[0013] This allows the estimation results of the state of the observation area estimated using multiple estimation modes to be weighted averaged using optimal weighting coefficients, and the state of the observation area can be accurately estimated from this weighted average.

[0014] In the medical image processing device according to the fourth aspect of the present invention, the plurality of estimation modes preferably include a foreground mode for estimating the state of the observation area based on the observation image, and a background mode for estimating the state of the observation area based on the observation image, the background mode being more accurate in estimating the state of the observation area than the foreground mode when distance information exceeds a threshold. Since the appearance of an observation image varies depending on the distance, it is preferable to apply an estimation mode (foreground mode, background mode) appropriate for the distance of the observation image.

[0015] In a medical image processing device according to a fifth aspect of the present invention, it is preferable that the device has a first learning model that inputs an observation image and estimates the state of the observation area corresponding to a close-up mode, and a second learning model that inputs an observation image and estimates the state of the observation area corresponding to a distant view mode, and that the processor estimates the state of the observation area using at least one of the first learning model and the second learning model.

[0016] In the medical image processing device according to the sixth aspect of the present invention, it is preferable that the multiple estimation modes include a detection mode for detecting lesions present in the observation area based on the observation image, and a differentiation mode for classifying lesions present in the observation area into two or more classes based on the observation image.

[0017] In the medical image processing device according to the seventh aspect of the present invention, the observation images preferably include normal light images captured using normal light and special light images captured using special light, and the processor preferably acquires distance information of the observation images as well as an observation mode indicating whether the observation image is a normal light image or a special light image, and selects the detection mode or the discrimination mode based on the distance information and the observation mode. Because the normal light image and the special light image look different, it is preferable that information on these images also be used to select the detection mode or the discrimination mode.

[0018] In the medical image processing apparatus according to the eighth aspect of the present invention, it is preferable that the processor selects the discrimination mode when the distance information is equal to or less than a threshold and the observation mode is a special light observation mode for observing a special light image.

[0019] In the medical image processing device according to the ninth aspect of the present invention, it is preferable that the processor acquires distance information corresponding to each of a plurality of small regions in the observation region, and estimates the state of each of the small regions in the observation region based on the observation image and the distance information corresponding to the plurality of small regions. The small region may be one pixel of the observation image, or may be an area of ​​multiple pixels.

[0020] In a medical image processing device according to a tenth aspect of the present invention, it is preferable that the device has a third learning model that inputs an observation image and estimates distance information, and the processor inputs the acquired observation image to the third learning model and acquires the distance information estimated by the third learning model. Since the distance information is acquired from the observation image, physical distance measurement using a laser beam for distance measurement or the like is not required, and distance information of the observation area in the observation image can be acquired even in an endoscope or the like for which it is difficult to install a new measuring instrument.

[0021] In the medical image processing device according to the 11th aspect of the present invention, it is preferable that the device has a fourth learning model that inputs at least one of a distance map indicating distance information of the observation area of ​​the observation image and the observation image, and outputs a weighting coefficient to be used in calculating the weighted average, and that the processor inputs at least one of the distance map and the observation image to the fourth learning model and obtains the weighting coefficient to be used in calculating the weighted average from the fourth learning model.

[0022] In the medical image processing apparatus according to the twelfth aspect of the present invention, it is preferable that the processor displays the estimated state of the observation region on a display device that displays the observation image.

[0023] In the medical image processing device according to the thirteenth aspect of the present invention, it is preferable that the processor detects lesions present in the observation area as the state of the observation area, classifies the lesions present in the observation area into two or more classes, recognizes treatment tools present in the observation area, or recognizes organs or parts present in the observation area.

[0024] The invention according to a fourteenth aspect is a method for operating a medical image processing device equipped with a processor, the method including the steps of: the processor acquiring an observation image of an observation area inside the body taken by an endoscope; the processor acquiring distance information relating to the distance from the endoscope to the observation area; and the processor estimating the state of the observation area based on the observation image and the distance information.

[0025] In the operating method of a medical image processing device according to the 15th aspect of the present invention, it is preferable that the step of estimating the state of the observation area includes the steps of selecting one of a plurality of estimation modes for estimating the state of the observation area based on the acquired distance information, and estimating the state of the observation area using the selected estimation mode.

[0026] In the operating method of a medical image processing device relating to the 16th aspect of the present invention, it is preferable that the step of acquiring distance information acquires distance information corresponding to each of a plurality of small areas in the observation area, and the step of estimating the state of the observation area estimates the state of each of the plurality of small areas in the observation area based on the observation image and the distance information corresponding to the plurality of small areas.

[0027] A seventeenth aspect of the invention is a medical image processing program that causes a computer to execute the method for operating a medical image processing apparatus according to any one of the fourteenth to sixteenth aspects.

[0028] An eighteenth aspect of the invention is a non-transitory computer-readable recording medium on which the medical image processing program according to the seventeenth aspect is recorded. [Effects of the Invention]

[0029] According to the present invention, it is possible to accurately estimate the state of an observation region from an observation image of the observation region inside a body taken by an endoscope, regardless of the photographing distance of the observation image. [Brief explanation of the drawings]

[0030] [Figure 1] FIG. 1 is a system configuration diagram of an endoscope system including a processor device that functions as a medical image processing device according to the present invention. [Figure 2] FIG. 2 is a block diagram showing an embodiment of the hardware configuration of a processor device that constitutes the endoscope system shown in FIG. [Figure 3] FIG. 3 is a functional block diagram showing a first embodiment of a main processor of the processor device shown in FIG. [Figure 4] FIG. 4 is a diagram showing an example of an observed image and a distance map acquired from the observed image. [Figure 5] FIG. 5 is a diagram showing an example of observation images of a distant view, a medium-close view, and a close view. [Figure 6] FIG. 6 is a block diagram illustrating a specific example of the processor shown in FIG. [Figure 7] FIG. 7 shows an example of a distant view observation image to which the distant view mode is applied and a foreground view observation image to which the foreground view mode is applied. [Figure 8] FIG. 8 is a diagram showing an example of an observation image to which mode switching is applied and a distance map showing distance information of the observation region of the observation image. [Figure 9] FIG. 9 is a diagram showing the relationship between the switching pattern between the foreground mode and the background mode and the estimation accuracy. [Figure 10] FIG. 10 is a diagram showing an example of a user interface when the user sets a switching pattern between the foreground mode and the background mode. [Figure 11] FIG. 11 is a diagram showing an example of weighting coefficients (near view ratio, distant view ratio) used for calculating the weighted average of the lesion area estimation result obtained in the near view mode and the lesion area estimation result obtained in the distant view mode. [Figure 12] FIG. 12 is a block diagram showing another example of the processor shown in FIG. [Figure 13] FIG. 13 is a functional block diagram showing a second embodiment of the main processor of the processor device shown in FIG. [Figure 14] FIG. 14 is a flowchart showing an embodiment of a method for operating a medical image processing apparatus according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0031] Hereinafter, preferred embodiments of a medical image processing apparatus, an operating method of a medical image processing apparatus, a medical image processing program, and a recording medium according to the present invention will be described with reference to the accompanying drawings.

[0032] [System Configuration] FIG. 1 is a system configuration diagram of an endoscope system including a processor device that functions as a medical image processing device according to the present invention.

[0033] In FIG. 1, the endoscope system 1 includes an endoscope 10, a processor device 20, a light source device 30, and a display device 40.

[0034] The endoscope 10 is also called an electronic endoscope or an endoscope. The endoscope 10 also includes a laparoscope.

[0035] The endoscope 10 captures an image of an observation region inside the body of a subject and acquires an endoscopic image (hereinafter referred to as an "observation image"), which is a medical image. The tip of the endoscope 10 has a built-in optical system (objective lens), an image sensor, etc., and image light from the observation region is incident on the image sensor via the objective lens. The image sensor converts the image light of the observation region incident on its imaging surface into an electrical signal and outputs an image signal representing the observation image.

[0036] The rear end of the endoscope 10 is provided with a video connector and a light guide connector that connect the endoscope 10 to the processor device 20 and the light source device 30. By attaching the video connector provided on the endoscope 10 to the processor device 20, an image signal representing an observation image captured by the endoscope 10 is sent to the processor device 20. Furthermore, by attaching the light guide connector provided on the endoscope 10 to the light source device 30, illumination light emitted from the light source device 30 is irradiated towards the observation area from an illumination window on the distal end surface of the endoscope 10 via the light guide connector and a light guide disposed inside the endoscope 10.

[0037] The light source device 30 supplies illumination light to the light guide of the endoscope 10 via the light guide connector attached to the endoscope 10. Depending on the observation mode selected by the user (for example, normal light observation mode, special light observation mode, etc.), the light source device 30 can emit normal light (white broadband light or multiple broadband lights with different wavelength bands) or one or multiple special lights (specific narrowband light, or light of various wavelength bands depending on the observation purpose, such as a combination of these).

[0038] When normal light is emitted from the light source device 30, the endoscope 10 can capture a normal light image (WL (White Light) image), and when special light is emitted from the light source device 30, the endoscope 10 can capture a special light image (BLI (Blue Light Imaging or Blue LASER Imaging) image). image, Alternatively, a Linked Color Imaging (LCI) image can be taken.

[0039] The special light used for BLI is an observation light with a high proportion of V (Violet) light, which is highly absorbed by superficial blood vessels, and a low proportion of G (Green) light, which is highly absorbed by middle-layer blood vessels. This makes it suitable for generating images (BLI images) that highlight the blood vessels and structures on the surface of the mucous membrane of the subject.

[0040] In addition, the special light used for LCI has a higher proportion of V light than the observation light used for WL, making it more suitable for capturing subtle changes in color tone than the observation light used for WL. LCI images are images that have been color-enhanced by also using R (Red) component signals to make reddish colors redder and whitish colors whiter, focusing on colors near the mucous membrane.

[0041] [Processor device] FIG. 2 is a block diagram showing an embodiment of the hardware configuration of a processor device that constitutes the endoscope system shown in FIG.

[0042] The processor device 20 shown in FIG. 2 is composed of an image acquisition unit 21, a processor 22, a memory 23, a display control unit 24, an input / output interface 25, and an operation unit 26.

[0043] The image acquisition unit 21 includes a connector to which the video connector of the endoscope 10 is connected, and acquires, via the connector, from the endoscope 10, an observation image captured by an image sensor disposed at the tip of the endoscope 10. The processor device 20 also acquires, via the connector to which the endoscope 10 is connected, a remote signal generated by an operation on a handheld operation unit of the endoscope 10. The remote signal includes a release signal that commands still image capture, an observation mode signal that indicates the observation mode, and the like.

[0044] The observed image may be a moving image or a still image captured in synchronization with a release signal.

[0045] The processor 22 is composed of a CPU (Central Processing Unit) and the like, and controls each part of the processor device 20 in an integrated manner, and also functions as a processing unit that performs image processing of the observation image acquired from the endoscope 10, AI (Artificial Intelligence) processing to estimate the state of the observation area from the observation image, processing to acquire distance information regarding the distance from the endoscope 10 to the observation area, and acquisition and storage processing of still images using a release signal acquired via the endoscope 10.

[0046] The memory 23 includes flash memory, ROM (Read-only Memory), RAM (Random Access Memory), a hard disk drive, etc. The flash memory, ROM, or hard disk drive is a non-volatile memory that stores various programs and the like executed by the processor 22. The RAM functions as a working area for processing by the processor 22, and also temporarily stores programs and the like stored in the flash memory, etc. The processor 22 may have a portion of the memory 23 (RAM) built in. Still images captured during endoscopic examination can also be saved in the memory 23.

[0047] The display control unit 24 generates a display image based on the observation image (video or still image) after image processing applied from the processor 22 and the estimated result of the state of the observation area processed by AI by the processor 22, and outputs the display image to the display device 40.

[0048] The input / output interface 25 includes a connection unit for wired and / or wireless connection to an external device, a communication unit connectable to a network, etc. By transmitting an observation image to an external device such as a personal computer connected via the input / output interface 25, the external device may be provided with some or all of the functions of the medical image processing apparatus according to the present invention.

[0049] A foot switch (not shown) is also connected to the input / output interface 25. The foot switch is an operation device placed at the surgeon's feet and operated by the foot, and an operation signal is sent to the processor device 20 by stepping on a pedal. The processor device 20 is connected to storage (not shown) via the input / output interface 25. The storage (not shown) is an external storage device connected to the processor device 20 via a LAN (Local Area Network) or the like, and is, for example, a file server of a system for filing medical images such as a PACS (Picture Archiving and Communication System), or a NAS (Network Attached Storage).

[0050] The operation unit 26 includes a power switch, switches for manually adjusting white balance, light intensity, zooming, etc., and a switch for setting an observation mode.

[0051] [First embodiment of the processor] FIG. 3 is a functional block diagram showing a first embodiment of a main processor of the processor device shown in FIG.

[0052] As shown in FIG. 3, the processor 22 includes an image acquisition unit 110, a distance information acquisition unit 112, and an observation area state estimation unit 114.

[0053] The image acquisition unit 110 is a part that acquires the observation image 100 acquired from the endoscope 10 by the image acquisition unit 21 of the processor device 20. The image acquisition unit 21 of the processor device 20 is a hardware part that acquires the observation image 100 from the endoscope 10, and the image acquisition unit 110 of the processor 22 is a software part that acquires the observation image 100 to be processed in the subsequent distance information acquisition unit 112 and state estimation unit 114. Therefore, the image acquisition unit 110 also includes cases where the observation image 100 that has been temporarily stored in the memory 23 is acquired from the memory 23.

[0054] The distance information acquisition unit 112 is a part that inputs the observation image 100 acquired by the image acquisition unit 110 and acquires distance information relating to the distance from the endoscope 10 (tip) to the observation region based on the input observation image 100.

[0055] The distance information acquisition unit 112 of this example uses a learning model (third learning model) that estimates distance information, inputs the observed image 100 to the third learning model, and acquires the distance information estimated by the third learning model.

[0056] FIG. 4 is a diagram showing an example of an observed image and a distance map acquired from the observed image.

[0057] 4A of Figure 4 shows an observed image input to the fourth learning model used by the distance information acquisition unit 112, and 4B of Figure 4 shows an example of a distance map showing the distance estimation result estimated by the fourth learning model.

[0058] In the distance map shown in FIG. 4B, pixel values ​​are assigned such that the farther the distance from the endoscope 10 to the observation region, the darker the pixel value, and the closer the distance, the whiter the pixel value.

[0059] The third learning model used by the distance information acquisition unit 112 can be a trained generator that has been trained in an unsupervised manner using a large number of observed images in a Generative Adversarial Network (GAN) having a generator and a discriminator. Note that the fourth learning model is not limited to a generator trained by a GAN, and may be one trained by, for example, a Variational Auto Encoder (VAE).

[0060] The observation area state estimation unit 114 receives the observation image 100 acquired by the image acquisition unit 110 and the distance information acquired (estimated) from the observation image 100 by the distance information acquisition unit 112, and estimates the state of the observation area based on the observation image 100 and the distance information.

[0061] FIG. 5 is a diagram showing an example of observation images of a distant view, a medium-close view, and a close view.

[0062] The mucosa of the observation area observed with an endoscope may look very different depending on the distance, even if it is the same area.

[0063] Figure 5 shows an example of how a lesion looks when observed from a distance, a medium distance, and a close distance. The characteristics of the lesion at each distance are as follows:

[0064] 5A in Fig. 5 is a distant view image. From this distant view image, an oval-shaped lesion can be observed that is more reddish than the surrounding area (high-density area in Fig. 5A).

[0065] 5B of Fig. 5 is a medium-to-close observation image, which is an image of a circular area surrounding the lesion included in the distant observation image of 5A of Fig. 5. From this medium-to-close observation image, a strong redness compared to the surrounding area (high-density areas in 5B of Fig. 5) and irregularities in the mucosal pattern can be observed.

[0066] 5C in Fig. 5 is a close-up image, which is an image of a circular area surrounding the lesion included in the medium-close-up image in 5B in Fig. 5. From this close-up image, irregularities in the mucosal pattern can be observed throughout the image.

[0067] As such, even the same lesion will exhibit different characteristics depending on the distance at which it is photographed, so it may be inappropriate to use the same AI to estimate the state of the observation area, including the detection and differentiation of lesions, at all distances.

[0068] Therefore, the observation area state estimation unit 114 selects one of multiple estimation modes for estimating the state of the observation area based on the distance information acquired from the distance information acquisition unit 112, and estimates the state of the observation area using the selected estimation mode, or weights and averages the estimation results of the state of the observation area estimated by each of the multiple estimation modes for estimating the state of the observation area according to the acquired distance information to obtain the final output.

[0069] Here, "estimating the state of the observation area" includes detecting (estimating) the presence or absence of a lesion in the observation area and / or the lesion area, classifying lesions in the observation area into two or more classes (e.g., neoplastic, non-neoplastic), recognizing a treatment tool in the observation area (recognizing the presence or absence of a treatment tool and / or the type of treatment tool), or recognizing an organ or part in the observation area. When the endoscope 10 is an upper gastrointestinal endoscope, recognizing an organ in the observation area means detecting (classifying) organs such as the esophagus, stomach, and duodenum. Furthermore, recognizing a part in the observation area means, for example, detecting each part of the stomach, such as the cardia, fundus, upper body, middle body, lower body, antrum, pyloric antrum, and pylorus, if the organ is the stomach.

[0070] In the following, in this example, a case where a lesion area present in the observation area is detected (estimated) will be described as "estimating the state of the observation area."

[0071] <Specific example of the processor according to the first embodiment> FIG. 6 is a block diagram illustrating a specific example of the processor shown in FIG.

[0072] 6, an observed image 100 is input to the processor 22. The processor 22 reads the input observed image 100 and the learning model 23A and necessary parameter groups from the learning model 23A, the first parameter group 23B, the second parameter group 23C, the third parameter group 23D, the fourth parameter group 23E, ..., which are stored in advance in the memory 23, and functions as the distance information acquisition unit 112 and the observation area state estimation unit 114 shown in FIG.

[0073] As the learning model 23A, one or more known learning models can be applied, and as the learning model used as the state estimation unit 114, for example, a convolution neural network (CNN) can be applied.

[0074] Furthermore, the first parameter group 23B, the second parameter group 23C, the third parameter group 23D, the fourth parameter group 23E, etc. are applied to one or more learning models 23A, and are, for example, filter coefficients of a filter used in the convolution layer of the CNN, weighting coefficients for the input of each layer of the CNN, etc.

[0075] The processor 22 applies a third group of parameters to the learning model 23A read from the memory 23 to use it as a trained third learning model, and by inputting the observed image 100 into this third learning model, it estimates and outputs distance information indicating the distance of the observation area.

[0076] In this example, the multiple estimation modes for estimating the lesion area include a foreground mode and a background mode, where the background mode is a mode that estimates the lesion area more accurately than the foreground mode when distance information exceeds a certain threshold.

[0077] When estimating a lesion area in the close-up mode, the processor 22 applies a first group of parameters to the learning model 23A read from the memory 23, thereby using it as a learning model (first learning model) for estimating a lesion area in the close-up observation image 100 having distance information less than a threshold, and estimates the lesion area in the close-up observation image 100 by inputting the observation image 100 into this first learning model.

[0078] On the other hand, when estimating a lesion area in distant view mode, the processor 22 applies a second parameter group to the same learning model 23A read from the memory 23, and uses it as a learning model (second learning model) for estimating a lesion area in a distant view observation image 100 having distance information exceeding a threshold, and estimates a lesion area in the distant view observation image 100 by inputting the observation image 100 into this second learning model.

[0079] The first parameter group is the observation image of the close view (Fig. 5 5C) and ground truth data (ground truth mask) showing the lesion area. Similarly, the second parameter group can be obtained by machine learning the untrained learning model 23A using a large number of training data sets that pair the ground truth data (ground truth mask) showing the lesion area. 5 The data can be acquired by machine learning an untrained learning model 23A using a large number of training data sets in which the training data are paired with the target data indicating the lesion area (see 5A of FIG. 1).

[0080] FIG. 7 shows an example of a distant view observation image to which the distant view mode is applied and a foreground view observation image to which the foreground view mode is applied.

[0081] When the processor 22 determines that the input observation image 100 is a close-up observation image 100 having distance information less than a threshold value based on the output of the third learning model that estimates distance information (particularly when the distance is uniform across the entire image as shown in 7B of Figure 7), it selects the close-up mode as the estimation mode for estimating the lesion area, and estimates the lesion area in the close-up observation image 100 using the first learning model corresponding to the selected close-up mode.

[0082] In addition, when the processor 22 determines that the input observation image 100 is a distant observation image 100 having distance information exceeding a threshold value based on the output of the third learning model that estimates distance information (particularly when the distance is uniform across the entire image as shown in 7A of Figure 7), it selects the distant view mode as the estimation mode for estimating the lesion area, and estimates the lesion area in the distant view observation image 100 using the second learning model corresponding to the selected distant view mode.

[0083] When the close-up mode is selected, the processor 22 uses a first learning model corresponding to the close-up mode to estimate the lesion area in the close-up observation image 100 and outputs the estimated result; when the distant-view mode is selected, the processor 22 uses a second learning model corresponding to the distant-view mode (in this example, the parameter group used for the learning model 23A is switched from the first parameter group 23B to the second parameter group 23C) to estimate the lesion area in the distant-view observation image 100 and outputs the estimated result.

[0084] The processor 22 is not limited to switching the estimation mode (learning model to be used) depending on whether the observed image is a close-up or a distant view, as described above, and outputting the estimated result of the lesion area. The processor 22 may also take a weighted average of the estimated results of the lesion area estimated using a plurality of estimation modes (in this example, the close-up mode and the distant view mode) depending on the distance information of the observed image 100, and output the weighted average as the final output.

[0085] FIG. 8 is a diagram showing an example of an observation image to which mode switching is applied and a distance map showing distance information of the observation region of the observation image.

[0086] The observation image shown in 8A of Fig. 8 includes a distant observation area in the upper left and a close-up observation area in the lower right. Note that the distance map shown in 8A of Fig. 8 assigns pixel values ​​that are darker the farther the distance from the endoscope to the observation area, and whiter the closer the distance is.

[0087] In the case of such an observation image, it is preferable to switch the estimation mode to be applied within the observation image between the foreground mode and the background mode, rather than applying the foreground mode or the background mode to the entire image.

[0088] In this example, processor 22 applies the distant view mode to the upper left area of ​​the observation image shown in 8A of FIG. 8, and the close view mode to the other areas, and integrates and outputs the estimation results of the lesion area estimated using both estimation modes.

[0089] The area of ​​the observation image to which the foreground mode or the background mode is applied includes one or more outlying areas depending on the distance information of the observation area.

[0090] FIG. 9 is a diagram showing the relationship between the switching pattern between the foreground mode and the background mode and the estimation accuracy.

[0091] FIG. 9C is a distance map showing the distance from the endoscope to the observation region estimated from the observation image.

[0092] Figure 9A shows a case in which a certain threshold (first threshold) is set for the distance information indicated by the distance map shown in Figure 9C, and a close-up mode is applied to areas of the entire observation image having distance information less than the first threshold, and a distant-view mode is applied to areas having distance information greater than the first threshold, thereby estimating the lesion area.

[0093] According to the switching pattern between the foreground mode and the background mode shown in 9A of FIG. 9, the area to which the background mode is applied is larger than the area to which the foreground mode is applied.

[0094] On the other hand, Figure 9B shows a case in which a threshold value (second threshold value) smaller than the first threshold value is set for the distance information indicated by the distance map shown in Figure 9C, and the foreground mode is applied to areas of the entire observation image having distance information less than the second threshold value, and the background mode is applied to areas having distance information exceeding the second threshold value, thereby estimating the lesion area.

[0095] According to the switching pattern between the foreground mode and the background mode shown in 9B of FIG. 9, the area to which the background mode is applied is smaller than the area to which the foreground mode is applied.

[0096] Furthermore, the estimation result of the lesion area when the switching pattern shown in 9A of FIG. 9 is applied is more accurate than the estimation result of the lesion area when the switching pattern shown in 9B of FIG. 9 is applied.

[0097] Therefore, in the case of the observation image of this example, it is preferable to apply the switching pattern shown in 9A of FIG. 9 (that is, to set an area to which the foreground mode is applied and an area to which the background mode is applied using the first threshold).

[0098] To estimate the lesion area more accurately, it is necessary to set an optimal switching pattern (i.e., an optimal threshold for dividing the area to which the close-up mode and the distant-view mode are applied). The optimal threshold can be determined in advance by trial and error, or it can be automatically estimated for each observed image using a learning model and the estimated threshold can be used.

[0099] In this case, the learning model can be generated by machine learning an untrained learning model using a large number of training data sets that pair observed images with thresholds (ground truth data) that produce highly accurate lesion area estimates for those observed images.

[0100] FIG. 10 is a diagram showing an example of a user interface when the user sets a switching pattern between the foreground mode and the background mode.

[0101] FIG. 10A is a diagram showing an example of a distance map showing distance information of an observation area of ​​an observation image.

[0102] As shown in 10B and 10C of FIG. 10, the processor 22 causes the display device 40 to display a distance map, and also causes an operation screen having a slider (triangle icon) for setting a threshold value adjacent to the distance map.

[0103] The user can set threshold value A shown in 10B of FIG. 10 and threshold value B shown in 10C of FIG. 10 by moving the slider to a position of any density (distance) using the operation unit 26 or a touch panel (not shown).

[0104] Also, as shown in 10B and 10C of Figure 10, when threshold A or threshold B is set by the user, processor 22 displays areas on the distance map to which the foreground mode is applied and areas to which the background mode is applied in a distinguishable manner by coloring or the like according to threshold A or threshold B.

[0105] This allows the user to check the switching pattern between the foreground mode and the background mode according to the threshold value, and to set the desired switching pattern.

[0106] Furthermore, in the above example, the area of ​​the observed image to which the close-up mode is applied and the area to which the distant view mode is applied are completely separated, and a learning model corresponding to each mode is used to obtain the estimated result of the lesion area. However, this is not limited to this, and the estimated result of the lesion area estimated in the close-up mode and the estimated result of the lesion area estimated in the distant view mode may be weighted and averaged according to the distance information to obtain the final output.

[0107] FIG. 11 is a diagram showing an example of weighting coefficients (near view ratio, distant view ratio) used for calculating the weighted average of the lesion area estimation result obtained in the near view mode and the lesion area estimation result obtained in the distant view mode.

[0108] 11C is a diagram showing an example of a distance map showing distance information of the observation area of ​​the observation image. The processor 22 calculates the foreground ratio and the background ratio in pixel units based on the distance map. Note that the foreground ratio and the background ratio are not limited to pixel units, and may be calculated in units of small areas larger than one pixel.

[0109] As shown in 11A of FIG. 11, the foreground ratio and the background ratio in this example each have a value in the range of 0.0 to 1.0, and the sum of the foreground ratio and the background ratio obtained for a certain pixel is 1.0.

[0110] Processor 22 uses a learning model (fourth learning model) that takes a distance map as input and outputs weighting coefficients (near view ratio, far view ratio) used to calculate the weighted average, inputs the distance map into the fourth learning model, and can obtain the near view ratio and far view ratio for each pixel of the observed image from the fourth learning model.

[0111] 11B of FIG. 11 is an image diagram showing the near view ratio and the far view ratio for each pixel of the observed image acquired based on the distance map of 11C of FIG.

[0112] In Figure 6, processor 22 applies a fourth parameter group to learning model 23A read from memory 23 to use it as a trained fourth learning model, inputs a distance map into the fourth learning model, and can obtain the foreground ratio and background ratio for each pixel of the observed image estimated by the fourth learning model.

[0113] The fourth learning model may be generated so that when an observation image is input instead of a distance map, the near view ratio and far view ratio are estimated and output for each pixel of the observation image, or when both an observation image and a distance map are input, the fourth learning model may be generated so that when an observation image and a distance map are input, the near view ratio and far view ratio are estimated and output for each pixel of the observation image.

[0114] In addition, processor 22 may obtain the foreground ratio and background ratio based on the distance information for each pixel from a lookup table in which the foreground ratio and background ratio are stored in advance according to the distance information, without using the fourth learning model.

[0115] When processor 22 acquires the results of the lesion area estimation in the foreground mode, the results of the lesion area estimation in the background mode, and the foreground ratio and background ratio for each pixel of the observed image, it calculates a weighted average based on the following formula:

[0116] [Number 1] Weighted average = (Estimated result in distance mode) x (Distant view rate) + (Estimated result from close-up mode) × (close-up rate) The processor 22 displays the lesion area estimation result (state of the observation area) obtained from the observation image 100 as described above on the display device 40 that displays the observation image 100. For example, when a lesion area is detected from the observation image 100, the lesion area estimation result can be displayed on the display device 40 by superimposing a bounding box or the like surrounding the lesion area on the observation image 100.

[0117] <Another specific example of the processor according to the first embodiment> FIG. 12 is a block diagram showing another example of the processor shown in FIG.

[0118] The processor 22 shown in Figure 6 reads out the learning model 23A from the memory 23 and applies the first parameter group 23B, the second parameter group 23C, the third parameter group 23D, the fourth parameter group 23E, ... to the learning model 23A to acquire distance information, a weighting coefficient acquisition unit acquires a weighting coefficient to be used for the weighted average of the estimation result of the lesion area estimated in the foreground mode and the estimation result of the lesion area estimated in the background mode, and switches between the first parameter group 23B and the second parameter group 23C to operate the state estimation unit in the foreground mode or the background mode. However, the processor 22 shown in Figure 12 differs from the processor 22 shown in Figure 6 in that it separately includes the distance information acquisition unit 120, the weighting coefficient acquisition unit 122, the first state estimation unit 124, the second state estimation unit 126, ...

[0119] Here, the first state estimation unit 124 and the second state estimation unit 126 use learning models (first learning model, second learning model) to which the first parameter group 23B and the second parameter group 23C shown in Fig. 6 are applied, respectively, and input the observed image 100 into the first learning model and the second learning model, thereby acquiring estimation results of the lesion area from the first learning model and the second learning model. Here, the first learning model can output estimation results of the lesion area with high detection accuracy for a close-up observation area for the input observation image 100, and the second learning model can output estimation results of the lesion area with high detection accuracy for a distant observation area for the input observation image 100.

[0120] Processor 22 selects and outputs either one of the lesion area estimation results estimated by first state estimation unit 124 and second state estimation unit 126 for the entire observation image 100 based on the distance information acquired by distance information acquisition unit 120, or integrates and outputs the lesion area estimation results selected on a pixel-by-pixel or small area-by-small area basis, or performs a weighted average of the lesion area estimation results estimated by first state estimation unit 124 and second state estimation unit 126 according to the distance information of observation image 100 to obtain the final output.

[0121] [Second embodiment of the processor] FIG. 13 is a functional block diagram showing a second embodiment of the main processor of the processor device shown in FIG.

[0122] As shown in FIG. 13, the processor 22 includes a distance information acquisition unit 130, an observation mode acquisition unit 132, a detection mode / discrimination mode selection unit 134, a first state estimation unit 136, and a second state estimation unit 138.

[0123] The distance information acquisition unit 130 is a part that inputs the observation image 100 and acquires distance information regarding the distance from the endoscope to the observation area based on the input observation image 100, and outputs the acquired distance information to the detection mode / discrimination mode selection unit 134.

[0124] The observation mode acquisition unit 132 is a part that acquires an observation mode indicating whether the observation image 100 is a normal light image or a special light image, and outputs an observation mode signal indicating the acquired current observation mode to the detection mode / discrimination mode selection unit 134.

[0125] The observation mode acquisition unit 132 can acquire the current observation mode by inputting an observation mode signal that indicates the observation mode, for example, a remote signal generated by a user operating a handheld operation unit of the endoscope 10.

[0126] The detection mode / discrimination mode selection unit 134 selects the detection mode or the discrimination mode based on the input distance information and observation mode.

[0127] Here, the detection mode is one of a plurality of estimation modes, which is a mode for detecting a lesion (lesion area) present in the observation area based on the observation image, and the discrimination mode is another of the plurality of estimation modes, which is a mode for classifying the lesion present in the observation area into two or more classes (e.g., neoplastic, non-neoplastic) based on the observation image. Note that the discrimination mode may include the detection mode, or may be a mode for individually classifying the lesion area detected in the detection mode.

[0128] The detection mode / discrimination mode selection unit 134 selects the discrimination mode when the input distance information is below the threshold and the observation mode is a special light observation mode in which a special light image is observed, and selects the detection mode in other cases.

[0129] When observing special light images captured using special light at a short distance where distance information is below a threshold, a user may wish to differentiate lesions rather than detect them. Accordingly, in the case of "close-distance photography and special light observation mode," it is preferable to select the differentiation mode. Furthermore, observation images captured under the "close-distance photography and special light observation mode" conditions clearly show the blood vessels and structures on the mucosal surface of the subject, making them suitable for differentiation.

[0130] The first state estimation unit 136 is a state estimation unit that operates when the detection mode is selected, and detects a lesion present in the observation area based on the input observation image 100. The first state estimation unit 136 can detect a lesion (lesion area) using a learning model for detecting lesions, as described above. The first state estimation unit 136 may also selectively obtain the results of lesion estimation estimated in the close-up mode and the distant-view mode based on distance information, or may perform a weighted average based on the distance information as the final output.

[0131] The second state estimation unit 138 is a state estimation unit that operates when the discrimination mode is selected, and classifies lesions present in the observation area into two or more classes based on the input observation image 100, and outputs the classified classes (discrimination results).

[0132] The second state estimation unit 138 uses a trained learning model that has undergone machine learning using a large number of data sets of learning data that pair special light images taken at close range with classification results (ground truth data) corresponding to those special light images, and can output classification results by inputting observed images (special light images taken at close range) into this learning model.The second state estimation unit 138 may also use a support vector machine (SVM), which is a type of machine learning model, to classify lesions.

[0133] The processor 22 selects the first state estimation unit 136 or the second state estimation unit 138 according to the detection mode or the differentiation mode selected by the detection mode / differentiation mode selection unit 134, and displays the lesion detection result or differentiation result estimated by the selected first state estimation unit 136 or second state estimation unit 138 on the display device 40 that also displays the observation image 100. This allows the user to check the lesion detection result or differentiation result on the screen of the display device 40.

[0134] [Operation method of medical image processing device] FIG. 14 is a flowchart showing an embodiment of a method for operating a medical image processing apparatus according to the present invention.

[0135] The method for operating a medical image processing device is, for example, a method for operating a medical image processing device equipped with the processor 22 shown in Figure 3, and the processor 22 executes the various processes shown below in accordance with the flowchart shown in Figure 14.

[0136] In FIG. 14, the processor 22 acquires an observation image of an observation region inside a body, captured by an endoscope, via the image acquisition unit 110 (step S10).

[0137] Next, the processor 22 acquires distance information relating to the distance from the endoscope to the observation region using the distance information acquisition unit 112 (step S12). In this example, the distance information acquisition unit 112 uses a learning model (third learning model) that estimates distance from an observed image, and acquires distance information from the third learning model by inputting the observed image to the third learning model.

[0138] The processor 22 estimates the state of the observation region (for example, a lesion region present in the observation region) based on the observation image acquired in step S10 and the distance information acquired in step S12 (step S14).

[0139] The observation area state estimation unit 114 of the processor 22 selects one of a plurality of estimation modes (for example, close-up mode, distant view mode) for estimating the lesion area based on the distance information, and estimates the lesion area using the selected estimation mode, or performs a weighted average of the estimation results of the lesion area estimated by each of the plurality of estimation modes according to the acquired distance information to provide the final output.

[0140] Furthermore, in step S12, when distance information corresponding to each of multiple small areas in the observation area including pixel units of the observation image is acquired, in step S14 of estimating the state of the observation area (lesion area), the state of each of the multiple small areas in the observation area is estimated based on the observation image and the distance information corresponding to each of the multiple small areas.

[0141] The processor 22 outputs the estimation result obtained in step S14 (step S16). The processor 22 preferably displays the estimation result on the display device 40 that displays the observed image.

[0142] Next, the processor 22 determines whether or not the user has finished observing the observation image (step S18). If the processor 22 determines that the observation has not finished, it transitions to step S10 and repeatedly executes the processes from step S10 to step S18. That is, if the observation image is a moving image, the processes from step S10 to step S18 are repeatedly executed for each frame or every few frames, thereby making it possible to acquire an estimation result for the observation image of the moving image in real time. Note that the determination of the end of observation can be made, for example, by detecting a user operation by the user to end the endoscopic examination.

[0143] When the processor 22 determines in step S18 that the observation has ended, it ends this processing.

[0144] [others] In this embodiment, distance information regarding the distance from the endoscope to the observation area is obtained from the observation image using AI, but this is not limited to this. For example, if the endoscope is equipped with a distance measurement unit that physically measures the distance between the tip of the endoscope and the object of observation using laser light or the like, the distance information measured by the distance measurement unit may be obtained.

[0145] Furthermore, in this embodiment, two estimation modes, a close-up mode and a distant view mode, have been described as multiple estimation modes that are applied depending on the distance, but this is not limited to this. For example, the state of the observation area for an observation image at a corresponding distance may be estimated using three or more estimation modes, including a medium-close-up mode that is applied to a medium-close-up view between a close-up and a distant view.

[0146] The hardware structure that executes various controls of the medical image processing apparatus according to the present invention is the following various processors: The various processors include a CPU (Central Processing Unit), which is a general-purpose processor that executes software (programs) and functions as various control units, a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), whose circuit configuration can be changed after manufacture, and a dedicated electric circuit, such as an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing specific processing.

[0147] A single processing unit may be configured with one of these various processors, or may be configured with two or more processors of the same or different types (for example, multiple FPGAs, or a combination of a CPU and an FPGA). Also, multiple control units may be configured with a single processor. Examples of multiple control units configured with a single processor include, first, a form in which one processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server, and this processor functions as multiple control units. Second, a form in which a processor is used to realize the functions of an entire system including multiple control units on a single IC (Integrated Circuit) chip, as typified by a system on chip (SoC). In this way, the various control units are configured with one or more of the above-mentioned various processors as a hardware structure.

[0148] The present invention also includes a medical image processing program that, when installed on a computer, causes the computer to function as the medical image processing device of the present invention, and a non-transitory, computer-readable recording medium on which this medical image processing program is recorded.

[0149] Furthermore, the present invention is not limited to the above-described embodiment, and it goes without saying that various modifications are possible without departing from the spirit of the present invention. [Explanation of symbols]

[0150] 1. Endoscopy system 10 Endoscopy 20 Processor unit 21 Image acquisition unit 22 processors 23 Memory 23A Learning Model 23B First parameter group 23C Second parameter group 23D Third parameter group 23E Fourth parameter group 24 Display control unit 25 Input / Output Interface 26 Control section 30 Light source device 40 Display device 100 observation images 110 Image acquisition unit 112 Distance information acquisition section 114 State Estimation Unit 120, 130 Distance information acquisition section 122 Weighting coefficient acquisition unit 124, 136 First state estimation unit 126, 138 Second state estimation unit 132 Observation mode acquisition unit 134 Detection mode / discrimination mode selection section S10 Step S12 Step S14 Step S16 Step S18 Step

Claims

1. A medical image processing device having a processor, The processor: Obtaining an observation image of an observation area inside the body using an endoscope; acquiring distance information relating to the distance from the endoscope to the observation region; estimating a state of the observation area based on the observation image and the distance information; the processor selects one of a plurality of estimation modes for estimating the state of the observation area based on the acquired distance information, and estimates the state of the observation area using the selected estimation mode. Medical imaging equipment.

2. A medical image processing device equipped with a processor, The processor: Obtaining an observation image of an observation area inside the body using an endoscope; acquiring distance information relating to the distance from the endoscope to the observation region; estimating a state of the observation area based on the observation image and the distance information; the processor performs a weighted average of the estimation results of the state of the observation area, which are estimated using a plurality of estimation modes for estimating the state of the observation area, according to the acquired distance information, to obtain a final output. Medical imaging equipment.

3. the plurality of estimation modes include a foreground mode in which a state of the observation area is estimated based on the observation image, and a background mode in which a state of the observation area is estimated based on the observation image, and when the distance information exceeds a threshold, the background mode estimates the state of the observation area with higher accuracy than the foreground mode. The medical image processing device according to claim 1 or 2.

4. a first learning model that receives the observation image and estimates a state of the observation area corresponding to the close view mode; and a second learning model that receives the observation image and estimates a state of the observation area corresponding to the distant view mode, the processor estimates a state of the observation area using at least one of the first learning model and the second learning model; The medical image processing device according to claim 3 .

5. the plurality of estimation modes include a detection mode for detecting a lesion present in the observation area based on the observation image, and a discrimination mode for classifying the lesion present in the observation area into two or more classes based on the observation image. The medical image processing device according to claim 1 or 2.

6. the observed image includes a normal light image captured using normal light and a special light image captured using special light, The processor: acquiring an observation mode indicating whether the observation image is the normal light image or the special light image together with the distance information of the observation image; selecting the detection mode or the discrimination mode based on the distance information and the observation mode; The medical image processing device according to claim 5 .

7. the processor selects the discrimination mode when the distance information is equal to or less than a threshold and the observation mode is a special light observation mode in which the special light image is observed. The medical image processing device according to claim 6 .

8. The processor: acquiring the distance information corresponding to each of a plurality of small regions in the observation region; estimating states of the plurality of small regions in the observation region based on the observation image and the distance information corresponding to the plurality of small regions; The medical image processing device according to claim 1 .

9. a third learning model that receives the observed image and estimates the distance information; the processor inputs the acquired observed image into the third learning model and acquires the distance information estimated by the third learning model. The medical image processing device according to claim 1 .

10. a fourth learning model that receives input of at least one of a distance map indicating the distance information of the observation region of the observation image and the observation image, and outputs a weighting coefficient used in calculating the weighted average; the processor inputs at least one of the distance map and the observed image into the fourth learning model, and obtains the weighting coefficients used in calculating the weighted average from the fourth learning model. The medical image processing device according to claim 2 .

11. the processor causes the estimated state of the observation area to be displayed on a display device that displays the observation image. The medical image processing device according to claim 1 .

12. the processor detects a lesion present in the observation area as the state of the observation area, classifies the lesion present in the observation area into two or more classes, recognizes a treatment tool present in the observation area, or recognizes an organ or a part present in the observation area. The medical image processing device according to claim 1 .

13. A method of operating a medical imaging device having a processor, comprising: a step of the processor acquiring an observation image of an observation region inside a body by an endoscope; the processor obtaining distance information relating to a distance from the endoscope to the observation region; the processor estimating a state of the observation area based on the observation image and the distance information; The step of estimating a state of the observation area includes: selecting one of a plurality of estimation modes for estimating a state of the observation area based on the acquired distance information; and estimating a state of the observation area using the selected estimation mode. A method of operating a medical imaging device.

14. A method of operating a medical imaging device having a processor, comprising: a step of the processor acquiring an observation image of an observation region inside a body by an endoscope; the processor obtaining distance information relating to a distance from the endoscope to the observation region; the processor estimating a state of the observation area based on the observation image and the distance information; The step of estimating a state of the observation area includes: a weighted average of the estimation results of the state of the observation area, which are estimated by a plurality of estimation modes for estimating the state of the observation area, in accordance with the acquired distance information, is obtained as a final output; A method of operating a medical imaging device.

15. The step of acquiring distance information includes acquiring the distance information corresponding to each of a plurality of small regions in the observation region, the step of estimating the state of the observation area includes estimating states of the plurality of small areas in the observation area based on the observation image and the distance information corresponding to the plurality of small areas; A method for operating a medical image processing apparatus according to claim 13 or 14.

16. A medical image processing program that causes a computer to execute the method for operating a medical image processing apparatus according to any one of claims 13 to 15.

17. A non-transitory computer-readable recording medium on which the medical image processing program according to claim 16 is recorded.

Citation Information

Patent Citations

  • Image processor, endoscope device, program and image processing method

    JP2014188222A

  • Processor for endoscopes, information processing unit, program, information processing method and learning model generation method

    JP2020156903A

  • Endoscope system, processor device, diagnosis assistance method, and computer program

    WO2021141048A1

  • Medical image processing device, endoscope system, medical image processing method, and program

    WO2021157487A1

  • Endoscope insertion assistance device, method, and non-temporary computer-readable medium having program stored therein

    WO2021205818A1