Medical image processing apparatus, hierarchical neural network, medical image processing method, and program
The hierarchical neural network addresses the challenge of high-resolution feature extraction for real-time medical image processing by using separate sub-networks for detection and classification, enhancing accuracy and efficiency in region-of-interest analysis.
Patent Information
- Application Number
- JP2024033022
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-05
- Publication Date
- 2025-09-18
AI Technical Summary
Existing medical image processing methods face challenges in achieving high accuracy and real-time processing for region-of-interest detection and classification, particularly in models designed for real-time processing where aggressive resolution reduction is necessary, leading to decreased accuracy in evaluating detailed structures.
A hierarchical neural network is employed, comprising a feature extraction network, a first sub-network for detection, and a second sub-network for classification, which extracts features at different resolutions to enhance detection and classification accuracy while allowing real-time processing.
The hierarchical neural network enables highly accurate and real-time detection and classification of regions of interest in medical images, improving the precision of lesion classification and reducing the need for redundant feature extraction processes.
Smart Images

Figure 2025135263000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a medical image processing device, a hierarchical neural network, a medical image processing method and a program, and more particularly to a technique for detecting and classifying regions of interest from medical images. [Background technology]
[0002] There is known a function that detects regions of interest such as lesions from medical images captured by medical imaging diagnostic devices such as endoscopes and ultrasound diagnostic devices, and simultaneously classifies the regions of interest into classes such as disease types. Such a function can be realized by training a neural network (NN) using images containing the regions of interest and position information and classification class information of the regions of interest (see Patent Documents 1 and 2).
[0003] During inference, feature extraction is performed on the input image by gradually reducing the resolution using neural network processing, and area detection and classification are performed from the resulting feature values. In the classification task, a portion of the feature value that corresponds to the detection area is extracted from this reduced-resolution feature value before classification processing is performed. As a result, the feature resolution of the processing target in the classification task is low, which can lead to a decrease in accuracy when it is necessary to evaluate detailed structures, such as in lesion classification. This problem is particularly pronounced in models designed for real-time processing, where aggressive reduction in resolution is necessary for high-speed processing. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] International Publication No. 2020 / 203552 [Patent Document 2] International Publication No. 2019 / 142243 Summary of the Invention [Problem to be solved by the invention]
[0005] One way to solve this problem is to extract the detection area from the input image and use the extracted image to train the classification task. However, extracting features from high-resolution input images using neural network processing is inefficient and may make real-time processing difficult.
[0006] The present invention has been made in consideration of the above circumstances, and aims to provide a medical image processing device, a hierarchical neural network, a medical image processing method, and a program that realize highly accurate, real-time processing of region-of-interest detection and class classification. [Means for solving the problem]
[0007] In order to achieve the above object, a medical image processing device according to a first aspect of the present disclosure is a medical image processing device that includes one or more processors that acquire medical images and one or more memories that store programs to be executed by the one or more processors, wherein the one or more processors process the medical images with a feature extraction network of a hierarchical neural network that includes a feature extraction network, a first sub-network, and a second sub-network, thereby extracting a first feature and a second feature having a relatively higher resolution than the first feature from the medical image, processing the first feature with the first sub-network of the hierarchical neural network to detect an area of interest contained in the medical image, and processing the second feature with the second sub-network of the hierarchical neural network to classify the area of interest.
[0008] In a medical image processing device according to a second aspect of the present disclosure, in the medical image processing device according to the first aspect, it is preferable that the second feature is an intermediate feature in the process of extracting the first feature from the medical image using a feature extraction network.
[0009] A medical image processing device according to a third aspect of the present disclosure is preferably a medical image processing device according to the first or second aspect, wherein one or more processors train a first sub-network using a first dataset and train a second sub-network using a second dataset, the first dataset including a pair of a first medical image and location information of a region of interest contained in the first medical image, and the second dataset including a pair of a second medical image and location information and a classification class label of a region of interest contained in the second medical image.
[0010] A medical image processing device according to a fourth aspect of the present disclosure is preferably a medical image processing device according to the third aspect, wherein one or more processors train a feature extraction network and a first sub-network using a first data set, and train a second sub-network using a second data set based on the trained feature extraction network and the trained first sub-network.
[0011] A medical image processing device according to a fifth aspect of the present disclosure is preferably a medical image processing device according to the fourth aspect, wherein the one or more processors pre-train the feature extraction network using a third dataset different from the first dataset and the second dataset before training the feature extraction network and the first sub-network using the first dataset.
[0012] In a medical image processing device according to a sixth aspect of the present disclosure, in a medical image processing device according to any of the first to fifth aspects, it is preferable that one or more processors notify the position information of the region of interest in a manner according to the classification result.
[0013] A medical image processing device according to a seventh aspect of the present disclosure is preferably the medical image processing device according to the sixth aspect, wherein the one or more processors add information based on the position information to the medical image and display it on the display.
[0014] In the medical image processing device according to the eighth aspect of the present disclosure, in the medical image processing device according to the sixth or seventh aspect, it is preferable that one or more processors not report the position information of the region of interest when the classification result is a specific class.
[0015] A medical image processing device according to a ninth aspect of the present disclosure is a medical image processing device according to the eighth aspect, wherein the classification class includes malignancy, and the one or more processors preferably do not report if the malignancy is relatively low.
[0016] A medical image processing device according to a tenth aspect of the present disclosure is a medical image processing device according to any one of the first to ninth aspects, wherein one or more processors preferably extract a portion of the second feature depending on the detection result of the region of interest, and process the extracted second feature in a second sub-network.
[0017] In a medical image processing device according to an eleventh aspect of the present disclosure, in the medical image processing device according to the tenth aspect, it is preferable that one or more processors align the spatial size of the extracted second feature to a fixed size and process the fixed-size second feature in a second sub-network.
[0018] In order to achieve the above object, the hierarchical neural network according to a twelfth aspect of the present disclosure is a hierarchical neural network including a feature extraction network that extracts a first feature and a second feature having a relatively higher resolution than the first feature from an input medical image, a first sub-network that detects an area of interest contained in the medical image from the input first feature, and a second sub-network that classifies the area of interest from the input second feature.
[0019] In order to achieve the above object, a medical image processing method according to a thirteenth aspect of the present disclosure is a medical image processing method executed by one or more processors, which acquires a medical image, processes the medical image with a feature extraction network of a hierarchical neural network including a feature extraction network, a first sub-network, and a second sub-network, thereby extracting a first feature and a second feature having a relatively higher resolution than the first feature from the medical image, processes the first feature with the first sub-network of the hierarchical neural network to detect an area of interest contained in the medical image, and processes the second feature with the second sub-network of the hierarchical neural network to classify the area of interest.
[0020] In order to achieve the above object, a program according to a 14th aspect of the present disclosure is a program that causes a computer to execute the medical image processing method according to the 13th aspect. A non-transitory computer-readable storage medium on which the program according to the 13th aspect is recorded is also included in the present disclosure. [Effects of the Invention]
[0021] According to the present invention, it is possible to realize highly accurate and real-time processable region of interest detection and classification. [Brief explanation of the drawings]
[0022] [Figure 1] FIG. 1 is a conceptual diagram of a network system that performs a general detection process. [Figure 2] Figure 2 is a conceptual diagram of a network system in which detection and classification are performed by separate neural networks. [Figure 3] FIG. 3 is a conceptual diagram of a network system according to the first embodiment. [Figure 4] FIG. 4 is a flow chart showing the steps of a medical image processing method. [Figure 5] FIG. 5 is a block diagram showing the configuration of a medical image processing apparatus. [Figure 6]FIG. 6 is a schematic diagram showing the overall configuration of an endoscope system including a medical image processing device. DETAILED DESCRIPTION OF THE INVENTION
[0023] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings. The same components are designated by the same reference numerals and redundant explanations will be omitted.
[0024] <General detection process> 1 is a conceptual diagram of a network system 10 that performs a general detection process. The network system 10 includes a feature extraction unit 12 and a detection processing unit 14.
[0025] The feature extraction unit 12 is a neural network (NN) that outputs the feature values of an image when the image is input. In the example shown in FIG. 1, the feature extraction unit 12 receives a medical image as input image IM and outputs the feature values of the input image IM. In FIG. 1, the arrows represent NN processing, FV1, FV2, FV3, and FV4 represent the feature values of the processing steps, and the widths of FV1, FV2, FV3, and FV4 represent the spatial resolution. The feature extraction unit 12 extracts the features of the input image IM while gradually reducing the resolution of the feature values. The feature values extracted by the feature extraction unit 12 are input to the detection processing unit 14.
[0026] The detection processing unit 14 is a neural network that, when an image feature is input, simultaneously outputs the position of a lesion, which is a region of interest in the image, and the class classification of the lesion. In the example shown in Fig. 1, the detection processing unit 14 receives the feature FV4 extracted by the feature extraction unit 12, and outputs a bounding box BB indicating the position of the lesion in the input image IM and the result of the lesion class classification, "Class: cancer."
[0027] In the network system 10, it is necessary to classify regions of interest based on relatively low-resolution features, which leads to a decrease in accuracy in tasks such as lesion classification, which require evaluation of detailed surface structures.
[0028] <Detection process that performs detection and classification separately> One way to increase the resolution during classification is to perform detection and classification using separate neural networks. Figure 2 is a conceptual diagram of a network system 20 in which detection and classification are performed using separate neural networks. The network system 20 includes a feature extraction unit 22, a detection processing unit 24, a cutout unit 26, and a classification processing unit 28.
[0029] The feature extraction unit 22 is a neural network (NN) that outputs the feature values of an image when an image is input. In the example shown in FIG. 2, the feature extraction unit 22 receives a medical image as input image IM and outputs the feature values of the input image IM. In FIG. 2, the solid arrows represent NN processing, FV11, FV12, FV13, and FV14 represent feature values in the processing steps, and the widths of FV11, FV12, FV13, and FV14 represent the spatial resolution. The feature extraction unit 22 extracts the features of the input image IM while gradually reducing the resolution of the feature values. The feature values extracted by the feature extraction unit 22 are input to the detection processing unit 24.
[0030] The detection processing unit 24 is a neural network that receives an image feature and outputs only the position of the lesion in the image. In the example shown in Fig. 2, the detection processing unit 24 receives the feature FV14 extracted by the feature extraction unit 22 and outputs a bounding box BB indicating the position of the lesion in the input image IM.
[0031] When an input image IM and a bounding box BB are input, the cutout unit 26 generates a cutout region CR by cutting out an area surrounded by the bounding box BB from the input image IM.
[0032] The classification processing unit 28 is a NN that, when a cut-out region of an image is input, outputs a class classification of the disease type of the lesion contained in the cut-out region. In the example shown in Figure 2, the cut-out region CR is input, and the result of the lesion classification, "Class: cancer," is output. In Figure 2, the solid arrows represent NN processing, FV21, FV22, and FV23 each represent feature amounts in the processing process, and the width of FV21, FV22, and FV23 represents the size of the spatial resolution.
[0033] The network system 20 solves the problem of resolution in classification processing. However, since the classification processing unit 28 needs to extract features from the input image IM anew, the calculation efficiency is poor and it is unsuitable for situations where real-time processing is required.
[0034] First Embodiment [Network system configuration] 3 is a conceptual diagram of a network system 30 according to the first embodiment. The network system 30 includes a feature extraction unit 22, a detection processing unit 24, a clipping unit 36, a resizing unit 38, and a classification processing unit 40.
[0035] The feature extraction unit 22 is a neural network (NN) that extracts a first feature and a second feature, which has a higher resolution than the first feature, from an input medical image. The second feature may be an intermediate feature in the process of extracting the first feature from the input image by the feature extraction unit 22. The resolution of the second feature may be determined in consideration of the trade-off between the accuracy of lesion classification and the time required for classification. The feature extraction unit 22 estimates only the position information of the lesion, which is the region of interest, through the detection process, while retaining the intermediate feature in the calculation process. In the example shown in Figure 3, FV14 corresponds to the first feature, and FV12 corresponds to the second feature.
[0036] The detection processor 24 is a neural network that receives an image feature and outputs only the position of the lesion in the image. In the example shown in Fig. 3, the detection processor 24 receives the feature FV14 as the first feature and outputs a bounding box BB that indicates the position of the lesion in the input image IM.
[0037] The cropping unit 36 crops out a portion of the second feature according to the detection result of the region of interest. When the feature of the image and the bounding box BB are input, the cropping unit 36 generates a feature map by cropping out the region corresponding to the bounding box BB from the feature map indicated by the feature. In the example shown in FIG. 3, the cropping unit 36 receives the feature map FM1 indicated by the feature FV12, which is the second feature, and the bounding box BB, and outputs a feature map FM2 by cropping out the region corresponding to the bounding box BB from the feature map FM1. The feature map FM1 is two-dimensional data having a width in the horizontal direction and a height in the vertical direction. The feature map FM1 indicates the feature of the input image IM reflecting position information within the input image IM. The feature map FM2 indicates the feature of the lesion in the input image IM.
[0038] The resizing unit 38 generates a feature map FM3 by adjusting the spatial size of the input feature map FM2 to a fixed size by interpolation or the like. The fixed size is, for example, a size suitable for input to the classification processing unit 40.
[0039] The classification processing unit 40 is a neural network (NN) that classifies the type of lesion based on input feature values. In the example shown in FIG. 3, the classification processing unit 40 receives a feature map FM3 and outputs the result of classifying the lesion included in the feature map FM3 as "Class: cancer." In FIG. 3, the solid arrows represent NN processing, FV31, FV32, and FV33 represent feature values in the processing steps, and the widths of FV31, FV32, and FV33 represent the resolution in the spatial direction. The classification processing unit 40 may classify the malignancy of the lesion.
[0040] Thus, the network system 30 includes a hierarchical neural network that includes the feature extraction unit 22, which is a feature extraction network, the detection processing unit 24, which is a first sub-network, and the classification processing unit 40, which is a second sub-network.
[0041] [Medical image processing method] FIG. 4 is a flowchart showing the steps of a medical image processing method using the network system 30.
[0042] In step S1, the network system 30 acquires a medical image. In the example shown in Fig. 3, an input image IM is acquired.
[0043] In step S2, the network system 30 extracts a first feature and a second feature having a resolution relatively higher than that of the first feature by processing the medical image acquired in step S1 using the feature extraction unit 22. In the example shown in Fig. 3, FV14 is extracted as the first feature and FV12 is extracted as the second feature.
[0044] In step S3, the network system 30 detects a lesion included in the medical image by processing the first feature extracted in step S2 with the detection processing unit 24. In the example shown in Fig. 3, a bounding box BB indicating the position of the lesion in the input image IM is output.
[0045] In step S4, network system 30 cuts out the region of the lesion detected in step S3 from the second feature amount extracted in step S2 using cutout unit 36. In the example shown in Fig. 3, feature map FM2 is cut out.
[0046] In step S5, the network system 30 resizes the feature map extracted in step S4 to a fixed size by the resizing unit 38. In the example shown in Fig. 3, a feature map FM3 is generated.
[0047] In step S6, the network system 30 classifies the lesion detected in step S3 by processing the feature map resized in step S5 using the classification processing unit 40. That is, the network system 30 classifies the lesion by processing the second feature amount. In the example shown in FIG. 3, the result of the lesion classification, "Class: cancer", is output.
[0048] As described above, the medical image processing method utilizes intermediate features from the feature extraction unit 22 rather than cutting out features from the input image IM. This method enables classification processing based on features with higher resolution than the network system 10, and eliminates the need to perform feature extraction from the input image using neural network processing, as in the network system 20. This allows for both high accuracy and real-time processing. Furthermore, the network system 30 performs classification processing after aligning the spatial size of the feature map, thereby reducing the impact of size differences in the region of interest.
[0049] Here, the second feature is an intermediate feature obtained in the process of extracting the first feature from the input image by the feature extraction unit 22, but the method for extracting the second feature is not limited to this example. For example, the feature extraction unit 22 may have a decoding process that increases resolution, such as a U-net, and generate the second feature by increasing the resolution of the first feature. Alternatively, the feature extraction unit 22 may have a NN that branches midway, and extract the first feature and the second feature separately.
[0050] [Learning of Hierarchical Neural Networks] For training of the hierarchical neural network consisting of the feature extraction unit 22, the detection processing unit 24, and the classification processing unit 40, a plurality of sets of training data sets, each of which is a combination of training images and correct answer data, are used.
[0051] The training images may be, for example, medical images, such as endoscopic images, CT (Computed Tomography) images, MRI (Magnetic Resonance Imaging) images, and ultrasound images.
[0052] The correct answer data includes position information of the region of interest included in the training image and classification label information of the region of interest.
[0053] The detection processing unit 24 may be trained using a first dataset, and the classification processing unit 28 may be trained using a second dataset different from the first dataset. The first dataset includes a pair of a first medical image and position information of a region of interest included in the first medical image, and the second dataset includes a pair of a second medical image and position information and classification class label of a region of interest included in the second medical image.
[0054] While the training of the detection processing unit 24 requires pairs of training images and position information of regions of interest, the training of the classification processing unit 40 also requires information on classification labels of regions of interest. When training the entire hierarchical neural network using the same data set, it is necessary to prepare correct answer data for all training images. On the other hand, if the detection processing unit 24 and the classification processing unit 40 are trained separately using different data sets, it is not necessary to assign classification labels to all data. Only the detection processing unit 24 can be trained using a training data set to which position information is assigned as correct answer data, and the classification processing unit 40 can be trained using only a training data set to which classification labels are assigned as correct answer data.
[0055] When learning is performed in two stages like this, it is desirable to first train the feature extraction unit 22 and the detection processing unit 24, and then train the classification processing unit 40 with the trained weights transferred to the feature extraction unit 22 and the detection processing unit 24. That is, the feature extraction unit 22 and the detection processing unit 24 are trained using a first data set, and the classification processing unit 28 is trained using a second data set based on the trained feature extraction unit 22 and the trained detection processing unit 24.
[0056] Before training the feature extraction unit 22 and the detection processing unit 24, the feature extraction unit 22 may be pre-trained using a third dataset different from the first dataset and the second dataset. The third dataset includes pairs of images and the results of image tasks. The images are not limited to medical images and may be general images. This has the advantage that it is easier to obtain a larger amount of data for general images than for medical images. Furthermore, the task is not limited to a detection task and may be a classification task.
[0057] [Configuration of medical image processing device] 5 is a block diagram showing the configuration of a medical image processing device 50 to which the network system 30 is applied. The medical image processing device 50 is realized by at least one computer. As shown in FIG. 5, the medical image processing device 50 includes a processor 52, a memory 54, an input device 56, and an output device 58.
[0058] The processor 52 acquires medical images and executes instructions stored in the memory 54. The hardware configuration of the processor 52 is made up of various processors as follows: The various processors include a CPU (Central Processing Unit), which is a general-purpose processor that executes software (programs) and functions as various functional units including the cropping unit 36 and the resizing unit 38, a GPU (Graphics Processing Unit), which is a processor specialized for image processing, a PLD (Programmable Logic Device), which is a processor whose circuit configuration can be changed after manufacturing such as an FPGA (Field Programmable Gate Array), and a dedicated electrical circuit, such as an ASIC (Application Specific Integrated Circuit), which is a processor having a circuit configuration designed specifically for executing specific processing.
[0059] A single processing unit may be configured with one of these various processors, or may be configured with two or more processors of the same or different types (e.g., multiple FPGAs, a combination of a CPU and an FPGA, or a combination of a CPU and a GPU). Also, multiple functional units may be configured with a single processor. Examples of multiple functional units configured with a single processor include, first, a configuration in which a single processor is configured with a combination of one or more CPUs and software, as typified by a client or server computer, and this processor operates as multiple functional units. Second, a configuration in which a processor is used to realize the functions of an entire system including multiple functional units on a single IC (Integrated Circuit) chip, as typified by an SoC (System on Chip). In this way, the various functional units are configured with one or more of the above-mentioned various processors as a hardware structure.
[0060] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.
[0061] The memory 54 stores instructions to be executed by the processor 52. The memory 54 also stores programs and weight parameters for operating the NNs of the feature extraction unit 22, the detection processing unit 24, and the classification processing unit 40. The memory 54 includes a random access memory (RAM) and a read-only memory (ROM), not shown. The processor 52 uses the RAM as a working area, executes software using various programs and parameters stored in the ROM, and performs various processes of the medical image processing device 50 by using parameters stored in the ROM, etc.
[0062] 4 is realized by the processor 52 executing a medical image processing program stored in the memory 54. The medical image processing program may be provided by a computer-readable non-transitory storage medium. In this case, the medical image processing device 50 may read the medical image processing program from the non-transitory storage medium and store it in the memory 54.
[0063] The medical image processing device 50 may train the hierarchical neural network by having the processor 52 execute a training program stored in the memory 54. The training data set may be stored in the memory 54. Note that the training of the hierarchical neural network may be performed by a computer different from the medical image processing device 50. In this case, it is sufficient that the program and weight parameters for operating the trained NN are stored in the memory 54.
[0064] The input device 56 may be, for example, a keyboard, a mouse, a touch panel, or other pointing devices, or a voice input device, or an appropriate combination thereof. A doctor can use the input device 56 to input various instructions and information to the medical image processing apparatus 50.
[0065] The output device 58 includes a display device. The display device may be, for example, a liquid crystal display, an organic electro-luminescence (OEL) display, a projector, or an appropriate combination of these. The input device 56 and the display device of the output device 58 may be integrated into one unit, such as a touch panel. The output device 58 may also include an audio output device, such as a speaker that outputs audio. The medical image processing device 50 can output the results of the lesion detection and classification processing to the output device 58 and notify the doctor of the results by image information or audio.
[0066] <Second embodiment> When the medical image processing device 50 notifies a physician of the results of lesion detection and classification processing using image information or audio, the medical image processing device 50 may notify the physician in different ways depending on the classification result. For example, when the medical image processing device 50 displays a graphic indicating the location of a lesion together with an input image on a display device, the medical image processing device 50 varies at least one of the graphic's color, line type, thickness, whether or not it blinks, the blinking cycle, and the number of blinks depending on the classification result. The medical image processing device 50 may also vary the type of graphic depending on the classification result. The type of graphic may be a rectangular frame (bounding box) surrounding the lesion, a circular frame surrounding the lesion, multiple parentheses surrounding the lesion, an arrow pointing to the lesion, or the like. Furthermore, when the medical image processing device 50 outputs audio indicating the location of a lesion to an audio output device, the medical image processing device varies at least one of the audio volume, pitch, whether or not it repeats, the repeat cycle, and the number of repeats depending on the classification result.
[0067] Furthermore, the medical image processing device 50 may suppress reporting only in the case of a specific classification. For example, in the case of lesion detection, suppressing reporting only in the case of lesions with a relatively low degree of malignancy can reduce the adverse effect of information with a relatively low level of importance interfering with a doctor's diagnosis.
[0068] <Third embodiment> An application example of a medical image processing device 50 that achieves both high accuracy and real-time processing for lesion detection and classification will be described below. Fig. 6 is a schematic diagram showing the overall configuration of an endoscope system 100 including the medical image processing device 50. As shown in Fig. 6, the endoscope system 100 includes the medical image processing device 50, an endoscope scope 110 which is an electronic endoscope, a light source device 111, an endoscope processor device 112, and a display device 113.
[0069] The endoscope 110 is a flexible endoscope, for example, used to capture time-series medical images. The endoscope 110 has an insertion section 120 that is inserted into a subject and has a distal end and a proximal end, a handheld operation section 121 that is connected to the proximal end of the insertion section 120 and that a doctor holds to perform various operations, and a universal cord 122 that is connected to the handheld operation section 121.
[0070] The insertion section 120 is formed in a long shape with a small diameter as a whole. The insertion section 120 is configured by sequentially arranging a flexible soft section 125 having flexibility from its base end side to its tip end side, a bending section 126 that can be bent by operating the hand-operated operation section 121, and a tip section 127 that incorporates an imaging optical system and an imaging element 128 (not shown).
[0071] The imaging element 128 is a CMOS (complementary metal oxide semiconductor) type or CCD (charge coupled device) type imaging element. Image light of the observed region is incident on the imaging surface of the imaging element 128 via an observation window (not shown) opened in the distal end surface of the distal end portion 127 and an objective lens (not shown) arranged behind the observation window. The imaging element 128 captures the image light of the observed region incident on its imaging surface and outputs an imaging signal.
[0072] The handheld operation unit 121 is provided with two types of bending operation knobs 129 used to bend the bending section 126, an air / water supply button 130 for air / water supply operation, and a suction button 131 for suction operation. The handheld operation unit 121 is also provided with a still image capture instruction unit 132 for issuing an instruction to capture a still image 139 of the observation site, and a treatment tool introduction port 133 for inserting a treatment tool (not shown) into a treatment tool insertion passage (not shown) that passes through the insertion section 120.
[0073] The universal cord 122 is a connection cord for connecting the endoscope 110 to the light source device 111. The universal cord 122 contains a light guide 135, a signal cable 136, and a fluid tube (not shown), which are inserted through the insertion section 120. The end of the universal cord 122 is provided with a connector 137a which is connected to the light source device 111, and a connector 137b which branches off from the connector 137a and is connected to the endoscope processor device 112.
[0074] By connecting the connector 137a to the light source device 111, the light guide 135 and the fluid tube are inserted into the light source device 111. As a result, the necessary illumination light, water, and gas are supplied from the light source device 111 to the endoscope 110 via the light guide 135 and the fluid tube. As a result, illumination light is irradiated toward the observation site from an illumination window (not shown) on the distal end surface of the tip portion 127. Furthermore, in response to pressing the air / water supply button 130 described above, gas or water is sprayed from an air / water supply nozzle (not shown) on the distal end surface of the tip portion 127 toward an observation window (not shown) on the distal end surface.
[0075] Connecting the connector 137b to the endoscope processor device 112 electrically connects the signal cable 136 to the endoscope processor device 112. As a result, an image signal of the observation site is output from the imaging element 128 of the endoscope 110 to the endoscope processor device 112 via the signal cable 136, and a control signal is output from the endoscope processor device 112 to the endoscope 110.
[0076] The light source device 111 supplies illumination light to the light guide 135 of the endoscope 110 via the connector 137a. The illumination light is selected from various wavelength bands depending on the purpose of observation, such as white light which is light of a white wavelength band or light of multiple wavelength bands, light of one or multiple specific wavelength bands, or a combination of these. Note that the specific wavelength band is a band narrower than the white wavelength band.
[0077] A first example of the specific wavelength band is, for example, the blue or green band of the visible range, and includes a wavelength band of 390 nm to 450 nm or a wavelength band of 530 nm to 550 nm, and the light of the first example has a peak wavelength within the wavelength band of 390 nm to 450 nm or a wavelength band of 530 nm to 550 nm.
[0078] A second example of the specific wavelength band is, for example, the red band in the visible range, and includes a wavelength band of 585 nm to 615 nm or a wavelength band of 610 nm to 730 nm, and the light in the second example has a peak wavelength within the wavelength band of 585 nm to 615 nm or a wavelength band of 610 nm to 730 nm.
[0079] A third example of the specific wavelength band includes a wavelength band where the absorption coefficients of oxygenated hemoglobin and reduced hemoglobin differ, and the light of the third example has a peak wavelength in the wavelength band where the absorption coefficients of oxygenated hemoglobin and reduced hemoglobin differ. The wavelength band of this third example includes a wavelength band of 400±10 nm, 440±10 nm, 470±10 nm, or a wavelength band of 600 nm or more and 750 nm or less, and the light of the third example has a peak wavelength in the wavelength band of 400±10 nm, 440±10 nm, 470±10 nm, or a wavelength band of 600 nm or more and 750 nm or less.
[0080] A fourth example of a specific wavelength band is the wavelength band (390 nm to 470 nm) of excitation light that is used for observing fluorescence emitted by fluorescent substances in living organisms (fluorescence observation) and that excites these fluorescent substances.
[0081] A fifth example of the specific wavelength band is the wavelength band of infrared light, which includes a wavelength band of 790 nm to 820 nm or a wavelength band of 905 nm to 970 nm, and the light in this fifth example has a peak wavelength in the wavelength band of 790 nm to 820 nm or a wavelength band of 905 nm to 970 nm.
[0082] The endoscope processor device 112 controls the operation of the endoscope 110 via the connector 137b and the signal cable 136. The endoscope processor device 112 also generates a moving image 138, which is a time-series medical image made up of time-series frame images including an image of a subject, based on an imaging signal acquired from the imaging element 128 of the endoscope 110 via the connector 137b and the signal cable 136. The frame rate of the moving image 138 is, for example, 30 fps (frames per second).
[0083] Furthermore, when the still image capturing instruction unit 132 is operated on the handheld operation unit 121 of the endoscope 110, the endoscope processor device 112 acquires one frame image from the video 138 in accordance with the timing of the capturing instruction, and generates the still image 139, in parallel with generating the video 138.
[0084] The endoscope processor device 112 outputs the generated moving images 138 and still images 139 to the display device 113 and the medical image processing device 50 in real time.
[0085] The endoscope processor device 112 may generate (acquire) a special light image having information of the specific wavelength band described above based on the normal light image obtained using the white light described above. The endoscope processor device 112 acquires the signal of the specific wavelength band by performing a calculation based on the RGB color information of red, green, and blue or the CMY color information of cyan, magenta, and yellow contained in the normal light image.
[0086] The endoscope processor device 112 may also generate a feature image, such as a publicly known oxygen saturation image, based on at least one of the normal light image obtained using the above-mentioned white light and the special light image obtained using light of the above-mentioned specific wavelength band (special light). In this case, the endoscope processor device 112 functions as a feature image generator. Note that the above-mentioned in-vivo image, normal light image, special light image, and video 138 or still image 139 including the feature image are all medical images obtained by imaging or measuring the human body for the purpose of image-based diagnosis or examination.
[0087] The display device 113 is connected to the endoscope processor device 112 and displays a moving image 138 and a still image 139 input from the endoscope processor device 112. The doctor performs operations such as moving forward and backward of the insertion section 120 while checking the moving image 138 displayed on the display device 113, and if a lesion or the like is found in the observed region, he operates the still image capturing instruction section 132 to capture a still image of the observed region, and also performs a diagnosis, a biopsy, or the like.
[0088] According to the endoscope system 100 to which the medical image processing device 50 is applied, lesions can be detected and classified in real time with high accuracy from the video 138 and still images 139 captured by the imaging element 128 of the endoscope 110.
[0089] <Other> While the present disclosure has been described above as a hierarchical neural network for detecting and classifying lesions in medical images, the present disclosure can be applied to a variety of images, not just medical images. For example, the technology disclosed herein can be applied to detecting regions of interest in images of structures such as bridges and classifying the regions of interest into cracks, peeling, rust, holes, etc.
[0090] The technical scope of the present invention is not limited to the scope described in the above embodiments. The configurations of the respective embodiments can be appropriately combined with each other without departing from the spirit of the present invention. [Explanation of symbols]
[0091] 10...Network system 12...Feature extraction unit 14...Detection processing unit 20...Network System 22...Feature extraction unit 24...Detection processing unit 26...Cutout section 28...Classification processing unit 30...Network System 36...Cutout section 38...Resize section 40...Classification processing unit 50...Medical image processing device 52...Processor 54...Memory 56...Input device 58...Output device 100...Endoscope system 110...Endoscope 111...Light source device 112...Endoscope processor device 113...Display device 120...insertion section 121...Hand control unit 122...Universal Code 125…Soft part 126...Bend 127...Tip 128...Image sensor 129...Bending control knob 130...Air and water supply button 131...Suction button 132...Still image capture instruction section 133...Treatment tool introduction port 135...Light guide 136...Signal cable 137a...Connector 137b...Connector 138...Video 139...Still image BB...Bounding box FM1, FM2, FM3... feature maps FV1, FV2, FV3, FV4...features FV11, FV12, FV13, FV14...Features FV21, FV22, FV23...Features S1 to S6: Steps in the medical image processing method
Claims
1. one or more processors for acquiring medical images; one or more memories in which programs to be executed by the one or more processors are stored; Equipped with The one or more processors: extracting a first feature amount and a second feature amount having a resolution relatively higher than that of the first feature amount from the medical image by processing the medical image with the feature extraction network of a hierarchical neural network including a feature extraction network, a first sub-network, and a second sub-network; detecting a region of interest included in the medical image by processing the first feature amount with the first sub-network of the hierarchical neural network; classifying the region of interest by processing the second feature amount with the second sub-network of the hierarchical neural network; Medical imaging equipment.
2. the second feature is an intermediate feature in a process of extracting the first feature from the medical image by the feature extraction network; The medical image processing device according to claim 1 .
3. the one or more processors train the first sub-network using a first data set and train the second sub-network using a second data set; the first data set includes a pair of a first medical image and position information of a region of interest included in the first medical image; the second data set includes a second medical image and a pair of position information and classification class label of a region of interest included in the second medical image; The medical image processing device according to claim 1 .
4. The one or more processors: training the feature extraction network and the first sub-network using the first data set; training the second sub-network based on the trained feature extraction network and the trained first sub-network using the second data set; The medical image processing apparatus according to claim 3 .
5. The one or more processors: prior to training the feature extraction network and the first sub-network using the first data set, pre-training the feature extraction network using a third data set different from the first data set and the second data set; The medical image processing apparatus according to claim 4 .
6. The one or more processors: notifying the position information of the attention area in a manner according to the result of the classification; The medical image processing device according to claim 1 .
7. The one or more processors: adding information based on the position information to the medical image and displaying the medical image on a display; The medical image processing apparatus according to claim 6 .
8. the one or more processors notify the position information of the region of interest when the result of the classification is a specific class. The medical image processing apparatus according to claim 6 .
9. The classes of the classification include malignancy grades, the one or more processors do not report if the malignancy level is relatively low; The medical image processing apparatus according to claim 8 .
10. The one or more processors: extracting a part of the second feature amount in accordance with the detection result of the region of interest; processing the extracted second feature quantity with the second sub-network; The medical image processing device according to any one of claims 1 to 9.
11. The one or more processors: adjusting the spatial size of the extracted second feature to a constant size; processing the second feature of the constant size in the second sub-network; The medical image processing apparatus of claim 10.
12. a feature extraction network that extracts a first feature amount and a second feature amount having a resolution relatively higher than that of the first feature amount from an input medical image; a first sub-network that detects a region of interest included in the medical image from an input first feature amount; a second sub-network that classifies the region of interest based on the input second feature amount; A hierarchical neural network including
13. 1. A medical image processing method executed by one or more processors, comprising: Acquire medical images, extracting a first feature amount and a second feature amount having a resolution relatively higher than that of the first feature amount from the medical image by processing the medical image with the feature extraction network of a hierarchical neural network including a feature extraction network, a first sub-network, and a second sub-network; detecting a region of interest included in the medical image by processing the first feature amount with the first sub-network of the hierarchical neural network; classifying the region of interest by processing the second feature amount with the second sub-network of the hierarchical neural network; Medical image processing methods.
14. A program that causes a computer to execute the medical image processing method according to claim 13.
Citation Information
Patent Citations
Image diagnosis support system and image diagnosis support method
WO2019142243A1
Line structure extraction device, method and program, and learned model
WO2020203552A1