Medical image observation support device

The medical image observation support device uses a trained neural network to classify cystoscopic images into Hunner-type interstitial cystitis, bladder pain syndrome, and bladder cancer, enhancing diagnostic accuracy by leveraging data augmentation and multiple network training.

JP7728002B2Active Publication Date: 2025-08-22NAT UNIV CORP TOKAI NAT HIGHER EDUCATION & RES SYST
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021178275
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-08-22
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

Existing medical imaging technologies struggle to accurately differentiate between Hunner-type interstitial cystitis, bladder pain syndrome, and bladder cancer, as these conditions often present similar cystoscopic images, leading to difficulties in diagnosis.

Method used

A medical image observation support device utilizing a classification unit with a trained convolutional neural network to categorize cystoscopic images into three groups: Hunner-type interstitial cystitis, bladder pain syndrome, and bladder cancer, incorporating a learning unit for training, a memory unit for storing data, and an additional data generation unit to enhance training data.

Benefits of technology

The device achieves accurate classification of cystoscopic images into the specified categories, improving diagnostic accuracy and reducing statistical recognition errors through multiple network training and data augmentation techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007728002000001
    Figure 0007728002000001
  • Figure 0007728002000002
    Figure 0007728002000002
  • Figure 0007728002000003
    Figure 0007728002000003
Patent Text Reader

Abstract

To provide an image observation support apparatus that can perform classification of Hunner type interstitial cystitis from an image filmed by a bladder endoscope using a machine learning method.SOLUTION: An image observation support apparatus 10 includes an image processing unit 34 including a learned network 36. The image processing unit 34 can classify endoscopic images read by an image reading unit 32 and acquired by a bladder endoscope into three classifications of Hunner type interstitial cystitis, bladder pain syndrome and bladder cancer.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a medical image analysis device that analyzes medical images using machine learning, and in particular to a device that analyzes images obtained by a cystoscope. [Background technology]

[0002] Computer-aided diagnosis (CAD) using medical images is widely used in the medical field. Medical images taken by various imaging devices are analyzed by a computer, and the analysis results are used to support observation. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-070609 [Non-patent literature]

[0004] [Non-Patent Document 1] Goodfellow, IJ, Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y., “Generative Adversarial Nets,” Advances in Neural Information Processing Systems 27 (NIPS) (2014)

[0005] An example of such a device is shown in Patent Document 1. Patent Document 1 discloses an image processing device that performs texture analysis on the overall shading of an image captured using an endoscope or the like, including blood vessels in the mucosal epithelium, and calculates blood vessel features, and uses these to output a diagnosis of non-tumor, adenoma, cancer, or SSA / P. Summary of the Invention [Problem to be solved by the invention]

[0006] On the other hand, in diagnostics using medical images, there are some tests in which it is extremely difficult to determine the presence or absence of a specific disease from the images. Diagnosis of Hunner-type interstitial cystitis (HIC) is one such example. Specifically, images taken with a cystoscope are used to diagnose whether a subject has Hunner-type interstitial cystitis, but few doctors can correctly diagnose Hunner-type interstitial cystitis in cystoscopy. Furthermore, there has not yet been any mechanical device that can determine the presence or absence of Hunner-type interstitial cystitis from images.

[0007] The present invention has been made from this perspective, and aims to provide a medical image observation support device that can use machine learning methods to (1) differentiate (distinguish between benign and malignant diseases) between Hannah type interstitial cystitis (a designated intractable disease that is a benign disease that significantly impairs quality of life) and bladder intraepithelial carcinoma (malignant tumor), which have similar cystoscopic images, from images taken with a cystoscope, and further (2) differentiate between Hannah type interstitial cystitis and bladder pain syndrome (no obvious abnormal findings in the bladder), which have similar clinical symptoms; in other words, (3) differentiate between the three groups of diseases: Hannah type interstitial cystitis, bladder pain syndrome, and bladder cancer. [Means for solving the problem]

[0008] The gist of the first invention of the present application is (a) a medical image observation support device that classifies medical images, (b) having a classification unit with a trained network, (c) the medical images are endoscopic images obtained by a cystoscope, and (d) the classification unit classifies the input endoscopic images into three categories: Hanna-type interstitial cystitis, bladder pain syndrome, and bladder cancer. [Effects of the Invention]

[0009] According to the medical image observation support device of the first invention, images obtained by a cystoscope are classified into three categories, Hanna-type interstitial cystitis, bladder pain syndrome, and bladder cancer, by a classification unit having a trained network.

[0010] Preferably, the medical image observation support device according to the second aspect of the present invention is characterized in that, in the medical image observation support device according to the first aspect of the present invention, (e) the device performs two classifications, namely, Hanna type interstitial cystitis and others, based on the results of the three classifications. In this way, more preferable classification results can be obtained than when the classification unit directly performs two classifications, namely, Hanna type interstitial cystitis and others.

[0011] Preferably, a medical image observation support device according to a third aspect of the present invention is characterized in that it is the medical image observation support device according to the first or second aspect of the present invention, further comprising: (f) a learning unit for training the network; (g) a memory unit for storing training data by the learning unit; and (h) an additional data generation unit for generating additional training data based on the training data stored in the memory unit, wherein (i) the learning unit trains the network using the additional training data in addition to training data previously stored in the memory unit. In this way, the additional training data is generated by the additional data generation unit based on the training data stored in the memory unit, and the network is trained by the learning unit using the additional training data in addition to the training data, so that more learning can be performed even when only training data is provided.

[0012] Preferably, the medical image observation support device according to the fourth aspect of the present invention is the medical image observation support device according to any one of the first to third aspects, characterized in that (j) the classification unit outputs the likelihood of the classification result. In this way, since information on the likelihood can be obtained in addition to the classification result, it becomes possible to refer to the classification result taking the likelihood into consideration.

[0013] Preferably, the medical image observation support device according to the fifth aspect of the present invention is characterized in that the medical image observation support device according to any one of the first to fourth aspects includes (k) a plurality of the classifiers, and (l) a judgment integration unit that determines a classification result based on the outputs of the plurality of classifiers, and (m) the plurality of classifiers include a corresponding plurality of the networks. In this way, the classification result can be determined based on the outputs of a plurality of classifiers, each having a plurality of trained networks, and statistical recognition errors caused by training can be suppressed.

[0014] Preferably, a medical image observation support system according to a sixth aspect of the present invention is characterized in that it includes an endoscope device and any one of the first to fifth aspects of the medical image observation support device, and when a release switch provided on the endoscope device is operated, an image is input from the endoscope device to the medical image observation support device. In this way, the process from capturing images by the endoscope device to classifying images by the medical image observation support device can be easily performed as a series of operations. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a diagram illustrating the configuration of a medical image observation support device according to an embodiment of the present invention; [Figure 2] FIG. 2 is a functional block diagram illustrating the main control functions of the assistance device. [Figure 3] FIG. 1 illustrates an example of a network. [Figure 4] FIG. 4 is a diagram illustrating dilated convolution in the network of FIG. 3. [Figure 5] 10 is a table summarizing the output results of the three classifications by the output unit. [Figure 6] Of the actual input images, the correct answer was Hanna type interstitial cystitis, and the images that the support device of the present invention judged to be Hanna type interstitial cystitis are shown. [Figure 7]Among the actual input images, the correct answer was Hanna type interstitial cystitis, and the support device of the present invention judged the image to be bladder pain syndrome. [Figure 8] Of the actual input images, the correct answer was Hannah type interstitial cystitis, and the support device of the present invention judged the image to be bladder cancer. [Figure 9] Of the actual input images, the correct answer was bladder pain syndrome, and the images that the support device of the present invention judged to be bladder pain syndrome are shown. [Figure 10] Among the actual input images, the correct answer was bladder pain syndrome, and the support device of the present invention judged the image to be Hanna-type interstitial cystitis. [Figure 11] Of the actual input images, the correct answer was bladder pain syndrome, and the support device of the present invention determined that the image was bladder cancer. [Figure 12] Of the actual input images, the correct answer was bladder cancer, and the image was determined to be bladder cancer by the support device of the present invention. [Figure 13] Among the actual input images, the correct answer was bladder cancer, and the support device of the present invention judged the image to be Hanna-type interstitial cystitis. [Figure 14] Among the actual input images, the correct answer was bladder cancer, and the support device of the present invention judged the image to be bladder pain syndrome. [Figure 15] 10 is a table summarizing the output results of the two classifications by the output unit. [Figure 16] 1 is a diagram showing a receiver operating characteristic (ROC) curve and AUC of the assistance device 10 of the present embodiment. [Figure 17] FIG. 10 is a diagram illustrating an outline of a medical image observation support system according to another embodiment of the present invention. [Figure 18] FIG. 10 is a functional block diagram illustrating the main parts of the control functions provided in the assistance device according to another embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0016] An embodiment of the present invention will now be described in detail with reference to the drawings. [Example]

[0017] FIG. 1 is a diagram illustrating the configuration of a medical image observation support device 10 (hereinafter simply referred to as support device 10) according to one embodiment of the present invention. As shown in FIG. 1, support device 10 of this embodiment includes a central processing unit (CPU) 12, a graphics processing unit (GPU) 13, a processing device specialized for image processing, a read-only memory (ROM) 14, a random access memory (RAM) 16, and a storage device 18. The CPU 12 and GPU 13 are so-called microcomputers that process and control electronic information based on a predetermined medical image analysis program pre-stored in ROM 14 while utilizing the temporary storage function of RAM 16. The medical image analysis program is supplied, for example, via a medium such as a CD-ROM or a network interface 30, which will be described later.

[0018] The storage device 18 is preferably a known storage device (storage medium) such as a hard disk drive or SSD (Solid State Drive) capable of storing information. It may also be a removable storage medium such as a small memory card or optical disk. The output interface 26 is for connecting a display device or the like as needed, and the input interface 28 is for connecting known input devices such as a keyboard or mouse, or for connecting a cable for capturing a predetermined image signal. When an image signal is input to the input interface 28, a converter, an encoder, or the like may also be provided as needed. The network interface 30 is for connecting the support device 10 to a network, enabling the support device 10 to communicate information with other computers via the network. A network cable, an antenna for a wireless network, or the like is connected to the network. Note that the output interface 26, the input interface 28, and the network interface 30 may be provided as needed and are not necessarily required.

[0019] Fig. 2 is a functional block diagram illustrating the main control functions of the support device 10. As shown in Fig. 2, the support device 10 is functionally configured to include an image reading unit 32, an image processing unit 34, an output unit 38, and a learning unit 40.

[0020] The image reading unit 32 inputs the image to be analyzed by the assistance device 10 to the image processing unit 34, which will be described later. Specifically, the image to be analyzed by the assistance device 10 is stored in advance in the storage device 18, and in response to an input operation by the operator, for example, the image to be analyzed is read from the storage device 18 and passed to the image processing unit 34. In this embodiment, an image of the inside of the bladder, which has been taken in advance by a cystoscope (bladder scope), is input.

[0021] The image processing unit 34 identifies an area of ​​interest in the image and classifies the area of ​​interest based on the image input from the image reading unit 32. The image processing unit 34 has a trained network 36 therein, which processes the input image and generates an output image.

[0022] The trained network 36 (hereinafter referred to as the network 36) has a configuration of a convolutional neural network (CNN). The trained network 36 calculates the feature amount of the input image by so-called deep learning.

[0023] The output unit 38 converts the output results from the image processing unit 34 into a predetermined format and outputs them to a file or another program, outputs them to a device such as a display device connected via the output interface 26, or transmits them to another computer via the network interface 30.

[0024] FIG. 3 is a diagram illustrating an example of a network 36. As described above, the network 36 is a network having a CNN-type structure. The rectangular layers 50a to 50s in FIG. 3 represent feature maps or images. In the following description, layers are referred to as layers 50 when they are not distinguished from one another. A block is formed by three to five layers 50 arranged horizontally in FIG. 3. Specifically, layers 50a to 50c form block 51a, layers 50d to 50g form block 51b, layers 50h to 50k form block 51c, layers 50l to 50o form block 51d, and layers 50p to 50t form block 51e. The numbers written below the leftmost layer 50 in each block 51 indicate the number of kernels in the layer 50 in that block.

[0025] Among the thick lines connecting blocks, the thick solid line 52 indicates 3x3 convolution. Here, 3x3 indicates the kernel size. The thick double line 54 indicates dilated convolution. The thick triple line 58 indicates global average pooling. This global average pooling 58 calculates and outputs the average of all pixel values ​​constituting the feature map in the immediately preceding layer 50s for each feature map. Note that for the leftmost layers 50a, 50d, 50h, 50l, and 50p of each block 51, those with multiple inputs combine these inputs.

[0026] 3, in the network 36 of this embodiment, two convolutions are performed in block 51a, three convolutions are performed in block 51b, two convolutions and one dilated convolution are performed in blocks 51c and 51d, and three convolutions and one all-level average pooling are performed in block 51e.

[0027] In addition, a residual path 56 is provided in each block 51 in the network 36, and dense pooling (64 to 70) is performed between each block 51, and input and output are also performed between blocks 51 other than the adjacent blocks 51.

[0028] Of these, the residual path 56 passes the contents of the layer 50 preceding a given layer 50 together with the output of the given layer 50 to the next layer 50. Specifically, for example, in block 51a, if we focus on layer 50b, the contents of layer 50a are convolved 52 and input to layer 50b. Meanwhile, the contents of layer 50b are convolved 52 and output from layer 50b to layer 50c. Here, the residual path 56 provided in block 51a combines the contents of layer 50a, which is the layer preceding layer 50b, with the output of layer 50b to form layer 50c. By providing this residual path 56, gradient loss in each layer is prevented.

[0029] Next, we will explain dense pooling. Dense pooling is a type of pooling that reduces the amount of data while retaining important elements from features (feature maps) obtained by convolution. Blocks 51a to 51c of network 36, feature maps or images that have undergone convolution 52 or dilated convolution, are used for subsequent convolutions using dense pooling.

[0030] Furthermore, the dense pooling in this embodiment uses mixed pooling. Specifically, the feature maps output from two convolutions undergo pooling with a kernel size of 2×2 (hereinafter referred to as 2×2 pooling, and the same applies to others) as indicated by the thin solid line 64, and are used as inputs for the next two convolutions (i.e., the first-stage output is used as input for the second-stage, the second-stage output is used as input for the third-stage, the third-stage output is used as input for the fourth-stage, and the fourth-stage output is used as input for the fifth-stage). Furthermore, as indicated by the thin dotted line 66, 4×4 pooling is performed, and the result is used as inputs for the next two convolutions (i.e., the first-stage output is used as input for the third-stage, the second-stage output is used as input for the fourth-stage, and the third-stage output is used as input for the fifth-stage). Furthermore, as indicated by the thin dashed line 68, 8×8 pooling is performed, and the result is used as inputs for the next two convolutions (i.e., the first-stage output is used as input for the fourth-stage, and the second-stage output is used as input for the fifth-stage). Furthermore, as shown by the thin dashed dotted line 70, 16x16 pooling is performed and the result is used as the input for the next two convolutions (i.e., the fifth stage for the output of the first stage). This makes it possible to repeat the convolution while obtaining information that may otherwise be lost, unlike when the output of two convolutions is used as the input for the next two convolutions after max pooling.

[0031] Fig. 4 is a diagram illustrating the dilated convolution represented by the thick double line 54 in Fig. 3. This dilated convolution is performed following the two convolutions 52 in blocks 51c and 51d in the network 36 in Fig. 3. The dilated convolution enables convolution processing that takes into account features that are spread out in the feature map.

[0032] In FIG. 4 , the thin solid line 72, thin dotted line 74, thin dashed line 76, and thin dashed-dotted line 78 all represent 3×3 convolutions, but each has a different dilation rate (DR): the thin solid line 72 corresponds to DR=1, the thin dotted line 74 corresponds to DR=2, the thin dashed line 76 corresponds to DR=4, and the thin dashed-dotted line 78 corresponds to DR=8. The dilation rate DR indicates the gap in the kernel in the convolution, with DR=1 indicating that the kernel is used as is. Furthermore, DR=n indicates that a kernel with (n−1) gaps between each element of the original kernel is used. In other words, the dilated convolution 54 of this embodiment simultaneously performs dilated convolutions with multiple different dilation rates. Dilated convolution allows for convolution of a large feature map without using pooling, and unlike pooling, it has the advantage of not reducing the size of the output feature map.

[0033] That is, as shown in FIG. 4, dilated convolution 54 integrates and outputs the output obtained by performing 1×1 convolution 58a on the input, the output obtained by performing 3×3 convolution 72 with DR=1 and performing 1×1 convolution 58b on the output, the output obtained by performing 3×3 convolution 74 with DR=2 and performing 1×1 convolution 58c on the output, the output obtained by performing 3×3 convolution 76 with DR=4 and performing 1×1 convolution 58d on the output, and the output obtained by performing 3×3 convolution 78 with DR=8 and performing 1×1 convolution 58e on the output.

[0034] 2, the learning unit 40 trains the network 36. Specifically, so-called supervised learning is performed using training images and corresponding ground truths. The learning unit 40 may train the network 36 in advance before using the assistance device 10, or may further train the network 36 during use.

[0035] The learning unit 40 (1) divides endoscopic images captured in advance and stored in the memory unit 18 into training images and evaluation images for evaluating the network after training at a fixed ratio. Then, (2) it uses the training images to train the network 36, and (3) it uses the trained network 36 to classify the evaluation images. The learning of the network 36 is completed by repeating this series of steps (1) to (3) a predetermined number of times. In this way, the training images are replaced with each repetition of the steps, allowing the network 36 to be trained optimally. Furthermore, since multiple evaluations can be performed, the evaluation results are less susceptible to the selection of the training images. Furthermore, because the above series of steps are independent of each other, data contamination, in which the same data is mixed in the training data and the evaluation data, does not occur.

[0036] Furthermore, the additional data generation unit 42 performs a so-called data augmentation process, which creates additional data by performing a predetermined process on images to be used as learning images by the learning unit 40 as described above, from among the endoscopic images previously stored in the storage unit 18. The additional data created in this way is used for training the network 36 in the learning unit 40, together with the learning images described above. In this embodiment, the additional data generation unit 42 generates a preset number of additional data.

[0037] Here, the predetermined processing performed by the additional data generating unit 42 includes, for example, the following processing: (1) Rotating the endoscopic image stored in the storage unit 18 (hereinafter referred to as the "original image") by an angle of ±30 degrees or less; (2) translating the original image in at least one of the up, down, left, and right directions by ±22 pixels or less (in other words, within 10% of the image width); (3) Enlarging or reducing the original image by a factor of 0.8 to 1.2; (4) Flip the original image vertically or horizontally; (5) Changing at least one of the hue, saturation, and brightness in the HSV color space. The additional data generating unit 42 generates a new image by performing at least one of the processing steps (1) to (5) on the original image. At this time, which processing step is applied is determined randomly. If, as a result of performing the processing steps (1) to (5), pixels that are outside the original image appear in the new image, those pixels are blacked out.

[0038] Alternatively, a generative learning model, GAN (Generative Adversarial Model), as disclosed in Non-Patent Document 1 above, may be used to learn cystoscope images and generate learning images.

[0039] In the training of the network 36 by the training unit 40, the mini-batch size, which is the number of data items in one training run, and the number of training epochs, which is the repetition index for the training process for all data, are appropriately determined. The additional data generation unit 42 generates additional data for each epoch.

[0040] (Experimental example) The following experiment was carried out to verify the effectiveness of this embodiment. In this experimental example, the support device 10 classifies the cases corresponding to the cystoscopy images into three categories: Hanna type interstitial cystitis, bladder pain syndrome, or bladder cancer, and then, based on the results of these three categories, classifies the cases into two categories: Hanna type interstitial cystitis and others.

[0041] First, cystoscopic images obtained from 681 patients are stored in the storage device 18. The cystoscopic images are taken by the cystoscope, and each image is 224 x 224 pixels in size. These images are stored in association with information about their classification by a specialist, i.e., information about whether they are Hanna-type interstitial cystitis, bladder pain syndrome, or bladder cancer.

[0042] Prior to the experiment, network 36 is trained by training unit 40 based on the cystoscope images of the above-mentioned 681 people stored in storage device 18 in advance and additional data generated by additional data generation unit 42.

[0043] First, the learning unit 40 divides the cystoscope images of 681 people stored in advance in the storage device 18 into learning images and evaluation images at a predetermined ratio, 80:20 in this experimental example.

[0044] Next, the additional data generation unit 42 generates additional data from the divided training images. Specifically, as described above, (1) the training image (original image) is rotated by an angle of ±30 degrees or less; (2) the original image is translated by ±22 pixels or less in at least one of the up, down, left, and right directions (in other words, within 10% of the image width); (3) the original image is enlarged or reduced by 0.8 to 1.2 times; (4) the original image is flipped vertically or horizontally; and (5) at least one of hue, saturation, and brightness is changed in the HSV color space. The additional data generation unit 42 generates a new image by performing at least one of the processes (1) to (5) on the original image. At this time, the correct answer data for the generated additional data is the original data. In this experimental example, the additional data generation by the additional data generation unit 42 is performed so that the total size of the training image and the additional data generated by the additional data generation unit 42 is 50 times the size of the training image.

[0045] The learning unit 40 uses the learning images and additional data to train the network 36. The learning images and additional data (hereinafter referred to as "input images") are input to the network 36 shown in FIG. 3. At this time, the input images are three images for each color element in the image, i.e., three images for each of the R element, G element, and B element, and are input to the network 36 as density values ​​for each of these color elements.

[0046] In this experimental example, the learning of the network 36 by the learning unit 40 was performed with a mini-batch size of 15, which is the number of data items in one learning session, and a learning epoch count, which is the repetition index for the learning process for all data, of 200. These values ​​may be set to values ​​at which the weight parameters are experimentally considered to have converged, and can be changed as appropriate. The learning rate, which is the amount by which the weight parameters are changed during learning, was initially set to 0.005 and gradually decreased after 100 epochs. This is because the initial value was set to a relatively large value to prevent the weights from reaching a local minimum solution, but the value was reduced to avoid divergence after repeated learning.

[0047] An image to be classified by the assistance device 10 of this embodiment is input to the network 36 trained in this manner. As described above, in this experimental example, the input image is 224 x 224 pixels in size and is input as an image for each of the R, G, and B color components. The input image input to the first layer 50a passes sequentially through each layer 50 of the network 36, undergoing the above-described processing. When reaching the final layer 50t of the network 36, global average pooling 58 is performed as described above. At this time, each image in layer 50s of block 51e is digitized, and the input image is evaluated for one of three classifications: Hanna-type interstitial cystitis, bladder pain syndrome, or bladder cancer. Note that if one side of the input image is less than 224 pixels, that portion may be blacked out or enlarged to 224 pixels. Furthermore, if one side is greater than 224 pixels, the image may be trimmed to 224 pixels or reduced to 224 pixels.

[0048] At this time, the network 36 is configured to output the likelihood of each of the three classifications for the input image. That is, for each of the three classifications, Hanna-type interstitial cystitis, bladder pain syndrome, and bladder cancer, the network 36 outputs the likelihood that the image is considered to be that of the three classifications.

[0049] The output unit 38 outputs which of the three categories the input image falls into based on the output of the network 36. Specifically, for example, it outputs that the input image falls into the category with the highest likelihood among the likelihoods of the three categories output.

[0050] Fig. 5 is a table summarizing the output results from output unit 38 when 646 of the 681 cystoscope images stored in storage device 18 were input to trained network 36. Fig. 5 shows the relationship between the correct answer and each of the three classifications classified by output unit 38. According to this, the correct answer rate for images that were identified as Hannah type interstitial cystitis was 82.9%, the correct answer rate for images that were identified as bladder pain syndrome was 93.1%, and the correct answer rate for images that were identified as bladder cancer was 86.5%.

[0051] 6 to 14 show actual input images for each classification result. Of these, FIGS. 6 to 8 show images for which the correct answer is Hanna-type interstitial cystitis, and of these, FIG. 6 is an image that the support device 10 of this embodiment also classified as Hanna-type interstitial cystitis. FIG. 7 is an image that the support device 10 of this embodiment classified as bladder pain syndrome. Also, FIG. 8 is an image that the support device 10 of this embodiment classified as bladder cancer. Note that in FIGS. 6 to 8, and in FIGS. 9 to 11 and 12 to 14 described below, the likelihood that served as the basis for the determination is shown together with the image. The output unit 38 can display the image, the determination result, and the likelihood for the determination result together.

[0052] Figures 9 to 11 show images for which the correct answer is bladder pain syndrome. Of these, Figure 9 is an image that the support device 10 of this embodiment also classified as bladder pain syndrome. Figure 10 is an image that the support device 10 of this embodiment classified as Hanna-type interstitial cystitis. Figure 11 is an image that the support device 10 of this embodiment classified as bladder cancer.

[0053] Figures 12 to 14 show images for which the correct answer is bladder cancer, and of these, Figure 12 is an image that the support device 10 of this embodiment also classified as bladder cancer. Figure 13 is an image that the support device 10 of this embodiment classified as Hanna-type interstitial cystitis. Figure 14 is an image that the support device 10 of this embodiment classified as bladder pain syndrome.

[0054] Furthermore, in this experimental example, the output unit 38 can classify the results of the three classifications, i.e., Hanna-type interstitial cystitis, bladder pain syndrome, and bladder cancer, that are determined to be bladder pain syndrome or bladder cancer as not being Hanna-type interstitial cystitis. In other words, the output unit can output two classifications, either Hanna-type interstitial cystitis or not Hanna-type interstitial cystitis, based on the results of the three classifications.

[0055] FIG. 15 is a table summarizing the output results of the two-category classification by the output unit 38. FIG. 15 shows the relationship between each of the two categories classified by the output unit 38 and the correct answer. According to this, the sensitivity of the support device 10 is the ratio of the number of images determined to be Hanna type interstitial cystitis to the total number of images for which the correct answer is Hanna type interstitial cystitis, that is, 165 / (165+34), or 82.9%. Furthermore, the specificity is the number of images determined not to be Hanna type interstitial cystitis to the total number of images for which the correct answer is not Hanna type interstitial cystitis, that is, 416 / (416+31), or 93.1%. According to this, the accuracy rate for the two-category classification is 89.9%, which is higher than when three-category classification is performed.

[0056] As described above, it can be seen that both the three-class and two-class judgments provide good judgment as to whether the cystoscope image is Hanna-type interstitial cystitis.

[0057] FIG. 16 shows the receiver operating characteristic (ROC) curves of the support device 10 of this embodiment. The solid line indicates Hanna-type interstitial cystitis (truth) versus the other case, i.e., bladder pain syndrome or bladder cancer (false). The dashed line indicates bladder pain syndrome (truth) versus the other case, i.e., bladder cancer or Hanna-type interstitial cystitis (false). The dashed-dotted line indicates bladder cancer (truth) versus the other case, i.e., Hanna-type interstitial cystitis or bladder pain syndrome (false). As shown in FIG. 16, all of the ROC curves indicate that the support device 10 of this embodiment performed well in classification. The AUC (the ratio of the area below each ROC curve) was 0.96 for the solid line, 0.98 for the dashed line, and 0.97 for the dashed-dotted line, indicating that the performance of the classifier (classification algorithm) of the support device 10 of this embodiment was good.

[0058] The support device 10 of this embodiment has an image processing unit 34 (classification unit) with a trained network 36, and the image processing unit 34 can classify the endoscopic images obtained by a cystoscope and read by the image reading unit 32 into three categories: Hanna-type interstitial cystitis, bladder pain syndrome, and bladder cancer.

[0059] Furthermore, according to the support device 10 of this embodiment, the output unit 38 performs two classifications, either Hanna type interstitial cystitis or others, based on the results of the three classifications, so that more favorable classification results can be obtained compared to when directly performing two classifications, either Hanna type interstitial cystitis or others.

[0060] Furthermore, the assistance device 10 of this embodiment includes a learning unit 40 for training the network, a memory unit 18 for storing training data by the learning unit 40, and an additional data generation unit 42 for generating additional training data based on the training data stored in the memory unit 18. The learning unit 40 trains the network 36 using the additional training data in addition to the training data previously stored in the memory unit 18, so that even when only training data is given, more learning can be performed.

[0061] Furthermore, according to the support device 10 of this embodiment, the image processing unit 34 outputs the likelihood of the classification result, so that it is possible to refer to the classification result taking the likelihood into consideration.

[0062] Next, another embodiment of the present invention will be described. In the following description, parts common to the above embodiment will be given the same reference numerals and description thereof will be omitted. [Example]

[0063] FIG. 17 is a diagram illustrating an outline of a medical image observation support system (hereinafter referred to as "support system") 100 including the support device 10 of the present invention, and corresponds to FIG. 2 in the above-described embodiment.

[0064] The support system 100 of this embodiment includes an endoscopic device 102 in addition to the support device 10 of Embodiment 1. While in the above-described Embodiment 1, an input image previously stored in the storage unit 18 is input to the image reading unit 32, this embodiment differs in that the input is directly from the endoscopic device 102. Specifically, the endoscopic device 102, i.e., the cystoscope 102, is provided with a release switch 104. When the practitioner presses the release switch 104 while capturing the affected area of ​​the patient with the cystoscope 102, the cystoscope image at that time is input to the image reading unit 32.

[0065] When an input image is input to image reading unit 32 from cystoscope 102, image reading unit 32 causes image processing unit 34 to read the image in the same manner as in the first embodiment described above, and image processing unit 34 performs three-category classification of the input image using trained network 36. Furthermore, output unit 38 outputs the three-category classification results for Hannah type interstitial cystitis, bladder pain syndrome, or bladder cancer, and also outputs the likelihood calculated by image processing unit 34 as needed, or outputs one of two classifications, whether it is Hannah type interstitial cystitis or not, based on the three-category classification results.

[0066] The support system 100 of this embodiment includes an endoscopic device 102 and a support device 10. When a release switch 104 provided on the endoscopic device 102 is operated, an image is input from the endoscopic device 102 to the support device 10. This allows the process from capturing an image by the endoscopic device 102 to classifying the image by the support device 10 to be easily performed as a series of operations. [Example]

[0067] This embodiment is another embodiment corresponding to the image processing unit 34 and the output unit 38 in the previous embodiment.

[0068] FIG. 18 is a block diagram illustrating an outline of the control functions of the image processing unit 134, the output unit 138, and the decision integration unit 139, which are used in place of the image processing unit 34 and the output unit 38 in FIG. 2 in the previous embodiment.

[0069] 18, the image processing unit 134 of this embodiment is configured to include a plurality (n in FIG. 18) of trained networks 36a, 36b, ..., 36n. The trained networks 36a, 36b, ..., 36n each correspond to the trained network 36 of the above-described embodiment, but the training data used for each network may be different or the same.

[0070] The output unit 138 of this embodiment includes output units 38a, 38b, ..., 38n corresponding to the trained networks 36a, 36b, ..., 36n, respectively. The output units 38a, 38b, ..., 38n correspond to the output unit 38 in the above-described embodiment. Therefore, the output units 38a, 38b, ..., 38n perform classification based on the feature values ​​calculated by the corresponding trained networks 36a, 36b, ..., 36n. Here, since the trained networks 36a, 36b, ..., 36n are trained using different training data, the individual output units 38a, 38b, ..., 38n may output different classifications for a common determination image. Note that the number of trained networks 36a, 36b, ..., 36n and the number of output units 38a, 38b, ..., 38n are not limited to a plurality (i.e., n≧2).

[0071] The decision integrating unit 139 determines a final decision result by integrating the results output from the multiple output units 38 a, 38 b, ..., 38 n. The decision integrating unit 139 may determine a final decision result by taking a majority vote of the output results, or may determine a final decision result by taking a majority vote that takes into account the likelihood calculated by each of the trained networks 36 a, 36 b, ..., 36 n.

[0072] (Experimental example) The inventors of the present application conducted an experiment to determine images of Hannah type interstitial cystitis from multiple cystoscope images using the support device 10 of this embodiment and the support device 10 of Example 1. In this experiment, 84 cystoscope images were prepared. The breakdown was 37 images of Hannah type interstitial cystitis, 36 images of bladder cancer, and 11 images of bladder pain syndrome. In the support device of this embodiment, 10 networks were used as the multiple networks 36a, 36b, ..., 36n. Furthermore, the determination by the integrated determination unit 139 was based on a majority vote of the output results from the 10 output units 38a, 38b, ..., i.e., the determination that was most frequently output as the output result was adopted.

[0073] According to the support device 10 of Example 1, the accuracy rate (correct diagnosis rate) of images determined to be Hannah type interstitial cystitis was 84.5%. On the other hand, according to the support device 10 of this example, the accuracy rate was 89.2%. Furthermore, the accuracy rate for all three types of images was 75.3% according to the support device 10 of Example 1, and 79.8% according to the support device 10 of this example. According to these experiments, it was found that the accuracy rate of the support device 10 of this example, i.e., the judgment by the judgment integration unit 139 based on the judgment results of the multiple networks 36a, 36b, ..., 36n, was 4% or more higher than that of the support device 10 of Example 1, i.e., the judgment by a single network 38.

[0074] According to the support device 10 of this embodiment, the judgment image can be judged using likelihoods calculated by multiple trained networks 36a, 36b, ..., 36n, which have been trained using different training data, thereby reducing statistical recognition errors that occur due to training.

[0075] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the present invention is not limited to this embodiment and can be implemented in other modes.

[0076] For example, the network 36 described above is composed of five blocks 51, but the number of blocks 51 is not limited to this. Furthermore, the number of layers 50 constituting each block 51 is not limited to that in the embodiment.

[0077] Furthermore, in the above embodiment, the likelihood is displayed by the output unit 38, but this is not essential.

[0078] In the above-described embodiment, when the image processing unit 34 and the output unit 38 perform three-way classification, the classification result with the highest likelihood is set as the classification result, but this is not limited to this. For example, if the difference in likelihood between each classification is within a certain threshold, a classification result indicating that classification is impossible may be output, or a definitive classification result may not be output, but the classification names that fall within the predetermined range and the likelihoods of those classifications may be displayed.

[0079] Furthermore, the parameters used when the learning unit 40 learns the network 36, i.e., the number of mini-batches and the number of epochs, can be changed as appropriate. Furthermore, the number of additional data generated by the additional data generation unit 42 can also be changed as appropriate.

[0080] In the above-described embodiment, the network 36 is a convolutional neural network (CNN), but is not limited to this. As long as the network functions as a decision-maker capable of three-way classification using a network and is capable of learning global features and local features, it may be a network with a different configuration, such as YOLO.

[0081] Furthermore, the above-described embodiment is not limited to a configuration in which the support device 10 and the endoscopic device 102 are directly connected by a dedicated cable. The support device 10 and the endoscopic device 102 may be installed in separate locations so that they can communicate with each other via a network and images captured by the endoscopic device 102 are transmitted to the support device 10 via the network. In such a configuration, the support system 100 may be configured so that multiple endoscopic devices 102 share one support device 10.

[0082] Furthermore, the determination by the assistance device 10 of this embodiment may be made in real time during the examination using the cystoscope 102, or may be made by taking only images during the cystoscopic examination and then transmitting the captured images to the assistance device 10 for determination. Alternatively, the images may be transmitted to a server via a communication line (information network), the server may make a determination, and the result may be transmitted to the endoscope device. Alternatively, the endoscopic images transmitted from the endoscope device may be stored in a server, and the diagnostic result may be presented offline for the stored images.

[0083] In addition, in the network 36 shown in FIG. 3, a 3×3 convolution is performed at the location indicated by the thick solid line 52, for example. In addition to this convolution, batch normalization, i.e., normalizing the same channel of all data in one mini-batch so that their mean is 0 and their variance is 1, may also be performed.

[0084] In addition, in the above-described embodiment, cystoscopic images are classified into three categories to identify images of Hannah type interstitial cystitis, but cystoscopic images may be classified into four or more categories to identify images of Hannah type interstitial cystitis.

[0085] The above is merely one embodiment, and although other examples will not be given, the present invention can be implemented in various forms with various modifications and improvements based on the knowledge of those skilled in the art within the scope of the present invention. [Explanation of symbols]

[0086] 10: Medical image observation support device 34, 134: Image processing unit (classification unit) 36, 36a, 36b, ..., 36n: Network 38, 138: Output section (classification section) 40: Learning Department 42: Additional data generation unit 100: Medical image observation support system 102: Endoscopic device (cystoscope) 104: Release switch 139: Judgment integration unit

Claims

1. A medical image observation support device that classifies medical images, a classifier having a trained network; the medical image is an endoscopic image obtained by a cystoscope, the classification unit classifies the input endoscopic image into three categories: Hannah type interstitial cystitis, bladder pain syndrome, and bladder cancer; A medical image observation support device characterized by:

2. Based on the results of the three classifications, the patient is classified into two categories: Hanna type interstitial cystitis and others.

2. The medical image observation support device according to claim 1, wherein:

3. a learning unit for learning the network; a storage unit that stores data for learning by the learning unit; an additional data generating unit that generates additional learning data based on the learning data stored in the storage unit, the learning unit causes the network to learn using the additional learning data in addition to the learning data stored in advance in a storage unit; 3. The medical image observation support device according to claim 1, wherein:

4. the classification unit outputs a likelihood of the classification result; 4. The medical image observation support device according to claim 1, wherein:

5. A plurality of the classifiers; a decision integration unit that determines a classification result based on outputs of the plurality of classification units; the plurality of classifiers have a corresponding plurality of the networks; 5. The medical image observation support device according to claim 1, wherein:

6. A medical image observation support system including an endoscope apparatus and the medical image observation support apparatus according to any one of claims 1 to 5, When a release switch provided on the endoscope device is operated, an image is input from the endoscope device to the medical image observation support device; A medical image observation support system characterized by the above.

Citation Information

Patent Citations

  • Endoscope device

    JP2000066114A

  • Method of diagnosing a lower urinary tract disorder

    JP2009066090A

  • Image processing apparatus and image processing method

    JP2017070609A

  • Information processing device, information processing method and program

    JP2021097728A

  • System and method for diagnostic support using pathological image of skin tissue

    WO2018221625A1