Computer program, information processing device, information processing method, and learning model generation method
A learning model-based system improves biliary endoscopy diagnosis accuracy by analyzing endoscopic images, leveraging data augmentation techniques to enhance training, effectively supporting non-specialist physicians.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2026-05-28
AI Technical Summary
The accuracy of biliary endoscopy diagnosis is low due to the shortage of specialists, and diagnosing biliary tract diseases requires high expertise, making it difficult for non-specialists to provide accurate assessments.
A computer program and information processing method that utilizes a learning model, such as a neural network, to analyze endoscopic images of the biliary tract, predicting the malignancy of tumors and findings of lesions, supported by data augmentation techniques like Depth-Anything and image conversion models like CycleGAN to enhance the training dataset, improving diagnostic accuracy.
Enhances the accuracy of biliary endoscopic diagnoses, allowing non-specialist physicians to achieve the same level of diagnostic support as specialists, addressing the shortage of skilled professionals.
Smart Images

Figure JP2024041080_28052026_PF_FP_ABST
Abstract
Description
Computer Program, Information Processing Apparatus, Information Processing Method, and Learning Model Generation Method
[0001] The present invention relates to a computer program, an information processing apparatus, an information processing method, and a learning model generation method.
[0002] In recent years, endoscopes have been widely used for diagnosing the inside of a living body such as the digestive tract. A doctor diagnoses digestive tract diseases using the images captured by an endoscope. Patent Document 1 discloses an apparatus for inserting a scope into the lumen of an organ such as the stomach or large intestine to examine the presence or absence of lesions, and particularly discloses an apparatus that can appropriately determine the comprehensiveness of observation according to the observation flow during an endoscopic examination.
[0003] Japanese Patent Application Laid-Open No. 2024-8815
[0004] Among the organs handled by doctors in the department of gastroenterology, biliary tract diseases tend to be discovered in an advanced stage, and since the biliary tract is deep inside the body, endoscopic diagnosis and treatment require high expertise. On the other hand, the number of specialists regarding biliary tract diseases is very small, and it is desired to solve the problem of shortage of specialists by improving the accuracy of endoscopic image diagnosis.
[0005] The present invention has been made in view of such circumstances, and an object thereof is to provide a computer program, an information processing apparatus, an information processing method, and a learning model generation method that can improve the accuracy of biliary endoscopy diagnosis.
[0006] Although the present application includes a plurality of means for solving the above problems, for example, the computer program acquires image data obtained by a choledochoscope, and when the image data is input, causes a computer to execute a process of inputting the acquired image data into a learning model that predicts information regarding the malignancy of a biliary tumor and the findings of a lesion part of the biliary tract, and predicting information regarding the malignancy of the biliary tumor and the findings of the lesion part of the biliary tract.
[0007] According to the present invention, the accuracy of biliary endoscopy diagnosis can be improved.
[0008] This figure shows an example of the configuration of the information processing system of this embodiment. This figure shows an example of the configuration of the learning model. This figure shows a first example of the learning method for the learning model. This figure shows a first example of the learning method for the learning model. This figure shows a second example of the learning method for the learning model. This figure shows a second example of the learning method for the learning model. This figure shows a second example of the learning method for the learning model. This figure shows an example of the original image input to the learning model and a visualized image that visualizes the basis for the prediction result of the learning model. This figure shows an example of the original image and a depth image. This figure shows an example of preprocessing for the original image. This figure shows an example of an unstained image and a stained image. This figure shows an example of an unstained image and a stained image. This figure shows an example of the configuration of the image conversion model. This figure shows an example of the configuration of the image conversion model. This figure shows an example of a pseudo-stained image. This figure shows an example of cutout. This figure shows an example of visualization of the basis for prediction using GradCAM. This figure shows an example of a prediction basis image that visualizes the basis for prediction results regarding information on findings of lesions in the biliary tract. This figure shows an example of a prediction basis image that visualizes the basis for prediction results regarding information on findings of lesions in the biliary tract. This figure shows an example of a prediction basis image that visualizes the basis for the prediction results regarding the findings of lesions in the biliary tract. This figure shows an example of a method for evaluating a learning model. This figure shows the first example of the prediction result screen of the learning model. This figure shows the second example of the prediction result screen of the learning model. This figure shows the third example of the prediction result screen of the learning model. This figure shows the fourth example of the prediction result screen of the learning model. This figure shows an example of the procedure for generating a learning model by an information processing device. This figure shows an example of the procedure for prediction processing by an information processing device.
[0009] Embodiments of the present invention will be described below. Figure 1 is a diagram showing an example of the configuration of the information processing system of this embodiment. The information processing system comprises an information processing device 50 and an endoscope device 20. The endoscope device 20 and a terminal device 10 are connected to the information processing device 50 via a communication network 1. The terminal device 10 is composed of, for example, a personal computer, a tablet terminal, etc., and can be used by a physician.
[0010] The endoscope device 20 comprises a main unit 21 and an endoscope 22. The main unit 21 includes a processor, a light source, a display device, and an input device (none of which are shown). The endoscope 22 includes a long insertion section that is inserted into the body and an operating section. The endoscope 22 is detachably connected to the main unit 21 via a universal cord. Light introduced from a light source provided in the main unit 21 via an optical fiber is irradiated from the tip of the insertion section of the endoscope 22, and an image of the body (endoscopic image) can be acquired by the imaging optical system of the endoscope 22. The main unit 21 can output the acquired image to the information processing device 50. In addition, the display device of the main unit 21 can display the acquired image, and the physician can operate the insertion section inside the body by operating the operating section of the endoscope 22 while observing the image.
[0011] When acquiring images of the bile duct using the endoscope 22, the endoscope 22 comprises, for example, a duodenal endoscope and a cholangioscope. The duodenal endoscope is inserted into the duodenum, and the cholangioscope, which is attached to the tip of the duodenal endoscope, is inserted into the bile duct. The endoscope 22 may be a thin-diameter endoscope. In the following description, the bile duct will be used as an example of a site in the body, but the scope of application of this embodiment is not limited to the bile duct.
[0012] The information processing device 50 includes a control unit 51 that controls the entire device, a communication unit 52, a memory 53, and a storage unit 54.
[0013] The control unit 51 may be configured by incorporating a required number of CPUs (Central Processing Units), MPUs (Micro-Processing Units), GPUs (Graphics Processing Units), etc. Alternatively, the control unit 51 may be configured by combining DSPs (Digital Signal Processors), FPGAs (Field-Programmable Gate Arrays), etc.
[0014] The communication unit 52 is equipped with a communication module and has the function of communicating with the endoscope device 20 and the terminal device 10 via the communication network 1.
[0015] The storage unit 54 can be made up of semiconductor memory or a hard disk, and stores a computer program 55 (program product), a learning model 56, an image conversion model 57, and other necessary information.
[0016] The computer program 55 can be read by a recording medium (e.g., an optically readable disc storage medium such as a CD-ROM) M using a recording medium reading unit (not shown) and stored in a storage unit 54. The computer program 55 may also be read by a recording medium such as a storage device (semiconductor memory such as an SSD (Solid State Drive)) connected by a standard for connecting to a computer (e.g., USB (Universal Serial Bus) or other standards) and stored in a storage unit 54. Alternatively, the computer program 55 may be downloaded from an external device via a communication unit 52 and stored in a storage unit 54. Details of the learning model 56 and the image conversion model 57 will be described later.
[0017] The memory 53 can be composed of semiconductor memory such as SRAM (Static Random Access Memory), DRAM (Dynamic Random Access Memory), or flash memory. The computer program 55 can be loaded into the memory 53, and the control unit 51 can execute the computer program 55. The control unit 51 can execute the processing defined in the computer program 55. In other words, the processing performed by the control unit 51 is also the processing performed by the computer program 55.
[0018] Figure 2 shows an example of the configuration of the learning model 56. The learning model 56 is, for example, a neural network model generated by deep learning, and can be constructed as a CNN (Convolutional Neural Network). The learning model 56 includes convolutional layers 561, pooling layers 562, and fully connected layers 563. The learning model 56 includes multiple sets of convolutional layers 561 and pooling layers 562. The convolutional layer 561 can extract local features within an image. Local features include, for example, features such as edges and color changes within the image. The convolutional layer 561 can extract image features while retaining information within the image. The pooling layer 562 removes positional information while maintaining the features extracted by the convolutional layer 561, so as not to be affected even if the position of the feature part moves within the image. The fully connected layer 563 has the function of classifying what the features are by aggregating the features extracted by the convolutional layer 561 into a single node. Furthermore, the learning model 56 is not limited to CNN; it may also be composed of a Vision Transformer using a Transformer, or an architecture such as Mamba.
[0019] As shown in Figure 2, when an image (for example, an image obtained by a cholangioscope) is input to the learning model 56, the learning model 56 can output whether the biliary tract tumor is "benign" or "malignant" (benign or malignant). The image data input to the learning model 56 can be an image obtained by extracting the required region from an captured image. The learning model 56 can also output the accuracy (probability) of the prediction result. For example, if it predicts that the biliary tract tumor is malignant, it can also output the accuracy of that prediction. Although not shown in the figure, the learning model 56 can also output that there is "no tumor" in the bile tract. For convenience, the images obtained by a cholangioscope also include images of the inside of the bile tract obtained by an endoscope without a cholangioscope.
[0020] Furthermore, when the aforementioned image is input to the learning model 56, it can output information regarding the findings of the lesion in the biliary tract. The information regarding the findings of the lesion in the biliary tract can output at least one of the following: redness and a redness evaluation category (for example, divided into 0 to 2), vascular atypia and an evaluation category of vascular atypia 0 to 2, irregular protrusion and an evaluation category of irregular protrusion 0 to 2, and irregularity of the mucosal surface and an evaluation category of irregularity of the mucosal surface 0 to 2. Evaluation categories 0 to 2 represent the degree of abnormality of the lesion, and for example, evaluation category 0 may be normal (no abnormality), evaluation category 1 may be "moderate", and evaluation category 2 may be "severe".
[0021] Redness includes, for example, abnormal blood vessels or prominent pit patterns. Vascular dysplasia includes, for example, abnormal blood vessels. Irregular elevations include, for example, lesions that are lumps larger than one-quarter the diameter of the duct, or nodules or polypoids smaller than one-quarter the diameter of the duct, or papillary projections. Irregularities of the mucosal surface include, for example, cases where the mucosal features are granular. Information regarding findings of lesions in the biliary tract is not limited to the examples given above. For example, it may also include strictures, ulceration, and scarring within the biliary tract.
[0022] As described above, the control unit 51 acquires image data obtained by a cholangioscopy and, upon inputting the image data, can input the acquired image data into a learning model 56 that predicts information regarding the benign or malignant nature of biliary tract tumors and the findings of lesions in the biliary tract.
[0023] This could improve the accuracy of biliary endoscopic diagnoses and potentially solve the problem of a shortage of specialist physicians.
[0024] Furthermore, the control unit 51 can input the acquired image data into the learning model 56 to predict the evaluation category of the information regarding the findings. The information regarding the findings includes at least one of the following: redness, vascular dysplasia, irregular elevation, irregular mucosal surface, stenosis, ulcer formation, and scarring.
[0025] This allows for the prediction of findings in biliary tract lesions, including their evaluation categories, which would typically be diagnosed by a specialist physician. As a result, non-specialist physicians can obtain the same level of support from the information processing device 50 as they would from a specialist physician's diagnosis.
[0026] Next, we will explain the learning (generation) method of the learning model 56.
[0027] Figures 3A and 3B show a first example of the learning method for the learning model 56. The control unit 51 uses images obtained by cholangioscopy, collected from multiple patients or subjects, as learning image data, and acquires training data using flags indicating whether the biliary tract tumors present in the images are benign or malignant as training data.
[0028] As shown in Figure 3A, the control unit 51 inputs training image data containing benign tumors into the training model 56 and trains the training model 56 by adjusting its parameters so that the prediction result output by the training model 56 as output data approaches the "benign" training data. Note that the training image data containing benign tumors also includes image data that does not contain tumors. For example, image data that does not contain tumors but includes redness is included in the training image data containing benign tumors. Using the dataset of training image data and training data, the training is repeated until the difference between the prediction result and the training data is within an acceptable range.
[0029] Furthermore, as shown in Figure 3B, the control unit 51 inputs training image data containing malignant tumors into the training model 56 and trains the training model 56 by adjusting its parameters so that the prediction result output by the training model 56 as output data approaches the "malignant" data used as training data. The training process is repeated using the training image data and training data dataset until the difference between the prediction result and the training data falls within an acceptable range.
[0030] By training the learning model 56 in the manner described above, the learning model 56 can, as shown in Figure 2, output whether a tumor of the bile duct is benign or malignant when it is input with an image obtained by a cholangioscopy.
[0031] Figures 4A, 4B, and 4C show a second example of the learning method for the learning model 56. The control unit 51 acquires training data using images obtained by cholangioscopy collected from multiple patients or subjects as learning image data, and information on the findings of lesions present in the images as training data. The information on the findings of lesions used as training data can be, for example, "no redness" (evaluation category 0), "redness present" (evaluation category 1), and "redness present" (evaluation category 2).
[0032] As shown in Figure 4A, the control unit 51 inputs training image data without redness (evaluation category 0) to the training model 56 and trains the training model 56 by adjusting the parameters of the training model 56 so that the prediction result output by the training model 56 as output data approaches the training data of "no redness (evaluation category 0)". Using the training image data and training data dataset, the training is repeated until the difference between the prediction result and the training data is within an acceptable range.
[0033] Furthermore, as shown in Figure 4B, the control unit 51 inputs training image data with redness (evaluation category 1) into the training model 56 and trains the training model 56 by adjusting the parameters of the training model 56 so that the prediction result output by the training model 56 as output data approaches the training data of "redness (evaluation category 1)". Using the dataset of training image data and training data, the training is repeated until the difference between the prediction result and the training data is within an acceptable range.
[0034] Furthermore, as shown in Figure 4C, the control unit 51 inputs training image data with redness (evaluation category 2) into the training model 56 and trains the training model 56 by adjusting its parameters so that the prediction result output by the training model 56 as output data approaches the training data of "redness present (evaluation category 2)". The training is repeated using a dataset of training image data and training data until the difference between the prediction result and the training data is within an acceptable range. Note that the training image data input to the training model 56 can be an image obtained by cutting out the required region from a captured image.
[0035] By training the learning model 56 in the manner described above, the learning model 56 can output redness (evaluation categories 0-2) when it receives an image obtained by cholangioscopy, as shown in Figure 2. Although not shown in the figure, the learning model can also be similarly trained on vascular atypia and evaluation categories 0-2 for vascular atypia, irregular protrusions and evaluation categories 0-2 for irregular protrusions, and irregular mucosal surface and evaluation categories 0-2 for irregular mucosal surface.
[0036] As described above, the control unit 51 acquires image data obtained by cholangioscopy, and training data including information on whether the biliary tract tumor is benign or malignant and the findings of the lesion in the biliary tract corresponding to the image data. Based on the acquired training data, it can generate a learning model 56 that predicts whether the biliary tract tumor is benign or malignant and the findings of the lesion in the biliary tract when image data is input.
[0037] Next, we will explain data augmentation of the dataset used to train the learning model 56, in order to improve the accuracy of the prediction results of the learning model 56 and to compensate for the small number of biliary endoscopic images used for training. First, we will explain image preprocessing using a technique called Depth-Anything.
[0038] Figure 5 shows an example of a raw image input to the learning model 56 and a visualized image that visualizes the basis for the learning model 56's prediction result. The raw image is an image obtained by cholangioscopy. In the raw image, the rectangular frame indicates the lesion. The visualized image can be generated using GradCAM or similar software. In the visualized image, the rectangular frame indicates the basis for the prediction. The visualization result shows that the model is reacting to the luminal portion of the bile duct, not the lesion. Thus, the learning model 56 does not use the actual lesion as the basis for its prediction result, but rather uses a part other than the lesion (the luminal portion) as the basis. Therefore, it is necessary to improve the prediction accuracy of the learning model 56.
[0039] Figure 6 shows an example of the original image and depth image. The depth image can be obtained by applying a technique called Depth-Anything to the original image. Depth generally indicates the distance from the camera to the subject, and Depth-Anything is a monocular depth estimation model that estimates depth from a single image. As shown in Figure 6, in the depth image, areas that are close are bright, and areas that are far away are dark.
[0040] Figure 7 shows an example of preprocessing applied to the original image. The control unit 51 calculates the average color of the original image based on the pixel values of the original image. Based on the depth image, the control unit 51 generates a converted image by replacing the pixels in the deep parts of the original image with the calculated average color. The parts of the original image that are not deep are left as they are.
[0041] As described above, image data is generated in which the pixels of the luminal portion of the image obtained by cholangioscopy are set to predetermined pixel values, and this generated image data can be included in the training data for generating (training) the learning model 56. The predetermined pixel values include the average color of the image obtained by cholangioscopy. In the example above, an example using Depth-Anything was described as a method for identifying parts other than the lesion (luminal portion), but the method for identifying the luminal portion is not limited to Depth-Anything, and other models or methods that can determine depth from a single image may be used.
[0042] When the learning model 56 makes a diagnosis of benign or malignant conditions based on the images obtained by the choledochoscope, there is a possibility of making an incorrect diagnosis by using the lumen of the bile duct as the basis for prediction instead of the normal mucosa or the diseased mucosa. Also, when there is a stenosis in the distal bile duct or the like (which is common in malignant tumors), it is not possible to insert the choledochoscope up to the intrahepatic bifurcation or the like. Therefore, for malignant tumors, there are few images of the intrahepatic bifurcation or the like. However, according to the present embodiment, when the learning model 56 inputs image data in which the pixels of the lumen portion of the images obtained by the choledochoscope are set to a predetermined pixel value, the learning model 56 is trained to predict information regarding the differentiation between benign and malignant of biliary tumors and the findings of the diseased portion of the bile duct. As a result, the brightness of the lumen portion can be increased so that there is no sense of incongruity in the entire image, preventing the learning model 56 from mistakenly using the lumen portion as the basis for prediction and improving the prediction accuracy of the learning model 56.
[0043] Also, when visualizing the basis for prediction by GradCAM of the learning model 56 learned (generated) using a learning dataset data-augmented using a technique called Depth-Anything, a heatmap is obtained that avoids obstacles such as guidewires shown in the image, and the technique of replacing the lumen portion with an average color (that is, the processing of the lumen in an endoscopic image) is useful.
[0044] Next, as an example of data augmentation, generation of pseudo-stained images will be described.
[0045] FIGS. 8A and 8B are diagrams showing an example of an unstained image and a stained image. The unstained image is an image showing the state of the living body before the staining solution is sprayed, and the stained image is an image showing the state of the living body after the staining solution is sprayed. The stained image obtained by the choledochoscope is obtained, for example, by a contrast method. The contrast method uses, for example, indigo carmine to emphasize and observe the unevenness of the tissue surface (lesion). However, it takes time and cost to spray the staining solution into the living body, imposing a large burden on the patient or the subject, and also increasing the burden on the medical staff. Also, the number of stained images that can be used for learning is limited. Therefore, in the present embodiment, a pseudo-stained image is generated by using the image conversion model 57 to expand the learning dataset.
[0046] FIGS. 9A and 9B are diagrams showing an example of the configuration of the image conversion model 57. The image conversion model 57 includes a first generator 571, a second generator 572, a first discriminator 573, and a second discriminator 574. The control unit 51 can generate the image conversion model 57 using the method of CycleGAN (Cycle Generative Adversarial Networks).
[0047] CycleGAN is a model that includes two GANs which are adversarial generative networks and has cyclicity in the two GANs. In FIG. 9A, the first generator 571 and the second discriminator 574 constitute the first GAN, and the second generator 572 and the first discriminator 573 constitute the second GAN. FIG. 9A shows the configuration during learning, and FIG. 9B shows the configuration during application. In FIGS. 9A and 9B, the unstained image and the stained image are actually captured images, the generated unstained image is an image generated by the second generator 572, and the generated stained image is an image generated by the first generator 571.
[0048] The first generator 571 generates a generated stained image based on the input unstained image. The generated stained image is not an image of a site actually stained with a staining solution, but an image imitating a stained image by image processing.
[0049] When the generated stained image generated by the first generator 571 is input, the second generator 572 generates a generated unstained image based on the input generated stained image. In the learning phase, various parameters of the image conversion model 57 are optimized so that the unstained image input to the first generator 571 approaches the generated unstained image generated by the second generator 572 as closely as possible.
[0050] The second discriminator 574 discriminates whether the generated stained image generated by the first generator 571 is a stained image (real object) of a site actually stained with a staining solution or a generated stained image (fake object) generated by the first generator 571. The stained image (real object) can be, for example, an image stained with a staining solution containing indigo carmine, but the dye is not limited to indigo carmine.
[0051] The second generator 572 generates an unstained image based on the input stained image. The generated unstained image is not an image of a part that was not actually stained with the staining solution, but rather an image that mimics an unstained image through image processing.
[0052] When the first generator 571 receives a generated unstained image produced by the second generator 572, it generates a generated stained image based on the input generated unstained image. In the learning phase, the various parameters of the image conversion model 57 are optimized so that the stained image input to the second generator 572 approaches the generated stained image produced by the first generator 571 as closely as possible.
[0053] The first discriminator 573 identifies whether the generated unstained image produced by the second generator 572 is a genuine unstained image of an area that is not actually stained, or a fake unstained image produced by the second generator 572.
[0054] The control unit 51 uses multiple unstained images and multiple stained images as training data to machine-learn the image conversion model 57. The unstained images and stained images may be endoscopic images taken of the same biological site or of different biological sites. Furthermore, the biological site is not limited to the biliary tract, but may also be endoscopic images taken of other sites that are relatively easy to obtain, such as the stomach or large intestine. In addition, the image conversion model 57 may generate different models for each different dye.
[0055] Specifically, the control unit 51 optimizes various parameters of the image conversion model 57, for example using backpropagation, so that the first generator 571 generates a generated stained image that is close enough to the real stained image to deceive the second discriminator 574. The control unit 51 also optimizes various parameters of the image conversion model 57, for example using backpropagation, so that the second generator 572 generates a generated unstained image that is close enough to the real unstained image to deceive the first discriminator 573.
[0056] Furthermore, the control unit 51 optimizes various parameters of the image conversion model 57 so that the second discriminator 574 can correctly identify the generated stained image produced by the first generator 571 as "fake". Also, the control unit 51 optimizes various parameters of the image conversion model 57 so that the first discriminator 573 can correctly identify the generated unstained image produced by the second generator 572 as "fake".
[0057] Furthermore, the control unit 51 inputs the generated stained image produced by the first generator 571 based on the unstained image to the second generator 572, and optimizes various parameters of the image conversion model 57 so that the generated unstained image produced by the second generator 572 becomes the unstained image input to the first generator 571.
[0058] Furthermore, the control unit 51 inputs the generated unstained image produced by the second generator 572 based on the stained image to the first generator 571, and optimizes various parameters of the image conversion model 57 so that the generated stained image produced by the first generator 571 becomes the stained image input to the second generator 572.
[0059] As described above, the first generator 571 of the image conversion model 57 can be used as a model for generating a generated stained image. As shown in Figure 9B, in the application phase, when an unstained image is input to the first generator 571, the first generator 571 outputs a generated stained image. That is, the control unit 51 can create a generated stained image of the bile duct by inputting an unstained image obtained by cholangioscopy to the first generator 571.
[0060] As described above, an image transformation model 57 can be trained (generated) using data from parts of the bile duct different from the bile duct, such as the stomach and large intestine (unstained and stained images). Using the trained image transformation model 57, pseudo-stained images can be generated from unstained images obtained by cholangioscopy. This allows for the expansion of the training dataset by generating pseudo-stained images, even when it is difficult to obtain stained images of the bile duct that have actually been stained.
[0061] Figure 10 shows an example of a pseudo-stained image. As shown in Figure 10, compared to an unstained image, the pseudo-stained image emphasizes irregularities caused by lesions on the tissue surface, and the boundaries of cancer cells become clearer. This can improve the prediction accuracy of the learning model 56.
[0062] As described above, the control unit 51 can input image data obtained by the cholangioscopy into the image conversion model 57 and obtain pseudo-stained image data of the bile duct that is converted by the image conversion model 57. The obtained pseudo-stained image data can be included as image data for training the learning model 56 to generate (learn).
[0063] Furthermore, the image conversion model 57 is designed to output stained image data of the stomach or large intestine stained with a staining solution when unstained image data of the stomach or large intestine is input. The bile duct is usually filled with physiological saline or bile, making it difficult to stain with a staining solution. In addition, while indigo carmine solution is commonly sprayed on the mucosal surface in gastrointestinal endoscopy and its safety has been confirmed, its safety for use in the bile duct is unknown, and its effects on the human body are not clear. For this reason, it is difficult to collect stained images of the bile duct. However, according to the image conversion model 57 of this embodiment, stained images of parts other than the bile duct, such as the stomach or large intestine, can be collected as training data, and the image conversion model 57 trained with such training data can generate pseudo-stained images from unstained images of the bile duct. This increases the dataset used to train the learning model 56, and allows the learning model 56 to be trained to improve its prediction accuracy.
[0064] Next, we will explain cutout as an example of data augmentation.
[0065] Figure 11 shows an example of cutout. As illustrated in Figure 5, the learning model 56 may use parts other than the lesion, such as dark areas in the image, as the basis for prediction results, rather than the actual lesion. This is presumed to be because, during the training of the learning model 56, a bias occurs between benign and malignant image data due to the frequent use of "benign" image data with few dark areas. Therefore, in this embodiment, as shown in Figure 11, one or more rectangular black images are embedded in the original image. The color of the embedded image is not limited to black, but may be a color close to black, such as dark gray. Also, it is not limited to rectangles, but may be other shapes, and the number of embedded images can be determined as appropriate. As shown in Figure 11, the learning model 56 is trained using image data in which black rectangles are embedded in the original image. As a result, the learning model 56 comes to recognize dark areas in images obtained by cholangioscopy as common general features and tends to exclude them from the basis for prediction results, thereby improving the prediction accuracy of the learning model 56.
[0066] As described above, the control unit 51 can generate image data in which pixels in multiple regions of the image obtained by the cholangioscopy are set to pixel values of a predetermined color. The predetermined color may be, for example, black or dark gray. The generated image data can be included in the training data for generating (training) the learning model 56.
[0067] Next, we will explain how to visualize the basis for predictions made by the learning model 56 regarding the benign or malignant nature of biliary tract tumors and the findings of lesions in the biliary tract.
[0068] Figure 12 shows an example of visualizing the basis for prediction using GradCAM. As shown in Figure 12, the feature maps of the convolutional layer 561 of the learning model 56 are M1, M2, M3, ..., Mi. i indicates the number of channels. By using i filters in the convolution operation, an i-channel feature map is generated. Preferably, the feature map used is, for example, the final layer of the convolutional layer, which is on the input side of the fully connected layer. This is because the positional information of the image is lost in the fully connected layer, while the final layer of the convolutional layer can abstract the features of the image well.
[0069] The degree to which feature maps M1, M2, M3, ..., Mi influence the prediction results of the learning model 56 is calculated using gradients. The gradients of feature maps M1, M2, M3, ..., Mi are denoted as g1, g2, g3, ..., gi. Gradient calculation is an operation that calculates how much the prediction results change when each element of the feature map changes slightly, and then smooths the result within that feature map.
[0070] Next, the importance of the feature maps is calculated using the gradient calculation results. Specifically, each of the feature maps M1, M2, M3, ..., Mi is multiplied by the gradients g1, g2, g3, ..., gi to perform a weighting calculation. By adding the weighted feature maps g1・M1, g2・M2, g3・M3, ..., gi・Mi and passing them through the ReLU (activation function), a heatmap can be generated that visualizes the basis for the prediction results. Note that by using the ReLU activation function, it is possible to hide parts that negatively affect the prediction results.
[0071] The feature areas indicated by the heatmap information can be represented by the magnitude of the heatmap value, indicating the level of the feature's intensity. The display method (e.g., color or density) can be changed according to the degree of influence on the prediction result. For example, areas with a high influence can be visualized with a reddish color, and areas with a low influence with a bluish color.
[0072] Figures 13A, 13B, and 13C show examples of prediction basis images that visualize the basis for prediction results regarding findings of lesions in the biliary tract. Figure 13A shows a pre-processed image obtained by cholangioscopy, which has been cropped to a predetermined size. Figure 13B shows a prediction basis image that visualizes the basis for prediction results of the learning model 56. The learning model 56 outputs the following as prediction results: redness: evaluation category 1, vascular dysplasia: evaluation category 2, mucosal surface irregularity: evaluation category 2, and irregular protrusion: evaluation category 2. The prediction basis image visualizes, using a heat map, which part of the image the learning model 56 used as the basis for its judgment for each prediction result. Figure 13C is a composite of the prediction basis images that visualized the information regarding the four findings shown in Figure 13B into a single prediction basis image. By compositing the prediction basis images, it becomes easier to understand which part of the abnormality is influencing the prediction result.
[0073] As described above, the control unit 51 can output a prediction basis image that visualizes the basis for the prediction of information regarding the findings. This makes it possible to understand the basis for the prediction result.
[0074] Furthermore, the control unit 51 can output a composite image by combining multiple prediction basis images, each of which visualizes the basis for the prediction of the information related to the findings. This makes it easier to understand which part of the abnormality has the greatest influence on the prediction result for findings in multiple lesions.
[0075] Next, we will explain how to evaluate whether the learning model 56 is making accurate predictions. While experiments can be used to determine whether the prediction results of the learning model 56 are correct, conducting experiments is time-consuming. In this embodiment, the accuracy of the prediction results of the learning model 56 is automatically calculated by a program. The evaluation method for the learning model 56 is described below.
[0076] Figure 14 shows an example of an evaluation method for the learning model 56. First, the control unit 51 extracts the region of interest using GradCAM. Specifically, the control unit 51 inputs the image obtained by the cholangioscopy into the learning model 56 and obtains the prediction basis image for the prediction result from the learning model 56. As shown in Figure 14, the control unit 51 converts the obtained prediction basis image into a binarized image. The control unit 51 encloses the region with high brightness (white area in Figure 14) of the converted binarized image with the largest rectangle and designates this largest rectangle region as the important region.
[0077] On the other hand, for the image input to the learning model 56, a specialist such as a doctor outlines the area where lesions or other abnormalities exist with a correct rectangle and sets this correct rectangle as the correct region. The control unit 51 calculates IoU between the critical region and the correct region. If the critical region is represented by Q and the correct region by P, IoU can be calculated using the formula IoU = {(P∩Q) / (P∪Q)}. IoU is an index that represents how much the correct region P and the critical region Q overlap. If there is a perfect match, IoU = 1.0, and if there is no overlap at all, IoU = 0.0. In other words, the closer the value of IoU is to 1, the more accurate the prediction result of the learning model 56 can be judged to be.
[0078] The control unit 51 can evaluate the performance of the learning model 56 by calculating IoU for all the image data obtained by the cholangioscopy. In this case, the performance of the learning model 56 can be evaluated using the average value of the IoU calculated for each image data.
[0079] As described above, the control unit 51 can acquire prediction basis images that visualize the basis for the prediction results of the learning model 56, extract prediction basis regions based on the acquired prediction basis images, and evaluate the learning model 56 based on the degree of agreement between the extracted prediction basis regions and the pre-generated correct answer regions. This makes it possible to determine whether the learning model 56 is correctly identifying the lesion.
[0080] Next, an example of displaying the prediction results of the learning model 56 will be described. The prediction results can be displayed on a terminal device 10 or the like.
[0081] Figure 15 shows a first example of the prediction result screen 100 of the learning model 56. The prediction result screen 100 displays the patient ID, image area 101, and prediction result area 102. Image area 101 displays an image obtained by cholangioscopy, and on this image, the findings of biliary tract tumors and lesions of the bile tract predicted by the learning model 56 are displayed as rectangular boxes. In the example in Figure 15, rectangular box A indicates a biliary tract tumor, rectangular box B indicates redness, rectangular box C indicates an irregular mucosal surface, and rectangular box D indicates an irregular elevation. For convenience, in the example in Figure 15, redness, irregular mucosal surface, and irregular elevation are schematically represented by ×, △, and □, respectively.
[0082] The prediction result area 102 displays the prediction results of the learning model 56. In the example in Figure 15, corresponding to rectangular boxes A to D, box A displays malignant tumor, a 95% probability of diagnosis, and a figure indicating a 95% probability of diagnosis. Box B displays redness, moderate severity (evaluation category 1), box C displays irregular mucosal surface, severe severity (evaluation category 2), and box D displays irregular elevation, severe severity (evaluation category 2). Note that the notation for the probability of diagnosis may be changed for each range of the probability of diagnosis, for example, "High Confidence" for a probability of diagnosis of 90% to 100%, and "Medium Confidence" for a probability of diagnosis of 70% to 89%. In addition, the severity (severe, moderate, no abnormality) and evaluation categories may be displayed as figures instead of text. For example, bar graphs, pie charts, and band graphs may be used.
[0083] As described above, the control unit 51 can display information regarding the benign or malignant nature of the findings predicted by the learning model 56, the accuracy of the prediction, and the findings predicted by the learning model 56, as well as the evaluation category of that information. For physicians who are not specialists, this provides support similar to that obtained from a specialist's diagnosis.
[0084] The first example shown in Figure 15 illustrates the application of processing by the learning model 56 to already captured images, but the method of applying the learning model 56 is not limited to the first example. For example, processing by the learning model 56 may be applied in real time during video recording with an endoscope. Alternatively, processing by the learning model 56 may be applied when the user performs a predetermined operation (for example, pressing the freeze button) during video recording with an endoscope.
[0085] The physician can move the cursor 103 displayed in the image area 101 to the desired rectangular box, select the rectangular box, and then operate the GradCAM icon 104 to display the prediction result screen 110 shown in Figure 16, which will be described later.
[0086] Figure 16 shows a second example of the prediction result screen 110 of the learning model 56. As shown in Figure 16, the prediction result screen 110 displays the prediction basis image 111 instead of the prediction result area 102. In the prediction result screen 100 shown in Figure 15, the doctor selected the rectangular box C indicating the irregularity of the mucosal surface, so the prediction basis image 111 shows the basis for the prediction of the irregularity of the mucosal surface. Specifically, the prediction basis image 111 displays a heat map showing which part of the image the learning model 56 used as the basis for its prediction. This allows the doctor to confirm the basis for the prediction of the irregularity of the mucosal surface.
[0087] Figure 17 shows a third example of the prediction result screen 120 of the learning model 56. The prediction result screen 120 shown in Figure 17 is displayed when a physician selects multiple rectangular boxes in the prediction result screen 100 shown in Figure 15. In the example in Figure 17, the physician selects rectangular boxes B, C, and D. The prediction basis image 121 displays a composite image created by combining the prediction basis images for redness, irregular mucosal surface, and irregular elevation, which correspond to rectangular boxes B, C, and D. By combining the prediction basis images, it becomes easier to understand which part of the abnormality is affecting the prediction result. If multiple pieces of information regarding findings of the biliary tract lesion are predicted, the composite image exemplified in Figure 17 may be displayed without operating the GradCAM icon 104 mentioned above.
[0088] Figure 18 shows a fourth example of the prediction result screen 130 of the learning model 56. As shown in Figure 18, the superimposed image area 131 displays an image in which the visualization information of GradCAM is reflected in the image obtained by cholangioscopy. Methods for reflecting GradCAM information in the imaging area include, for example, (1) displaying important areas of GradCAM as rectangles, or (2) displaying GradCAM in real time. When displaying in real time, it is expected that the location of lesions can be easily found based on GradCAM during real-time imaging. In the example of Figure 18, an image is displayed in which the composite image exemplified in Figure 17 is superimposed on the image obtained by cholangioscopy (original image). When superimposing the composite image, only the part with the highest brightness in the prediction basis image (heatmap) (i.e., the black part in Figure 18, which is the most important part as the prediction basis) may be superimposed. This can improve the visibility of the superimposed image.
[0089] As described above, the control unit 51 can superimpose and display on the image obtained by the cholangioscopy at least one of the following: visualization information of the basis for predicting whether the lesion is benign or malignant, and visualization information of the basis for predicting the findings of the lesion in the biliary tract. This makes it easier to understand the correspondence between the original image and the composite image.
[0090] Next, we will explain the processing performed by the information processing device 50.
[0091] Figure 19 shows an example of the procedure for generating a learning model 56 by the information processing device 50. The control unit 51 acquires training data including image data obtained by cholangioscopy, image data obtained by preprocessing the said image data, and pseudo-stained image data generated based on the said image data (S11). The preprocessing includes the processes shown in Figures 7 and 11 above.
[0092] The control unit 51 acquires training data including information on whether the biliary tract tumor is benign or malignant, and information on the findings of the lesion in the biliary tract, corresponding to the acquired image data (S12). Based on the acquired training data, the control unit 51 generates a learning model 56 that predicts whether the biliary tract tumor is benign or malignant, and information on the findings of the lesion in the biliary tract, when image data is input (S13). The control unit 51 stores the generated learning model 56 in the storage unit 54 (S14) and terminates the process.
[0093] Figure 20 shows an example of the procedure for prediction processing by the information processing device 50. The control unit 51 acquires image data obtained by the cholangioscopy (S21) and inputs the acquired image data into the learning model 56 (S22). The control unit 51 predicts whether the biliary tract tumor is benign or malignant and information regarding the findings of the lesion in the biliary tract (S23), outputs the prediction result (S24), and terminates the process.
[0094] In this embodiment, the Depth-Anything processing performed during the training of the learning model 56 may be performed on the image data input to the learning model 56 when using the learning model 56 to predict whether a biliary tract tumor is benign or malignant, and information regarding the findings of the lesion in the biliary tract.
[0095] (Note 1) The computer program acquires image data obtained by cholangioscopy and, upon inputting the image data, inputs the acquired image data into a learning model that predicts information regarding the benign or malignant nature of biliary tract tumors and the findings of lesions in the biliary tract, causing the computer to perform the process of predicting information regarding the benign or malignant nature of biliary tract tumors and the findings of lesions in the biliary tract.
[0096] (Note 2) The computer program inputs the acquired image data into the learning model as described in Note 1 and causes the computer to perform a process to predict the evaluation category of the information regarding the findings.
[0097] (Note 3) In Note 1 or Note 2, the computer program includes at least one of the following in relation to the findings: redness, vascular dysplasia, irregular elevation, irregular mucosal surface, stenosis, ulceration, and scarring.
[0098] (Note 4) The computer program causes the computer to perform a process in any one of Notes 1 to 3 to output a prediction basis image that visualizes the prediction basis for the information regarding the findings.
[0099] (Note 5) The computer program causes the computer to perform a process in any one of Notes 1 to 4 to output a composite image, which is a composite image formed by combining multiple prediction basis images that visualize the prediction basis for each piece of information related to the findings.
[0100] (Note 6) The computer program instructs the computer to perform a process in any one of Notes 1 to 5 to superimpose and display on the image obtained by the cholangioscopy at least one of the visualization information of the other predictive basis for benign or malignant and the visualization information of the predictive basis for the information regarding the findings.
[0101] (Note 7) The computer program causes the computer to perform a process in any one of Notes 1 to 6 that displays the distinction between benign and malignant and the accuracy of the prediction made by the learning model, as well as information regarding the findings predicted by the learning model and the evaluation category of the information.
[0102] (Note 8) In any one of Notes 1 to 7, the computer program is trained to predict whether a biliary tract tumor is benign or malignant and to predict information regarding the findings of the lesion in the biliary tract when image data obtained by a cholangioscopy, in which the pixels of the luminal portion are set to predetermined pixel values, is input.
[0103] (Note 9) The information processing device includes a control unit, which acquires image data obtained by a cholangioscopy and, upon input of the image data, inputs the acquired image data to a learning model that predicts information regarding the benign or malignant nature of biliary tract tumors and the findings of lesions in the biliary tract to predict information regarding the benign or malignant nature of biliary tract tumors and the findings of lesions in the biliary tract.
[0104] (Note 10) The information processing method involves acquiring image data obtained by cholangioscopy, and then inputting the acquired image data into a learning model that predicts whether biliary tract tumors are benign or malignant and the findings of lesions in the biliary tract, in order to predict whether biliary tract tumors are benign or malignant and the findings of lesions in the biliary tract.
[0105] (Note 11) The learning model generation method involves acquiring image data obtained by cholangioscopy, and training data including information on whether the biliary tract tumor is benign or malignant and the findings of the lesion in the biliary tract corresponding to the image data, and generating a learning model based on the acquired training data to predict whether the biliary tract tumor is benign or malignant and the findings of the lesion in the biliary tract when image data is input.
[0106] (Note 12) The learning model generation method is as described in Note 11, which generates image data in which the pixels of the luminal portion of the image obtained by cholangioscopy are set to predetermined pixel values, and the generated image data is included in the training data.
[0107] (Note 13) In the learning model generation method, as stated in Note 12, the predetermined pixel value includes the average color of the image obtained by the cholangioscope.
[0108] (Note 14) The learning model generation method is as follows: In any one of Notes 11 to 13, image data obtained by cholangioscopy is input to an image conversion model, and pseudo-stained image data obtained by the image conversion model, which simulates staining the bile duct, is acquired, and the acquired pseudo-stained image data is included in the training data.
[0109] (Note 15) In the learning model generation method described in Note 14, the image conversion model is generated such that when unstained image data of the stomach or large intestine is input, it outputs stained image data of the stomach or large intestine stained with a staining solution.
[0110] (Note 16) The learning model generation method involves generating image data in which pixels in multiple regions of an image obtained by cholangioscopy are set to pixel values of a predetermined color, as described in any one of Notes 11 to 15, and using the generated image data as the image data included in the training data.
[0111] (Note 17) The learning model generation method involves obtaining a prediction basis image that visualizes the basis for the prediction result of the learning model, extracting a prediction basis region based on the obtained prediction basis image, and evaluating the learning model based on the degree of agreement between the extracted prediction basis region and the pre-generated correct answer region.
[0112] 1 Communication network 10 Terminal device 20 Endoscope device 21 Main unit 22 Endoscope 50 Information processing device 51 Control unit 52 Communication unit 53 Memory 54 Storage unit 55 Computer program 56 Learning model 561 Convolutional layer 562 Pooling layer 563 Fully connected layer 57 Image transformation model 571 First generator 572 Second generator 573 First discriminator 574 Second discriminator
Claims
1. A computer program that uses a learning model to predict whether a biliary tract tumor is benign or malignant and the findings of the lesion in the bile duct, based on the image data obtained by a cholangioscopy.
2. The computer program according to claim 1, which causes a computer to perform a process of inputting acquired image data into the learning model and predicting the evaluation category of the information regarding the findings.
3. The computer program according to claim 1, wherein the information relating to the findings includes at least one of redness, vascular dysplasia, irregular elevation, irregular mucosal surface, stenosis, ulceration, and scarring.
4. A computer program according to any one of claims 1 to 3, which causes a computer to perform a process that outputs a prediction basis image that visualizes the prediction basis for the information relating to the aforementioned findings.
5. A computer program according to any one of claims 1 to 3, which causes a computer to perform a process to output a composite image obtained by combining multiple prediction basis images, each of which visualizes the prediction basis for the information relating to the findings.
6. A computer program according to any one of claims 1 to 3, which causes a computer to perform a process of superimposing and displaying, on an image obtained by cholangioscopy, at least one of the visualization information of the other predictive basis for benign or malignant status and the visualization information of the predictive basis for the information regarding the findings.
7. A computer program according to any one of claims 1 to 3, which causes a computer to perform a process that displays the distinction between benign and malignant and the accuracy of the prediction made by the learning model, as well as information regarding the findings predicted by the learning model and the evaluation category of the information.
8. The computer program according to any one of claims 1 to 3, wherein the learning model is trained to predict whether a biliary tract tumor is benign or malignant and the findings of a lesion in the biliary tract when image data obtained by a cholangioscopy, in which the pixels of the luminal portion are set to predetermined pixel values, is input.
9. An information processing device comprising a control unit, the control unit acquires image data obtained by a cholangioscopy, and, upon inputting the image data, inputs the acquired image data into a learning model that predicts information regarding the benign or malignant nature of biliary tract tumors and the findings of lesions in the biliary tract, thereby predicting information regarding the benign or malignant nature of biliary tract tumors and the findings of lesions in the biliary tract.
10. An information processing method that obtains image data obtained by cholangioscopy, and then inputs the obtained image data into a learning model that predicts information regarding the benign or malignant nature of biliary tract tumors and the findings of lesions in the biliary tract.
11. A method for generating a learning model, comprising: acquiring image data obtained by cholangioscopy, and training data including information on whether the biliary tract tumor is benign or malignant and the findings of the lesion in the biliary tract corresponding to the image data; and generating a learning model based on the acquired training data to predict whether the biliary tract tumor is benign or malignant and the findings of the lesion in the biliary tract when image data is input.
12. A method for generating a learning model according to claim 11, comprising generating image data in which pixels of the luminal portion of an image obtained by a cholangioscopy are set to predetermined pixel values, and using the generated image data as image data included in the training data.
13. The method for generating a learning model according to claim 12, wherein the predetermined pixel values include the average color of the image obtained by cholangioscopy.
14. A method for generating a learning model according to any one of claims 11 to 13, comprising inputting image data obtained by a cholangioscope into an image conversion model, obtaining pseudo-stained image data of the bile duct that is pseudo-stained by the image conversion model, and including the obtained pseudo-stained image data as image data in the training data.
15. The learning model generation method according to claim 14, wherein the image conversion model is generated to output stained image data of the stomach or large intestine stained with a staining solution when unstained image data of the stomach or large intestine is input.
16. A method for generating a learning model according to any one of claims 11 to 13, wherein image data is generated in which pixels in multiple regions of an image obtained by a cholangioscope are set to pixel values of a predetermined color, and the generated image data is included in the training data.
17. A method for generating a learning model according to any one of claims 11 to 13, comprising: obtaining a prediction basis image that visualizes the basis for the prediction result of the learning model; extracting a prediction basis region based on the obtained prediction basis image; and evaluating the learning model based on the degree of agreement between the extracted prediction basis region and the pre-generated correct answer region.