Training device, image processing device, training method, image processing method, and storage medium

A multitask learning model for lesion detection in medical images uses biopsy data to enhance accuracy in lesion region and malignancy inference, addressing the incomplete ground truth issue and achieving precise lesion detection and malignancy assessment.

WO2026078862A1PCT designated stage Publication Date: 2026-04-16NEC CORP
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-11
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

Existing machine learning models struggle to accurately learn and infer the malignancy of lesion regions in medical images due to the lack of complete ground truth data for the entire lesion area, as accurate malignancy determination is typically only available from biopsy samples.

Method used

A multitask learning approach is employed to train a machine learning model that outputs both lesion region and malignancy inference results, using ground truth data from biopsy samples to enhance accuracy within the lesion area, with loss calculation focused on the sampled regions.

Benefits of technology

The model achieves high-accuracy inference of lesion regions and malignancy levels, even when complete ground truth data is unavailable, by integrating inference results and focusing on biopsy-determined areas, supporting accurate lesion detection and malignancy assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024036392_16042026_PF_FP_ABST
    Figure JP2024036392_16042026_PF_FP_ABST
Patent Text Reader

Abstract

A training device 1X mainly comprises an acquisition means 30X and a training means 33X. The acquisition means 30X acquires a medical image including a lesion region, first ground truth data indicating the lesion region in the medical image, and second ground truth data indicating the degree of malignancy of a lesion in a partial region of the lesion region. The learning means 33X executes, on the basis of the medical image, the first ground truth data, and the second ground truth data, machine learning of a machine learning model that, when an image is inputted, outputs a first estimation result related to a lesion region in the inputted image and a second inference result related to the degree of malignancy of a lesion in the inputted image.
Need to check novelty before this filing date? Find Prior Art

Description

Learning device, image processing device, learning method, image processing method, and storage medium

[0001] The present disclosure relates to the technical field of a learning device, an image processing device, a learning method, an image processing method, and a storage medium for machine learning of a model for making inferences regarding lesions or inference using such a model.

[0002] Conventionally, techniques related to the detection of lesion regions in endoscopic images using machine learning models have been known. For example, Patent Document 1 discloses an endoscopic examination support device that extracts still images from endoscopic videos and detects lesion regions using a machine learning model based on a neural network.

[0003] International Publication WO2024 / 121886

[0004] In the learning of a machine learning model for inferring the malignancy of a lesion in a lesion region in an endoscopic image, correct answer data indicating the malignancy of the lesion region is required. On the other hand, in general, the accurate malignancy of a lesion can only be obtained for a part of the region sampled from the lesion region, so there is a problem that it is difficult to learn the malignancy within the lesion region.

[0005] One of the objectives of the present disclosure is, in view of the above-described problems, to perform learning of a machine learning model for making inferences regarding the malignancy in a lesion region of a medical image, or to present information regarding the malignancy in a lesion region using a machine learning model learned in such a manner, and to provide a learning device, an image processing device, a learning method, an image processing method, and a storage medium.

[0006] One aspect of the learning device includes: acquisition means for acquiring a medical image including a lesion region, first correct answer data indicating the lesion region in the medical image, and second correct answer data indicating the malignancy of a lesion in a part of the lesion region; and learning means for executing machine learning of a machine learning model that outputs a first inference result regarding the lesion region in the input image and a second inference result regarding the malignancy of the lesion in the input image based on the medical image, the first correct answer data, and the second correct answer data when an image is input.

[0007] One embodiment of an image processing apparatus is an image processing apparatus comprising: an acquisition means for acquiring a medical image of a subject; a machine learning model that has been machine-learned to output a first inference result regarding a lesion area in the input image and a second inference result regarding the malignancy of the lesion in the input image when an image is input; an inference means for acquiring the first and second inference results regarding the medical image based on the medical image; an integration means for generating an integrated inference result by integrating the first and second inference results; and a display control means for displaying information based on the integrated inference result on a display device.

[0008] One aspect of the learning method is a learning method in which a computer acquires a medical image including a lesion region, first ground truth data indicating the lesion region in the medical image, and second ground truth data indicating the malignancy of a lesion in a part of the lesion region, and when an image is input, performs machine learning of a machine learning model that outputs a first inference result regarding the lesion region in the input image and a second inference result regarding the malignancy of the lesion in the input image, based on the medical image, the first ground truth data, and the second ground truth data.

[0009] One aspect of the image processing method is an image processing method which involves acquiring a medical image of a subject, a machine learning model that has been trained to output a first inference result regarding a lesion area in the input image and a second inference result regarding the malignancy of the lesion in the input image when an image is input, and the medical image, acquiring the first and second inference results regarding the medical image, generating an integrated inference result by integrating the first and second inference results, and displaying information based on the integrated inference result on a display device.

[0010] One embodiment of a storage medium is a storage medium that stores a program that causes a computer to perform machine learning on a machine learning model, which acquires a medical image including a lesion region, first ground truth data indicating the lesion region in the medical image, and second ground truth data indicating the degree of malignancy of a lesion in a part of the lesion region, and when an image is input, outputs a first inference result regarding the lesion region in the input image and a second inference result regarding the degree of malignancy of the lesion in the input image, based on the medical image, the first ground truth data, and the second ground truth data.

[0011] Another embodiment of the storage medium is a storage medium that stores a program that causes a computer to perform the following processes: acquiring a medical image of a subject, a machine learning model that has been machine-trained to output a first inference result regarding a lesion area in the input image and a second inference result regarding the malignancy of the lesion in the input image when the image is input, based on the medical image, acquiring the first and second inference results regarding the medical image, generating an integrated inference result by integrating the first and second inference results, and displaying information based on the integrated inference result on a display device.

[0012] One example of the effects of this disclosure is that it is possible to train a machine learning model that performs inferences about the malignancy of lesion areas in medical images, or to present information about the malignancy of lesion areas using such a trained machine learning model.

[0013] This shows the general configuration of the learning system. This shows the hardware configuration of the learning device. This is a functional block diagram of the learning device related to learning the lesion inference model. This is a diagram that schematically shows the processing performed in the model execution unit. This is a diagram that clearly shows the relationship between the lesion region and the sampling region in the second ground truth image. This is an example of a flowchart showing the overview of the processing performed by the learning device. This shows the general configuration of the endoscopic examination system. This shows the hardware configuration of the image processing device. This is an example of a functional block of the processor of the image processing device. This shows an example of the display shown by the display device during endoscopic examination. This is an example of a flowchart showing the overview of the processing performed by the image processing device during endoscopic examination. This is a block diagram of the learning device. This is an example of a flowchart executed by the learning device. This is a block diagram of the image processing device. This is an example of a flowchart executed by the image processing device.

[0014] The following describes embodiments of the learning device, image processing device, learning method, image processing method, and storage medium with reference to the drawings.

[0015] <First Embodiment> (1-1) System Configuration Diagram 1 shows the schematic configuration of the learning system 100. As shown in Figure 1, the learning system 100 is a system that performs machine learning on a machine learning model that makes inferences about lesion areas and malignancy in endoscopic images taken of an organ to be examined (also called a "subject") using an endoscope, and mainly comprises a learning device 1 and a storage device 2.

[0016] The endoscopes covered in this disclosure include, for example, pharyngeal endoscopes, bronchoscopes, upper gastrointestinal endoscopes, duodenal endoscopes, small bowel endoscopes, colonoscopes, capsule endoscopes, thoracos

[0017] The learning device 1 performs multitask learning, learning both the task of inferring image regions suspected of being lesions (also called "lesion regions") from endoscopic images and the task of inferring the malignancy of lesions in endoscopic images using a single model. Hereafter, the machine learning model on which multitask learning is performed in this embodiment will also be called the "lesion inference model." Based on the training data D1 stored in the storage device 2, the learning device 1 performs machine learning (training) of the lesion inference model and updates the model information D2 stored in the storage device 2.

[0018] The memory device 2 is a memory that stores various information necessary for processing the learning device 1. The memory device 2 stores training data D1 and model information D2.

[0019] Training data D1 is data used for machine learning of the lesion inference model by the learning device 1. Training data D1 has multiple records, each record containing a training image, which is an endoscopic image input to the lesion inference model in machine learning, and ground truth data, which indicates the correct answer that the lesion inference model should output when the training image is input to the lesion inference model.

[0020] Here, the lesion inference model outputs a first inference result showing the result for the task of inferring the lesion region in the input endoscopic image when an endoscopic image is input, and a second inference result showing the result for the task of inferring the malignancy of the lesion in the input endoscopic image. In this embodiment, as an example, the first inference result is a heatmap image (confidence map) that shows the confidence level of each pixel in the endoscopic image input to the lesion inference model as being a lesion region, ranging from 0 to 1. The second inference result is a heatmap image that shows the malignancy level of each pixel in the endoscopic image input to the lesion inference model as being ranging from 0 to 1.

[0021] The ground truth data for each record in training data D1 includes a first ground truth image representing the correct first inference result when the endoscopic image of that record is input to the lesion inference model, and a second ground truth image representing the correct second inference result. In this embodiment, as an example, the first ground truth image is assumed to be a binary mask image in which pixels in the lesion area are set to 1 and pixels in the background area that are not lesion areas are set to 0. The second ground truth image is assumed to be a mask image in which pixels other than the area that was actually collected and whose malignancy was determined by biopsy (also called the "collection area") are set to 0. The pixels in the collection area of ​​the second ground truth image are assigned pixel values ​​that represent the malignancy based on the biopsy determination. The first ground truth image is an example of first ground truth data, and the second ground truth image is an example of second ground truth data.

[0022] Model information D2 is information about the lesion inference model and includes the parameters of the lesion inference model. The parameters of the lesion inference model are updated by machine learning performed by the learning device 1. Here, the lesion inference model is a model (engine) that, when an endoscopic image is input to the lesion inference model, outputs a first inference result representing the lesion area in the input image and a second inference result representing the malignancy level in the input image. In other words, the lesion inference model is a model that has learned the relationship between the input image (i.e., the endoscopic image) and the lesion area and malignancy level in that image.

[0023] A lesion inference model is a machine learning model (including statistical models; the same applies hereinafter) having any architecture used for multitasking learning, such as a neural network or a support vector machine. For example, a lesion inference model includes a common hidden layer that performs feature extraction, a head that outputs a first inference result based on the feature data output by the hidden layer, and a head that outputs a second inference result based on the feature data output by the hidden layer. When the lesion inference model is composed of a neural network, the model information D2 includes various parameters (including hyperparameters), such as the layer structure, the neuron structure of each layer, the number and size of filters in each layer, and the weights of each element of each filter.

[0024] The conditions detected as lesion areas by the lesion inference model are exemplified below (a) to (f): (a) Head and neck: pharyngeal cancer, malignant lymphoma, papilloma (b) Esophagus: esophageal cancer, esophagitis, hiatal hernia, Barrett's esophagus, esophageal varices, achalasia, submucosal tumor of the esophagus, benign tumor of the esophagus (c) Stomach: gastric cancer, gastritis, gastric ulcer, gastric polyp, gastric tumor (d) Duodenum: duodenal cancer, duodenal ulcer, duodenitis, duodenal tumor, duodenal lymphoma (e) Small intestine: small intestinal cancer, small intestinal neoplastic disease, small intestinal inflammatory disease, small intestinal vascular disease (f) Large intestine: colon cancer, colonic neoplastic disease, colonic inflammatory disease, colonic polyp, colonic polyposis, Crohn's disease, colitis, intestinal tuberculosis, hemorrhoids

[0025] The storage device 2 may be an external storage device such as a hard disk connected to or built into the learning device 1, a storage medium such as flash memory, or a server device that communicates data with the learning device 1. Furthermore, the storage device 2 may be composed of multiple storage devices, with each of the above-mentioned storage units distributed among them.

[0026] The configuration of the learning system 100 shown in Figure 1 is an example, and various modifications may be made. For example, the learning device 1 and the storage device 2 may be implemented by the same device. In another example, the learning device 1 may be composed of multiple devices. In this case, the multiple devices constituting the learning device 1 exchange information necessary to execute pre-assigned processes between the devices by direct wired or wireless communication or by communication via a network.

[0027] (1-2) Hardware Configuration Diagram 2 shows an example of the hardware configuration of the learning device 1. The learning device 1 includes a processor 11, a memory 12, and an interface 13 as hardware. The processor 11, the memory 12, and the interface 13 are connected via a data bus 19.

[0028] The processor 11 functions as a controller (arithmetic unit) that controls the entire learning device 1 by executing programs stored in memory 12. The processor 11 is, for example, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or a TPU (Tensor Processing Unit). The processor 11 may be composed of multiple processors. The processor 11 is an example of a computer.

[0029] Memory 12 is composed of various volatile and non-volatile memories such as RAM (Random Access Memory), ROM (Read Only Memory), and flash memory. The program executed by the learning device 1 is stored in memory 12. Some of the information stored in memory 12 may be stored in an external storage device such as a storage device 2 that can communicate with the learning device 1, or in a storage medium that is detachable from the learning device 1. Furthermore, memory 12 may store information stored in storage device 2 instead.

[0030] Interface 13 is an interface for electrically connecting the learning device 1 with other devices. These interfaces may be wireless interfaces such as network adapters for wirelessly transmitting and receiving data with other devices, or they may be hardware interfaces for connecting with other devices via cables, etc.

[0031] (1-3) Learning the Lesion Inference Model Next, the learning of the lesion inference model will be explained in detail. In general terms, the learning device 1 uses training data D1, which includes a first ground truth image for the lesion area and a second ground truth image for the malignancy grade, to perform multi-task learning of a lesion inference model that outputs a first inference result for the lesion area and a second inference result for the malignancy grade. Generally, accurate malignancy can only be determined from the biopsy area, so it is not possible to prepare ground truth data that represents the malignancy grade of the entire lesion area, and there is a problem in that it is difficult to learn hierarchical malignancy within the lesion area. Taking the above into consideration, the learning device 1 uses training data D1 to perform multi-task learning of a lesion inference model that outputs a first inference result for the lesion area and a second inference result for the malignancy grade with high accuracy. This makes it possible to suitably learn a model for obtaining highly accurate inference results regarding malignancy within the lesion area, even when using ground truth data that represents the malignancy grade of only a part of the lesion area.

[0032] Figure 3 is a functional block diagram of the learning device 1 related to the learning of a lesion inference model. Functionally, the processor 11 of the learning device 1 has a model execution unit 31, a loss calculation unit 32, and a parameter update unit 33. In Figure 3, blocks where data is exchanged are connected by solid lines, but the combination of blocks where data is exchanged is not limited to Figure 3. The same applies to the diagrams of other functional blocks described later.

[0033] The model execution unit 31 extracts training images to be used for learning from the training data D1 and inputs them into the lesion inference model based on the model information D2 to obtain the first inference result and the second inference result output by the lesion inference model. The model execution unit 31 supplies the obtained first inference result and the second inference result to the loss calculation unit 32.

[0034] The loss calculation unit 32 obtains ground truth data corresponding to the training images input to the lesion inference model from the training data D1, and calculates a loss based on the difference between the inference result output by the lesion inference model supplied by the model execution unit 31 and the obtained ground truth data. In this case, the loss calculation unit 32 calculates a loss based on the difference between the first ground truth image and the first inference result of the obtained ground truth data, and the difference between the second ground truth image and the second inference result of the ground truth data. The difference here corresponds to the difference (sum of) the pixel values ​​between corresponding pixels. The loss calculation unit 32 may calculate the above loss based on any loss function such as mean squared error, mean absolute error, square root of mean squared error, mean squared logarithmic error, Huber loss, Poisson loss, hinge loss, or Kullback-Leibler divergence. For example, the loss calculation unit 32 calculates the loss between the first ground truth image and the first inference result, and the loss between the second ground truth image and the second inference result, respectively, and obtains the sum of these losses as the final loss. The loss calculation unit 32 supplies the calculated loss to the parameter update unit 33.

[0035] Preferably, the loss calculation unit 32, in calculating the loss based on the difference between the second ground truth image and the second inference result, does not include the difference in areas other than the sampling area of ​​the second ground truth image in the loss calculation. In other words, the loss calculation unit 32 calculates the loss without including the difference between the second inference result and the second ground truth image in areas other than the sampling area. This makes it possible to train the lesion inference model to learn the degree of malignancy for sampling areas where the degree of malignancy has been determined by biopsy.

[0036] The parameter update unit 33 determines the parameters of the lesion inference model so as to minimize the loss calculated by the loss calculation unit 32. In this case, the parameter update unit 33 may determine the updated parameters using any parameter update method used in machine learning, such as gradient descent or backpropagation. The parameter update unit 33 then updates the model information D2 by reflecting the determined parameters of the lesion inference model in the model information D2. The parameter update unit 33 performs the update of the model information D2 for each record of the training data D1.

[0037] The model execution unit 31, the loss calculation unit 32, and the parameter update unit 33 can be implemented, for example, by the processor 11 executing a program. Alternatively, the necessary programs may be recorded on any non-volatile storage medium and installed as needed to implement each component. At least a portion of these components may be implemented not only by software programs, but also by any combination of hardware, firmware, and software. At least a portion of these components may also be implemented using a user-programmable integrated circuit, such as an FPGA (Field-Programmable Gate Array) or a microcontroller. In this case, the program composed of the above components may be implemented using this integrated circuit. At least a portion of each component may also be composed of an ASSP (Application Specific Standard Produce), an ASIC (Application Specific Integrated Circuit), or a quantum processor (quantum computer control chip). Thus, each component may be implemented by various hardware. The same applies to other embodiments described later. Furthermore, each of these components may be implemented by the collaboration of multiple computers, for example, using cloud computing technology.

[0038] Figure 4 shows an overview of the process for updating the parameters of the lesion inference model. As shown in Figure 4, the lesion inference model functions as a feature extraction unit 310, a first head unit 311, and a second head unit 312 when the parameters of the model information D2 are applied.

[0039] The feature extraction unit 310 corresponds to an intermediate layer (backbone) common to the two tasks performed by the lesion inference model (a task related to the lesion region and a task related to malignancy), and generates feature data of the training images input to the lesion inference model. For example, the feature extraction unit 310 generates a feature map by performing a convolution operation on the training image, which is the input image. The feature map is a map of feature quantities (feature vectors) and corresponds to a tensor obtained by convolution. The feature extraction unit 310 supplies the generated feature data to the first head unit 311 and the second head unit 312.

[0040] The first head unit 311 corresponds to a separate output layer (head) for tasks related to lesion regions and generates a first inference result from feature data supplied from the feature extraction unit 310. The second head unit 312 corresponds to a separate output layer (head) for tasks related to malignancy and generates a second inference result from feature data supplied from the feature extraction unit 310.

[0041] Subsequently, the loss calculation unit 32 calculates a loss based on the difference between the first inference result output by the first head unit 311 and the first ground truth image corresponding to the input training image, and the difference between the second inference result output by the second head unit 312 and the second ground truth image corresponding to the input training image. The parameter update unit 33 updates the model information D2 based on the calculated loss.

[0042] Here, the first ground truth image is a mask image in which each pixel in the training image input to the lesion inference model is represented by a binary value indicating whether or not it is a lesion area. In this case, white areas represent lesion areas, and black areas represent non-lesion areas. The second ground truth image is a mask image in which only the pixel values ​​in the area where a biopsy was performed are set to a range of 0 to 1 according to the degree of malignancy, and the pixel values ​​in the other areas are set to 0.

[0043] Here, preferably, the loss calculation unit 32 does not include in the loss calculation the difference between the second ground truth image and the second inference result in areas other than the acquisition area of ​​the second ground truth image.

[0044] FIG. 5 is a diagram showing the relationship between the lesion region and the sampling region in the second correct image. In FIG. 5, the region surrounded by the broken line represents the lesion region, and the hatched region within the lesion region represents the sampling region. Here, the region within the lesion region other than the sampling region is called the non-sampling region.

[0045] As shown in FIG. 5, the sampling region corresponds to a part of the region within the lesion region, and since malignancy determination is not performed in other non-sampling regions, the pixel value is 0 in the second correct image (i.e., the malignancy is not substantially set). And the sampling region is the region where the difference between the second correct image and the second inference result is included in the calculation of the loss (the loss-included region in FIG. 5), and the non-sampling region, like the non-lesion region, is the region where the difference between the second correct image and the second inference result is not included in the calculation of the loss (the loss-excluded region in FIG. 5). By calculating the loss in this way, the learning device 1 can train the lesion inference model so as to learn the malignancy with high accuracy for the sampling region whose malignancy is determined by biopsy.

[0046] (1-4) Processing Flow FIG. 6 is an example of a flowchart showing an outline of the processing executed by the learning device 1.

[0047] First, the learning device 1 acquires a set of a learning image and correct data corresponding to one record used for learning the lesion inference model from the training data D1 stored in the storage device 2 (step S11). Here, the correct data includes a first correct image representing the correct answer of the first inference result and a second correct image representing the correct answer of the second inference result.

[0048] Next, the learning device 1 acquires the inference result of the lesion inference model based on the model information D2 stored in the storage device 2 and the learning image acquired in step S11 (step S12). In this case, the learning device 1 acquires the inference result output by the lesion inference model by inputting the learning image acquired in step S11 to the lesion inference model reflecting the model information D2. The learning device 1 acquires, as the inference result, a first inference result regarding the lesion region and a second inference result regarding the malignancy.

[0049] Then, the learning device 1 calculates a loss based on the inference result of the lesion inference model and the ground truth data acquired in step S11 (step S13). In this case, the learning device 1 calculates a loss based on the difference between the first inference result and the first ground truth image, and the difference between the second inference result and the second ground truth image. Next, the learning device 1 updates the parameters of the lesion inference model stored in the model information D2 based on the loss calculated in step S13 (step S14). In this case, the learning device 1 determines the updated parameters of the lesion inference model so as to minimize the loss.

[0050] Then, the learning device 1 determines whether or not to terminate the learning process (step S15). For example, the learning device 1 determines that the learning process should be terminated if it has trained the lesion inference model using all records of the training data D1, or if other predetermined learning termination conditions are met. If the learning device 1 determines that the learning process should be terminated (step S15; Yes), it terminates the flowchart process. On the other hand, if the learning device 1 determines that the learning process should not be terminated (step S15; No), it returns to step S11.

[0051] <Second Embodiment> A second embodiment will be described in which the lesion inference model developed using machine learning in the first embodiment is used in an endoscopic examination. (2-1) System Configuration Figure 7 shows the schematic configuration of the endoscopic examination system 200. As shown in Figure 7, the endoscopic examination system 200 is a system that determines the suspected lesion area and malignancy based on endoscopic images taken with an endoscope and presents the determination result. The endoscopic examination system 200 mainly comprises an image processing device 1A, an endoscope scope 3 connected to the image processing device 1A and handled by an examiner such as a doctor performing the examination or treatment, and a display device 4.

[0052] The image processing device 1A acquires images (also called "endoscopic image Ia") taken by the endoscope scope 3 in a time series from the endoscope scope 3 and displays a screen based on the endoscopic image Ia on the display device 4. The endoscopic image Ia is an image taken at a predetermined frame period during at least one of the insertion or withdrawal process of the endoscope scope 3 into the patient. In this embodiment, the image processing device 1A performs inference regarding the lesion area and malignancy for each endoscopic image Ia in the time series and presents information regarding the lesion area and malignancy for endoscopic image Ia in which the lesion area exists. The image processing device 1A may also function as the learning device 1 of the first embodiment and perform machine learning processing of the lesion inference model before the endoscopic examination.

[0053] The endoscope scope 3 mainly comprises an operating unit 36 ​​for the examiner to make predetermined inputs, a flexible shaft 37 that is inserted into the organ to be imaged of the patient, a tip 38 that incorporates an imaging unit such as a miniature image sensor, and a connecting unit 39 for connecting to the image processing device 1A.

[0054] The configuration of the endoscopic examination system 200 shown in Figure 7 is an example, and various modifications may be made. For example, the image processing device 1A may be configured integrally with the display device 4. In another example, the image processing device 1A may be composed of multiple devices.

[0055] (2-2) Hardware Configuration Diagram 8 shows the hardware configuration of the image processing device 1A. The image processing device 1A mainly includes a processor 21, memory 22, interface 23, input unit 24, light source unit 25, and sound output unit 26. Each of these elements is connected via a data bus 19.

[0056] The processor 21 executes predetermined processes by running programs and other data stored in the memory 22. The processor 21 is a processor such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), or TPU (Tensor Processing Unit). The processor 21 may be composed of multiple processors. The processor 21 is an example of a computer.

[0057] Memory 22 is composed of various volatile memories used as working memory, such as RAM (Random Access Memory) and ROM (Read Only Memory), and non-volatile memory that stores information necessary for processing by the image processing device 1A. Memory 22 may also include an external storage device such as a hard disk connected to or built into the image processing device 1A, or it may include a storage medium such as a removable flash memory. The memory 22 stores a program for the image processing device 1A to execute each of the processes in this embodiment.

[0058] Furthermore, the memory 22 stores model information D2, which is information about a lesion inference model that has been previously machine-learned based on the first embodiment. Model information D2 includes the parameters of the lesion inference model obtained by machine learning. The memory 22 may also optionally include other information necessary for the image processing device 1A to perform each of the processes in this embodiment.

[0059] Interface 23 performs interface operations between the image processing device 1A and an external device. For example, interface 23 supplies display information "Ib" generated by the processor 21 to the display device 4. Interface 23 also supplies light generated by the light source unit 25 to the endoscope scope 3. Interface 23 also supplies an electrical signal indicating the endoscopic image Ia supplied from the endoscope scope 3 to the processor 21. Interface 23 may be a communication interface such as a network adapter for wired or wireless communication with an external device, or it may be a hardware interface compliant with USB, SATA, etc.

[0060] The input unit 24 generates input signals based on operations performed by the examiner. The input unit 24 can be, for example, a button, touch panel, remote controller, or voice input device. The light source unit 25 generates light to be supplied to the tip 38 of the endoscope scope 3. The light source unit 25 may also incorporate a pump for supplying water or air to the endoscope scope 3. The sound output unit 26 outputs sound based on the control of the processor 21.

[0061] (2-3) Functional block diagram 9 shows an example of the functional blocks of the processor 21 of the image processing device 1A. Functionally, the processor 21 has an endoscopic image acquisition unit 40, a lesion inference unit 41, an integration unit 42, and an output control unit 43. In Figure 9, blocks that exchange data are connected by solid lines, but the combination of blocks that exchange data is not limited to this. The same applies to the diagrams of other functional blocks described later.

[0062] The endoscopic image acquisition unit 40 acquires endoscopic images Ia captured by the endoscope scope 3 via the interface 23 at predetermined intervals. The endoscopic image acquisition unit 40 then supplies the acquired endoscopic images Ia to the lesion inference unit 41 and the output control unit 43, respectively. The subsequent processing units then perform the processing described below, using the time interval at which the endoscopic image acquisition unit 40 acquires the endoscopic images Ia as a period.

[0063] The lesion inference unit 41 performs inferences regarding the lesion area and malignancy degree in the endoscopic image Ia supplied from the endoscopic image acquisition unit 40, based on the model information D2. In this case, the lesion inference unit 41 inputs the endoscopic image Ia into a lesion inference model configured by referring to the model information D2, and acquires the first inference result and the second inference result output by the lesion inference model. The lesion inference unit 41 supplies the acquired first inference result and second inference result to the integration unit 42 and the output control unit 43.

[0064] The integration unit 42 generates an inference result (also called the "integrated inference result") by integrating the first inference result and the second inference result. The integrated inference result corresponds to a heat map image that represents the degree of malignancy of the lesion in each pixel of the target endoscopic image Ia, taking into account the inference result regarding the lesion area. The integration unit 42 supplies the generated integrated inference result to the output control unit 43.

[0065] Here, we will explain a specific example of generating integrated inference results. For example, suppose the first inference result is a heatmap image (confidence map) that represents the confidence level of each pixel in the target endoscopic image Ia being a lesion area using a value range of 0 to 1, and the second inference result is a heatmap image that represents the malignancy level of each pixel in the target endoscopic image Ia using a value range of 0 to 1.

[0066] In this case, in the first integration example, the integration unit 42 applies a masking process to the heatmap image corresponding to the second inference result for the lesion region based on the heatmap image corresponding to the first inference result, and generates the heatmap image after the masking process as the integrated inference result. Specifically, the integration unit 42 converts the heatmap image corresponding to the first inference result into a binary mask image that indicates whether or not it is a lesion region based on a predetermined threshold, and performs masking on the heatmap image corresponding to the second inference result using the binary mask image. That is, in this case, the integration unit 42 converts the heatmap image corresponding to the second inference result so that pixels determined to be non-lesion regions in the binary mask image are set to 0. Then, the integration unit 42 acquires the heatmap image obtained by the masking process as the integrated inference result. In this case, the integrated inference result is a heatmap image that represents the degree of malignancy indicated by the second inference result for pixels that were inferred to be lesion regions based on the first inference result.

[0067] In the second integration example, the integration unit 42 generates a heatmap image by multiplying the values ​​of the heatmap image corresponding to the first inference result and the heatmap image corresponding to the second inference result at corresponding pixels. The integration unit 42 then acquires the generated heatmap image as the integrated inference result. In this case, the integrated inference result corresponds to a heatmap image obtained by correcting the malignancy degree shown in the second inference result based on the reliability of the lesion area based on the first inference result.

[0068] The output control unit 43 generates display information Ib based on the latest endoscopic image Ia supplied from the endoscopic image acquisition unit 40, the first and second inference results supplied from the lesion inference unit 41, and the integrated inference result supplied from the integration unit 42. The output control unit 43 then supplies the generated display information Ib to the display device 4, causing the latest endoscopic image Ia and the inference results regarding the lesion to be displayed on the display device 4. If the output control unit 43 determines that a lesion area exists based on at least one of the first inference result or the integrated inference result, it may control the sound output of the sound output unit 26 to output a warning sound or voice guidance to notify the user that a lesion area has been detected. For example, the output control unit 43 may determine that a lesion area has been detected if the number of pixels in the heatmap image shown by the first inference result that have a reliability of being a lesion area of ​​a predetermined threshold or higher is a predetermined number or higher.

[0069] (2-4) Display Example Figure 10 shows an example of a display shown by the display device 4 during an endoscopic examination. The output control unit 43 generates display information Ib based on the latest endoscopic image Ia supplied from the endoscopic image acquisition unit 40, the first and second inference results supplied from the lesion inference unit 41, and the integrated inference result supplied from the integration unit 42. The output control unit 43 then supplies the generated display information Ib to the display device 4, thereby displaying the above-mentioned display screen on the display device 4.

[0070] In this example, the output control unit 43 of the image processing device 1A provides a real-time image display area 70, a lesion malignancy display area 71, and an individual confidence map display area 72 on the display screen.

[0071] Here, the output control unit 43 displays a moving image representing the latest endoscopic image Ia in the real-time image display area 70. Furthermore, in the lesion malignancy display area 71, the output control unit 43 displays a heatmap image corresponding to the integrated inference result. In addition, in the individual confidence map display area 72, the output control unit 43 displays side by side a heatmap image corresponding to the first inference result (lesion area confidence map in Figure 10) and a heatmap image corresponding to the second inference result (malignancy confidence map in Figure 10). The heatmap image in the lesion malignancy display area 71 corresponds to an integrated image of the lesion area confidence map and malignancy confidence map on the individual confidence map display area 72.

[0072] In this display example, when an endoscopic image Ia containing a lesion is obtained, the image processing device 1A can present the examiner with a heatmap image that accurately represents the degree of malignancy in the lesion, thereby supporting the examiner's decision-making.

[0073] (2-5) The processing flow diagram 11 is an example of a flowchart that shows an overview of the processing performed by the image processing device 1A during an endoscopic examination.

[0074] First, the image processing device 1A acquires the endoscopic image Ia (step S21). In this case, the image processing device 1A receives the endoscopic image Ia from the endoscope scope 3 via the interface 13.

[0075] Next, the image processing device 1A inputs the endoscopic image Ia acquired in step S11 into the lesion inference model and obtains a first inference result and a second inference result from the lesion inference model (step S22). In this case, the image processing device 1A inputs the endoscopic image Ia into the lesion inference model configured by referring to the model information D2 and obtains the first inference result and the second inference result output from the lesion inference model.

[0076] Next, the image processing device 1A generates an integrated inference result by combining the first inference result and the second inference result (step S23). Then, the image processing device 1A displays information based on the endoscopic image Ia acquired in step S21 and the integrated inference result generated in step S23 (step S24).

[0077] Then, after step S24, the image processing device 1A determines whether or not the endoscopic examination has been completed (step S25). For example, the image processing device 1A determines that the endoscopic examination has been completed when it detects a predetermined input to the input unit 24 or the operation unit 36. If the image processing device 1A determines that the endoscopic examination has been completed (step S25; Yes), it terminates the processing of the flowchart. On the other hand, if the image processing device 1A determines that the endoscopic examination has not been completed (step S25; No), it returns to step S21.

[0078] In the first and second embodiments, the image processing device trained a machine learning model using endoscopic images obtained in an endoscopic examination as input and performed inference using the machine learning model. Alternatively, the image processing device may train a machine learning model (lesion inference model) using arbitrary medical images as input and perform inference using the machine learning model. In this case, "medical image" refers to an image obtained by an examination targeting any organ of the subject. Examples of medical images include CT images obtained by CT examinations, MRI images obtained by MRI examinations, endoscopic images obtained by endoscopic examinations, images obtained by X-ray examinations, images obtained by ultrasound (echo) examinations, and images obtained by other examinations. Thus, even when using arbitrary medical images, the image processing device can, based on the first embodiment, train a lesion inference model using medical images as input, or, based on the second embodiment, obtain inference results regarding lesions by inputting medical images into a trained lesion inference model.

[0079] <Third Embodiment> Figure 12 is a block diagram of the learning device 1X. The learning device 1X mainly comprises an acquisition means 30X and a learning means 33X. The learning device 1X can be, for example, the learning device 1 of the first embodiment. The learning device 1X may be composed of multiple devices.

[0080] The acquisition means 30X acquires a medical image including the lesion region, first ground truth data indicating the lesion region in the medical image, and second ground truth data indicating the malignancy of the lesion in a part of the lesion region. The acquisition means 30X is implemented, for example, in the first embodiment, by a model execution unit 31 that acquires a training image from training data D1, and a loss calculation unit 32 that acquires the first ground truth image and the second ground truth image from training data D1.

[0081] The learning means 33X performs machine learning on a machine learning model that outputs a first inference result regarding the lesion area in the input image and a second inference result regarding the malignancy of the lesion in the input image, based on the medical image, the first ground truth data, and the second ground truth data. The learning means 33X is implemented, for example, by the model execution unit 31, the loss calculation unit 32, and the parameter update unit 33 in the first embodiment.

[0082] Figure 13 is an example of a flowchart executed by the learning device 1X. The acquisition means 30X acquires a medical image including a lesion region, first ground truth data indicating the lesion region in the medical image, and second ground truth data indicating the malignancy of a part of the lesion region (step S31). The learning means 33X, upon receiving an image, performs machine learning on a machine learning model that outputs a first inference result regarding the lesion region in the input image and a second inference result regarding the malignancy of the lesion in the input image, based on the medical image, the first ground truth data, and the second ground truth data (step S32).

[0083] According to the third embodiment, the learning device 1X can learn a machine learning model that can suitably acquire information regarding the degree of malignancy within the lesion area.

[0084] <Fourth Embodiment> Figure 14 is a block diagram of the image processing device 1Y. The image processing device 1Y mainly includes an acquisition means 40Y, an inference means 41Y, an integration means 42Y, and a display control means 43Y. The image processing device 1Y can be, for example, the image processing device 1A in the second embodiment. The image processing device 1Y may be composed of multiple devices.

[0085] The acquisition means 40Y acquires medical images of the subject. The acquisition means 40Y can be the medical image acquisition unit 40 in the second embodiment. The inference means 41Y acquires first and second inference results regarding the medical image based on the medical image and a machine learning model that has been machine-trained to output a first inference result regarding the lesion area in the input image and a second inference result regarding the malignancy of the lesion in the input image when an image is input. The inference means 41Y can be the lesion inference unit 41 in the second embodiment.

[0086] The integration means 42Y generates an integrated inference result by integrating the first inference result and the second inference result. The integration means 42Y can be, for example, the integration unit 42 in the second embodiment. The display control means 43Y causes the information based on the integrated inference result to be displayed on the display device. The display control means 43Y can be the output control unit 43 in the second embodiment.

[0087] Figure 15 is an example of a flowchart executed by the image processing device 1Y. The acquisition means 40Y acquires a medical image of a subject (step S41). The inference means 41Y acquires a first inference result and a second inference result regarding the medical image based on the medical image and a machine learning model that has been machine-learned to output a first inference result regarding the lesion area in the input image and a second inference result regarding the malignancy of the lesion in the input image when an image is input (step S42). The integration means 42Y generates an integrated inference result by integrating the first inference result and the second inference result (step S43). The display control means 43Y causes the information based on the integrated inference result to be displayed on the display device (step S44).

[0088] According to the fourth embodiment, the image processing device 1Y can suitably present to the user the inference results regarding the lesion area and malignancy in the acquired medical image.

[0089] In each of the embodiments described above, the program can be stored using various types of non-transitory computer-readable medium and supplied to a computer, such as a processor. Non-transitory computer-readable mediums include various types of tangible storage mediums. Examples of non-transitory computer-readable mediums include magnetic storage mediums (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage mediums (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). The program may also be supplied to the computer by various types of transient computer-readable mediums. Examples of transient computer-readable mediums include electrical signals, optical signals, and electromagnetic waves. Transitory computer-readable mediums can supply the program to the computer via wired communication channels such as electric wires and optical fibers, or via wireless communication channels.

[0090] Furthermore, some or all of the above embodiments (including modifications, the same applies hereinafter) may also be described as follows, but are not limited to the following. Moreover, not limited to the devices, methods, and storage media described in the following appendices, some or all of the configurations described in the appendices may also be subordinate to various hardware, software, various recording means for recording software, or systems, without departing from the above embodiments.

[0091] [Note 1] A learning device comprising: an acquisition means for acquiring a medical image including a lesion region, first ground truth data indicating the lesion region in the medical image, and second ground truth data indicating the malignancy of a lesion in a part of the lesion region; and a learning means for performing machine learning on a machine learning model that outputs a first inference result regarding the lesion region in the input image and a second inference result regarding the malignancy of the lesion in the input image when an image is input, based on the medical image, the first ground truth data, and the second ground truth data. [Note 2] The learning device according to Note 1, wherein the machine learning is multi-task learning of the machine learning model. [Note 3] The learning device according to Note 1, wherein the learning means determines the parameters of the machine learning model to minimize the loss based on the difference between the first inference result output by the machine learning model when the medical image is input and the first ground truth data, and the difference between the second inference result output by the machine learning model when the medical image is input and the second ground truth data. [Note 4] The learning device according to Note 3, wherein the learning means calculates the loss without including the difference between the second inference result and the second correct data in areas other than the partial area in the loss. [Note 5] The learning device according to Note 1, wherein the first correct data is a mask image representing the lesion area, and the second correct data is a mask image representing the degree of malignancy for the partial area. [Note 6] An image processing device having: an acquisition means for acquiring a medical image of a subject; a machine learning model that has been machine-learned to output a first inference result regarding the lesion area in the input image and a second inference result regarding the degree of malignancy of the lesion in the input image when an image is input; an inference means for acquiring the first inference result and the second inference result regarding the medical image based on the medical image; an integration means for generating an integrated inference result by integrating the first inference result and the second inference result; and a display control means for displaying information based on the integrated inference result on a display device.[Note 7] The inference means obtains a first image as a first inference result showing the reliability of each pixel of the lesion region in the medical image, and obtains a second image as a second inference result showing the malignancy of each pixel in the medical image, the integration means generates a heatmap image based on the first image and the second image as the integrated inference result, and the display control means causes the heatmap image to be displayed on the display device, as described in Note 6. [Note 8] The integration means generates a second image, which has been masked for the lesion region based on the first image, as the integrated inference result, and the display control means causes the masked second image to be displayed on the display device, as described in Note 7. [Note 9] The integration means generates an image obtained by multiplying the first image and the second image by corresponding pixels as the integrated inference result, and the display control means causes the image generated as the integrated inference result to be displayed on the display device, as described in Note 7. [Note 10] A learning method comprising: a computer acquiring a medical image including a lesion region, first ground truth data indicating the lesion region in the medical image, and second ground truth data indicating the malignancy of a lesion in a part of the lesion region; and performing machine learning of a machine learning model that outputs a first inference result regarding the lesion region in the input image and a second inference result regarding the malignancy of the lesion in the input image when an image is input, based on the medical image, the first ground truth data, and the second ground truth data. [Note 11] An image processing method comprising: acquiring a medical image of a subject; a machine learning model that has been trained to output a first inference result regarding a lesion area in the input image and a second inference result regarding the malignancy of the lesion in the input image when the image is input; the medical image; acquiring the first and second inference results regarding the medical image; generating an integrated inference result by integrating the first and second inference results; and displaying information based on the integrated inference result on a display device.[Note 12] A storage medium storing a program that causes a computer to perform machine learning on a machine learning model that acquires a medical image including a lesion region, first ground truth data indicating the lesion region in the medical image, and second ground truth data indicating the degree of malignancy of the lesion in a part of the lesion region, and when an image is input, outputs a first inference result regarding the lesion region in the input image and a second inference result regarding the degree of malignancy of the lesion in the input image, based on the medical image, the first ground truth data and the second ground truth data. [Note 13] A storage medium storing a program that causes a computer to execute a process to obtain the first and second inference results relating to the medical image, based on the medical image, the first and second inference results relating to the medical image, generate an integrated inference result by integrating the first and second inference results, and display the information based on the integrated inference result on a display device.

[0092] Although the present invention has been described above with reference to embodiments, the present invention is not limited to the above embodiments. Various modifications to the structure and details of the present invention can be made that are understandable to those skilled in the art within the scope of the present invention. That is, the present invention naturally includes the full disclosure, including the claims, and various modifications and alterations that those skilled in the art could make in accordance with the technical idea. Furthermore, each disclosure of the above-mentioned patent documents and other references is incorporated herein by reference.

[0093] 1, 1X Learning device 1A, 1Y Image processing device 2 Storage device 11, 21 Processor 12, 22 Memory 13, 23 Interface 100 Learning system 200 Endoscopy system D1 Training data D2 Model information

Claims

1. A learning device comprising: an acquisition means for acquiring a medical image including a lesion region, first ground truth data indicating the lesion region in the medical image, and second ground truth data indicating the malignancy of a lesion in a part of the lesion region; and a learning means for performing machine learning on a machine learning model that outputs a first inference result regarding the lesion region in the input image and a second inference result regarding the malignancy of the lesion in the input image, based on the medical image, the first ground truth data, and the second ground truth data, when an image is input.

2. The learning device according to claim 1, wherein the machine learning is multi-task learning of the machine learning model.

3. The learning device according to claim 1, wherein the learning means determines the parameters of the machine learning model to minimize the loss based on the difference between the first inference result output by the machine learning model when the medical image is input and the first ground truth data, and the difference between the second inference result output by the machine learning model when the medical image is input and the second ground truth data.

4. The learning device according to claim 3, wherein the learning means calculates the loss without including the difference between the second inference result and the second correct answer data in areas other than the aforementioned partial area.

5. The learning device according to claim 1, wherein the first correct answer data is a mask image representing the lesion region, and the second correct answer data is a mask image representing the degree of malignancy for a part of the region.

6. An image processing apparatus comprising: an acquisition means for acquiring a medical image of a subject; a machine learning model that, when an image is input, outputs a first inference result regarding a lesion area in the input image and a second inference result regarding the malignancy of the lesion in the input image; an inference means for acquiring the first and second inference results regarding the medical image based on the medical image; an integration means for generating an integrated inference result by integrating the first and second inference results; and a display control means for displaying information based on the integrated inference result on a display device.

7. The image processing apparatus according to claim 6, wherein the inference means obtains a first image as a first inference result showing the reliability of each pixel of the lesion region in the medical image, and obtains a second image as a second inference result showing the malignancy of each pixel in the medical image, the integration means generates a heatmap image based on the first image and the second image as the integrated inference result, and the display control means causes the heatmap image to be displayed on the display device.

8. The image processing apparatus according to claim 7, wherein the integration means generates a second image, which is obtained by masking the lesion region based on the first image, as the integrated inference result, and the display control means causes the second image, which has been masked, to be displayed on the display device.

9. The image processing apparatus according to claim 7, wherein the integration means generates an image obtained by multiplying the first image and the second image pixel by pixel as the integrated inference result, and the display control means causes the display device to display the image generated as the integrated inference result.

10. A learning method comprising: a computer acquiring a medical image including a lesion region, first ground truth data indicating the lesion region in the medical image, and second ground truth data indicating the malignancy of a lesion in a portion of the lesion region; and performing machine learning of a machine learning model that outputs a first inference result regarding the lesion region in the input image and a second inference result regarding the malignancy of the lesion in the input image, based on the medical image, the first ground truth data, and the second ground truth data.

11. An image processing method comprising: acquiring a medical image of a subject; a machine learning model that has been trained to output a first inference result regarding a lesion area in the input image and a second inference result regarding the malignancy of the lesion in the input image when the image is input; the medical image; acquiring the first and second inference results regarding the medical image; generating an integrated inference result by integrating the first and second inference results; and displaying information based on the integrated inference result on a display device.

12. A storage medium storing a program that causes a computer to perform machine learning on a machine learning model that acquires a medical image including a lesion region, first ground truth data indicating the lesion region in the medical image, and second ground truth data indicating the degree of malignancy of a lesion in a part of the lesion region, and when an image is input, outputs a first inference result regarding the lesion region in the input image and a second inference result regarding the degree of malignancy of the lesion in the input image, based on the medical image, the first ground truth data, and the second ground truth data.

13. A storage medium storing a program that causes a computer to execute a process to obtain the first and second inference results relating to the medical image, based on the medical image, the first and second inference results relating to the medical image, generate an integrated inference result by integrating the first and second inference results, and display information based on the integrated inference result on a display device.

Citation Information

Patent Citations

  • Computer classification of biological tissues

    JP2021520959A

  • Biopsy prediction and guidance with ultrasound imaging and related devices, systems, and methods

    JP2021528158A

  • Method and system for diagnosis of covid-19 using artificial intelligence

    US20220039767A1

  • Diagnostic imaging assistance system and diagnostic imaging assistance device

    WO2020026349A1

  • Medical assistance system, medical assistance device, and medical assistance method

    WO2020121906A1