Lesion identification method and lesion identification program
A dual-model approach combining block-level and pixel-level analysis improves lesion identification accuracy in medical images by reducing both omission and false positives.
Patent Information
- Application Number
- PCT/JP2024/000331
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-17
AI Technical Summary
Current machine learning models for identifying lesions in medical images either miss lesions (omission) or incorrectly identify non-lesion regions as lesions (false positives), lacking a highly accurate method for lesion identification.
A dual-model approach using a first machine learning model to identify lesion probability in image blocks and a second model to confirm pixel-level lesions, combining local and global image features for accurate lesion detection.
Enhances lesion identification accuracy by reducing both omission and false positives, providing a more precise method for identifying lesions in medical images.
Smart Images

Figure JP2024000331_17072025_PF_FP_ABST
Abstract
Description
Lesion identification method and lesion identification program
[0001] The present invention relates to a lesion identification method and a lesion identification program.
[0002] Medical images obtained by CT (Computed Tomography) and MRI (Magnetic Resonance Imaging) are widely used to diagnose various diseases. Diagnostic imaging using medical images requires doctors to interpret a large number of images, placing a heavy burden on them. Therefore, there is a demand for technology that can somehow support doctors' diagnostic work using computers.
[0003] An example of such a technique is a technique for processing medical images using a trained model (machine learning model) generated by machine learning. For example, a method has been proposed for segmenting an object region image containing an object region from an object image using a segmentation model trained by a machine learning process.
[0004] As an example of image recognition technology, an image processing device has been proposed that generates an attribute score map representing the attributes of regions in the input image for each attribute based on the processing results of the input image using a hierarchical neural network, and then integrates each attribute score map to generate and output recognition results for the recognition target.
[0005] International Publication No. 2021 / 202204 Japanese Patent Application Laid-Open No. 2022-173399
[0006] There is a technology that uses machine learning models to identify whether an area on a medical image is a lesion area. For example, there is a machine learning model that divides a medical image into image blocks of a certain size and identifies whether each image block is a lesion area. While this machine learning model can identify lesions with few omissions, it may erroneously identify non-lesion areas as lesions.
[0007] Another example is a machine learning model that, when a single medical image is input, identifies whether or not each pixel in the image is a lesion. This machine learning model identifies lesions based on the image features of the entire medical image, such as the shape of the lesion and its position relative to the organ, and therefore is less likely to mistakenly identify a non-lesion area as a lesion compared to the aforementioned machine learning model. On the other hand, it is more likely to mistakenly identify a lesion area as a non-lesion compared to the aforementioned machine learning model.
[0008] As described above, each machine learning model for identifying lesions has its own advantages and disadvantages, and there is a problem that there is no highly accurate machine learning model with few disadvantages. In one aspect, the present invention aims to provide a lesion identification method and a lesion identification program that can identify lesion areas in medical images with high accuracy.
[0009] In one proposal, a lesion identification method is provided in which a computer acquires a first tomographic image of the inside of a first human body, inputs the first tomographic image into a first machine learning model that divides the input tomographic image into first image blocks of a fixed size and identifies whether or not each image block is an area of a specific lesion, thereby acquiring a first probability that each first image block included in the first tomographic image is an area of a specific lesion, generates first probability data including a second probability for each pixel included in the first tomographic image based on the acquired first probability, and inputs the first tomographic image and the first probability data into a second machine learning model that identifies whether or not each pixel included in the input tomographic image is an area of a specific lesion.
[0010] In addition, one proposal provides a lesion identification program that causes a computer to execute processing similar to the lesion identification method described above.
[0011] In one aspect, regions of lesions in medical images can be identified with high accuracy. These and other objects, features, and advantages of the present invention will become apparent from the following description taken in conjunction with the accompanying drawings illustrating preferred embodiments of the present invention by way of example.
[0012] FIG. 1 is a diagram illustrating an example of the configuration and processing of a lesion identification device according to a first embodiment. FIG. 2 is a diagram illustrating an example of the configuration of a diagnosis support system according to a second embodiment. FIG. 3 is a diagram illustrating an example of the hardware configuration of a lesion identification device. FIG. 4 is a diagram illustrating a first example of a lesion identification model. FIG. 5 is a diagram illustrating a second example of a lesion identification model. FIG. 6 is a diagram illustrating lesion identification processing according to the second embodiment. FIG. 7 is a diagram illustrating an example of the configuration of processing functions provided in each device of the diagnosis support system. FIG. 8 is a diagram illustrating an example of data stored in a learning data storage unit. FIG. 9 is a diagram illustrating an example of data stored in a model data storage unit. A flowchart illustrating an example of learning processing by a learning processing device. A flowchart illustrating an example of score map generation processing. A flowchart illustrating an example of lesion identification processing by a lesion identification device. A diagram illustrating an example of a screen display of lesion identification results. A diagram illustrating lesion identification processing according to a third embodiment. A flowchart illustrating an example of learning processing according to the third embodiment. A flowchart illustrating an example of lesion identification processing according to the third embodiment.
[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. [First Embodiment] Fig. 1 is a diagram showing an example of the configuration and processing of a lesion identification device according to a first embodiment. The lesion identification device 1 shown in Fig. 1 is an information processing device that identifies whether an image region included in a tomographic image 2 captured inside a human body is a region of a specific lesion (lesion region). The tomographic image 2 is a medical image captured by a CT device, MRI device, or the like. In the following description, it is assumed, as an example, that the tomographic images 2 are images captured at regular intervals in the height direction of the human body.
[0014] The lesion identification device 1 includes a storage unit 1a and a processing unit 1b. The storage unit 1a is a storage area allocated in a storage device (not shown) included in the lesion identification device 1. The storage unit 1a stores, for example, model parameters indicating a machine learning model that executes the lesion identification process. The processing unit 1b is, for example, a processor. In this case, the following processing of the processing unit 1b is realized, for example, by the processor executing a predetermined program. The processing unit 1b executes the following lesion identification process.
[0015] When the processing unit 1b acquires the tomographic image 2, it divides the tomographic image 2 into image blocks of a fixed size and inputs each divided image block to the machine learning model 4a to acquire the probability that each image block is a lesion area. This machine learning model 4a is a trained model that identifies whether or not each image block of the input tomographic image is a lesion area. The machine learning model 4a is formed using, for example, a neural network. In this case, the processing unit 1b can acquire a value indicating the probability that an image block is a lesion area from the final layer (output layer) of the neural network.
[0016] Next, based on the acquired probabilities, the processing unit 1b generates probability data 3 including a probability for each pixel included in the tomographic image 2. For example, the processing unit 1b sets the probability of each pixel included in a certain image block in the tomographic image 2 to the probability value acquired from the machine learning model 4a for that image block. Note that the probability data 3 may be generated as, for example, a probability image (probability map) in which a probability value is mapped for each pixel in the tomographic image 2.
[0017] Next, the processing unit 1b inputs the tomographic image 2 and the probability data 3 into a machine learning model 4b, which identifies whether or not each pixel included in the input tomographic image is a lesion area, and thereby identifies whether or not each pixel included in the tomographic image 2 is a lesion area. This machine learning model 4b uses, for example, multiple tomographic images of one or more human bodies and probability data corresponding to each tomographic image as input data for learning, and generates the model by machine learning using labels for each pixel in these multiple tomographic images as ground truth data. The learning probability data is generated based on the probability calculated using the first machine learning model 4a for each image block obtained by dividing each of the multiple tomographic images into the above-mentioned fixed size. The label indicates whether or not each pixel in the multiple tomographic images is a lesion area.
[0018] Here, the machine learning model 4a identifies lesions based on local image features in the tomographic image 2. Therefore, it can be said that the machine learning model 4a identifies lesions based on simple image features such as brightness of the image. While such a machine learning model 4a can identify lesion areas with little omission, it may erroneously identify non-lesion areas as lesion areas. The above probability data 3 can be said to be the identification results for each pixel that reflect the advantages and disadvantages of such a machine learning model 4a.
[0019] On the other hand, as a comparative example, consider a machine learning model that, when a single tomographic image 2 is input, identifies whether or not each pixel of that tomographic image 2 is a lesion. Compared to the above-described machine learning model 4a, this machine learning model is able to identify lesion areas in units of smaller areas within the tomographic image 2. Furthermore, because this machine learning model identifies lesion areas based on the image features of the entire tomographic image 2, such as the shape of the lesion and its positional relationship with the organ, it is less likely to erroneously identify a non-lesion area as a lesion area compared to machine learning model 4a. However, it is more likely to erroneously identify a lesion area as a non-lesion area compared to machine learning model 4a.
[0020] In this embodiment, machine learning model 4b receives not only tomographic image 2 but also probability data 3, allowing lesion identification processing that combines the advantages of the comparative example and machine learning model 4a. That is, probability data 3 reflecting the identification results of machine learning model 4a, which can identify lesion areas with minimal omissions, is input to machine learning model 4b along with tomographic image 2. This reduces the likelihood that a lesion area will be mistakenly identified as a non-lesion area during lesion identification processing by machine learning model 4b, thereby minimizing the disadvantages of the comparative example. Therefore, lesion areas can be identified with high accuracy.
[0021] Second Embodiment Next, a system capable of detecting liver lesion regions from CT images will be described. Fig. 2 is a diagram showing an example of the configuration of a diagnosis support system according to the second embodiment. The diagnosis support system shown in Fig. 2 is a system that supports image diagnosis work using CT imaging, and includes CT devices 11 and 21, a learning processing device 12, and a lesion identification device 22. Note that the lesion identification device 22 is an example of the lesion identification device 1 shown in Fig. 1.
[0022] The CT devices 11 and 21 capture CT images of the human body. In this embodiment, the CT devices 11 and 21 capture a predetermined number of axial tomographic images of the abdominal region including the liver while changing the position (slice position) in the height direction of the human body (direction perpendicular to the axial plane) at predetermined intervals.
[0023] Lesion identification device 22 detects a lesion area from each tomographic image captured by CT device 21. In this embodiment, it is assumed that a cancer area in the liver is detected as the lesion area. Lesion identification device 22 also detects the lesion area using a lesion identification model generated by machine learning. Furthermore, lesion identification device 22, for example, displays information indicating the lesion area identification result on a display device. In this way, lesion identification device 22 supports the image diagnosis work of a user (e.g., a radiologist).
[0024] The learning processing device 12 generates, by machine learning, a lesion identification model to be used by the lesion identification device 22. The learning processing device 12 performs machine learning to generate a lesion identification model using each tomographic image captured by the CT device 11 as learning data. Data (model parameters) indicating the lesion identification model generated by the learning processing device 12 is read into the lesion identification device 22, for example, via a network or via a portable recording medium.
[0025] Note that captured images may be input to the learning processing device 12 and the lesion classification device 22 from the same CT device. Furthermore, the learning processing device 12 may acquire captured images from the CT device via a recording medium, rather than directly. Furthermore, the learning processing device 12 and the lesion classification device 22 may be the same information processing device.
[0026] Fig. 3 is a diagram showing an example of the hardware configuration of a lesion identification device 22. The lesion identification device 22 is realized, for example, as a computer as shown in Fig. 3. As shown in Fig. 3, the lesion identification device 22 includes a processor 201, a random access memory (RAM) 202, a hard disk drive (HDD) 203, a graphics processing unit (GPU) 204, an input interface (I / F) 205, a reader 206, and a communication interface (I / F) 207.
[0027] The processor 201 (processor circuit) comprehensively controls the entire lesion identification device 22. The processor 201 is, for example, a central processing unit (CPU), a micro processing unit (MPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or a programmable logic device (PLD). The processor 201 may also be a combination of two or more elements selected from the CPU, MPU, DSP, ASIC, and PLD. The processor 201 is an example of the processing unit 1b shown in FIG. 1.
[0028] The RAM 202 is used as the main storage device of the lesion identification device 22. The RAM 202 temporarily stores at least a part of the OS (Operating System) program and application programs executed by the processor 201. The RAM 202 also stores various data necessary for processing by the processor 201.
[0029] The HDD 203 is used as an auxiliary storage device for the lesion identification device 22. The OS program, application programs, and various data are stored in the HDD 203. Note that other types of non-volatile storage devices, such as a solid state drive (SSD), can also be used as the auxiliary storage device.
[0030] A display device 204a is connected to the GPU 204. The GPU 204 displays an image on the display device 204a in accordance with an instruction from the processor 201. The display device 204a may be a liquid crystal display or an organic EL (Electroluminescence) display.
[0031] An input device 205a is connected to the input interface 205. The input interface 205 transmits a signal output from the input device 205a to the processor 201. The input device 205a may be a keyboard or a pointing device. The pointing device may be a mouse, a touch panel, a tablet, a touch pad, a trackball, or the like.
[0032] A portable recording medium 206a is detachably attached to the reading device 206. The reading device 206 reads data recorded on the portable recording medium 206a and transmits the data to the processor 201. The portable recording medium 206a may be an optical disk, a semiconductor memory, or the like.
[0033] The communication interface 207 transmits and receives data to and from other devices such as the CT device 21 via the network. The above hardware configuration can realize the processing functions of the lesion identification device 22. The learning processing device 12 can also be realized as a computer with the hardware configuration shown in FIG.
[0034] Next, a comparative example of a lesion identification model generated by machine learning will be described using Figures 4 and 5. Figure 4 is a diagram showing a first example of a lesion identification model. The lesion identification model 30 shown in Figure 4 is a trained model that identifies lesions using the segmentation method. Hereinafter, this lesion identification model 30 may be referred to as the "segmentation model 30." Upon receiving a tomographic image 32 as input, the segmentation model 30 identifies whether or not a specific lesion region exists (in this embodiment, whether or not the region is liver cancer) for each pixel of the tomographic image 32. Here, it is assumed that the tomographic image 32 is an axial plane tomographic image.
[0035] This segmentation model 30 is generated by machine learning (e.g., deep learning) using the same axial plane tomographic image group 31 as training data. A label indicating whether or not a pixel is a lesion area is added to each tomographic image used as training data, and these labels are used as ground truth data during machine learning. Machine learning is performed on a tomographic image basis using such training data to generate the segmentation model 30.
[0036] The segmentation model 30 may be a lesion identification model that identifies three or more lesion types for each pixel, including a type indicating no lesion. For example, when identifying lung opacities, the segmentation model 30 may identify each pixel as one of six opacity types: ground-glass opacity, infiltrate, honeycomb lung, emphysema, other opacities, and normal (no opacity). In this case, a label indicating which of the above opacity types each pixel belongs to is added to each tomographic image used as training data. Lesion identification processing using such a segmentation model 30 is effective for diagnosing, for example, diffuse lung disease.
[0037] FIG. 5 is a diagram showing a second example of a lesion identification model. The lesion identification model 40 shown in FIG. 5 is a trained model that identifies lesions using a classification method. Hereinafter, this lesion identification model 40 may be referred to as the "classification model 40." Upon receiving a tomographic image 42 as input, the classification model 40 divides the tomographic image 42 into patch images (image blocks) of a certain size and uses each divided patch image as a unit to identify whether or not it is a specific lesion area (in this embodiment, whether or not it is liver cancer). Note that, like the tomographic image 32 in FIG. 4, the tomographic image 42 is assumed to be a tomographic image of an axial plane.
[0038] This Classification model 40 is generated by machine learning (e.g., deep learning) using a group 41 of axial plane tomographic images as training data. However, unlike the segmentation method of Fig. 4 , each tomographic image used as training data is divided into the above-mentioned patch images, and a label indicating whether or not it is a lesion area is added to each divided patch image. Then, each patch image is used as training data, and machine learning is performed on a patch image basis using the labels as correct answer data, thereby generating the Classification model 40.
[0039] The classification model 40 may be a lesion classification model that classifies three or more lesion types, including a type indicating a non-lesion, for each patch image. For example, the classification model 40 may classify each patch image as one of the six lung opacity types described above. In this case, a label indicating which of the above opacity types each patch image belongs to is added to each cross-sectional image used as training data.
[0040] In this way, the Classification model 40 identifies lesions using patch images each containing multiple pixels as a unit. On the other hand, the Segmentation model 30 identifies lesions using pixels as units, and therefore has the advantage that, compared to the Classification model 40, it can identify lesions using smaller regions in a tomographic image as units.
[0041] Furthermore, the classification model 40 identifies lesions based on simple image features such as brightness of the image. Therefore, the classification model 40 can, for example, identify lesions that appear black in a tomographic image with little omission. However, the classification model 40 may erroneously identify regions that appear black in a tomographic image but are not lesions as lesions. In other words, the classification model 40 is characterized by a low incidence of missed lesion detections, but a tendency for overdetection.
[0042] On the other hand, the segmentation model 30 identifies lesions based on the image features of the entire tomographic image, such as the shape of the lesion and its positional relationship with organs. Therefore, compared to the classification model 40, which identifies lesions based on local image features within a patch image, the segmentation model 30 can accurately identify lesions that appear black in a tomographic image with fewer errors. However, the segmentation model 30 may identify lesions that appear black in a tomographic image as not being lesions. In other words, the segmentation model 30 is less likely to overdetect lesions than the classification model 40, but is more likely to miss detections.
[0043] 6 is a diagram showing the lesion identification process in the second embodiment. In consideration of the above-described features of the classification model and the segmentation model, the lesion identification device 22 in this embodiment executes the lesion identification process using both the classification model and the segmentation model.
[0044] As shown in Fig. 6, the lesion identification device 22 first divides an input tomographic image 51 into patch images 52 of a certain size. The lesion identification device 22 inputs each of the divided patch images 52 into a classification model 61 to execute lesion identification processing. Similar to Fig. 5, this classification model 61 is generated by machine learning using a group of tomographic images as training data, and identifies whether or not each patch image 52 is a lesion.
[0045] The lesion identification device 22 calculates a score indicating the probability that each patch image 52 is a lesion by performing lesion identification processing using a classification model 61. The classification model 61 is, for example, a machine learning model using a neural network. In this case, the score is output from the final layer (output layer) of the neural network. The score, for example, takes a value between 0 and 1.
[0046] Lesion identification device 22 generates score map 53 using the scores calculated for each patch image 52. The score calculated for each patch image 52 is mapped to each pixel corresponding to that patch image 52 within the area of score map 53. Score map 53 is data in which a score is mapped for each pixel of input tomographic image 51. As described above, the score is a probability, and therefore score map 53 can also be considered a probability image corresponding to tomographic image 51.
[0047] Next, the lesion identification device 22 inputs the input tomographic image 51 and the generated score map 53 into the segmentation model 62, and performs lesion identification processing for each pixel of the tomographic image 51. This segmentation model 62 uses the tomographic image and the score map as input data for learning, and is generated by machine learning using, as ground truth data, a label indicating whether or not each pixel of the tomographic image is a lesion. The identification result output from the segmentation model 62 for each pixel of the tomographic image 51 becomes the final identification result.
[0048] In the above process, the score of each pixel in the score map 53 indicates the probability of it being a lesion calculated using the classification model 61. The score of the score map 53 is likely to show high values in lesion areas, but it may also show high values in non-lesion areas. However, by using both the score map 53 and the tomographic image 51 as input data and executing lesion identification processing using the segmentation model 62, it becomes more likely that pixels that show high scores despite not being lesions can be correctly identified as non-lesions.
[0049] 4, the segmentation model that receives only a tomographic image as input is prone to overlooking lesion detection. However, by inputting score map 53, which reflects the classification results of classification model 61, which is less likely to overlook detection, into segmentation model 62 along with tomographic image 51, the possibility of erroneously classifying pixels of a lesion as not being a lesion is reduced.
[0050] Therefore, the processing by the lesion identification device 22 described above makes it possible to identify with high accuracy whether or not each pixel in the tomographic image 51 is a lesion, compared to when only the Classification model or only the Segmentation model is used.
[0051] 7 is a diagram showing an example of the configuration of the processing functions of each device in the diagnosis support system. First, the learning processing device 12 includes a learning data storage unit 110, a model data storage unit 120, a classification model generation unit 131, a score map generation unit 132, and a segmentation model generation unit 133.
[0052] The learning data storage unit 110 and the model data storage unit 120 are storage areas allocated in a storage device (not shown) included in the learning processing device 12. The learning data storage unit 110 stores learning data used when generating the Classification model 61 and the Segmentation model 62 by machine learning. The model data storage unit 120 stores model data (model parameters) that indicate the Classification model 61 and the Segmentation model 62 generated by machine learning.
[0053] The processing of the classification model generation unit 131, the score map generation unit 132, and the segmentation model generation unit 133 is realized, for example, by a processor (not shown) included in the learning processing device 12 executing a predetermined application program.
[0054] The classification model generation unit 131 generates the classification model 61 by performing machine learning using the learning data for generating the classification model 61 stored in the learning data storage unit 110. The classification model generation unit 131 stores classification model data 61' indicating the generated classification model 61 in the model data storage unit 120. When the classification model 61 is formed by a neural network, the classification model data 61' includes weighting coefficients between nodes on the neural network.
[0055] The score map generating unit 132 generates a score map using the generated Classification model 61 for each tomographic image included in the learning data for generating the Segmentation model 62 stored in the learning data storage unit 110 .
[0056] The segmentation model generation unit 133 generates the segmentation model 62 by performing machine learning using the training data for generating the segmentation model 62 stored in the training data storage unit 110 and the generated score map. The segmentation model generation unit 133 stores segmentation model data 62' indicating the generated segmentation model 62 in the model data storage unit 120. When the segmentation model 62 is formed by a neural network, the segmentation model data 62' includes weighting coefficients between nodes on the neural network.
[0057] The classification model data 61' and segmentation model data 62' generated by the learning processing device 12 are transferred to the lesion identification device 22 via a network or a portable recording medium.
[0058] Next, lesion identification device 22 includes a model data storage unit 210, an organ region identification unit 221, a score map generation unit 222, and a lesion identification processing unit 223. Model data storage unit 210 is a storage area allocated in a storage device provided in lesion identification device 22, such as RAM 202 or HDD 203. Model data storage unit 210 stores data of the organ region identification model, as well as Classification model data 61' and Segmentation model data 62' passed from lesion identification device 22. The organ region identification model is a trained model for identifying an organ region (in this embodiment, the liver) from a set of tomographic images, and is generated in advance by machine learning.
[0059] The processes of the organ region specifying unit 221, the score map generating unit 222, and the lesion identification processing unit 223 are realized, for example, by the processor 201 executing a predetermined application program.
[0060] The organ region identification unit 221 inputs the input tomographic image set to the organ region identification model and identifies the organ region from the tomographic image set. By this processing, tomographic images including the organ region are extracted from the tomographic image set, and the organ region is identified from the extracted tomographic images.
[0061] The score map generation unit 222 inputs a tomographic image including an organ region into the Classification model 61 based on the Classification model data 61′, and calculates a score for each patch image. The score map generation unit 222 generates a score map corresponding to the tomographic image based on the calculated score.
[0062] The lesion identification processing unit 223 inputs the tomographic image including the organ region and the corresponding score map into the segmentation model 62 based on the segmentation model data 62', executes lesion identification processing, and outputs the lesion identification result for each pixel.
[0063] The classification model generation unit 131, the score map generation unit 132, and the segmentation model generation unit 133 may be provided in the same device as the organ region identification unit 221, the score map generation unit 222, and the lesion identification processing unit 223. In other words, the learning process of the classification model 61 and the segmentation model 62 and the lesion identification process using the classification model 61 and the segmentation model 62 may be executed in the same device.
[0064] 8 is a diagram showing an example of data stored in the learning data storage unit 110. The learning data storage unit 110 stores learning data tables 111 and 112. The learning data table 111 stores learning data used to generate the Classification model 61. The learning data table 111 stores multiple sets of learning data for each of multiple cases, each set including image data as input data and correct answer data.
[0065] The image data is data of patch images obtained by dividing a tomographic image into predetermined sizes, and the image data group corresponding to a case is a group of image data of patch images obtained from each of multiple tomographic images obtained by a single imaging of one person. Note that for each patch image, the tomographic image from which it was divided can be identified. In the example of Figure 8, the first four digits of the file name of the image data of the patch image are the same as the file name of the image data of the tomographic image. The correct answer data is a label indicating whether the corresponding patch image is a lesion or not.
[0066] The training data table 111 also registers scores calculated using the generated Classification model 61. The scores are registered in association with the image data of the patch images from which the scores were calculated. As described above, the scores take values between 0 and 1.
[0067] The training data table 112 stores training data used to generate the segmentation model 62. The training data table 112 stores multiple sets of training data including image data, correct answer data, and a score map for each of multiple cases.
[0068] The image data and score map are used as input data during machine learning. The image data is data of tomographic images, and the image data groups corresponding to cases are image data groups of multiple tomographic images obtained by a single imaging session of one person. These tomographic images are the tomographic images from which the patch images registered in the training data table 111 were divided. The ground truth data is data in which a label indicating whether or not each pixel of the corresponding tomographic image is associated with a lesion. A score map is generated for each tomographic image based on the scores registered in the training data table 111 and registered in the training data table 112.
[0069] 9 is a diagram showing an example of data stored in the model data storage unit 210. The model data storage unit 210 stores a model data table 211. In the model data table 211, model structure data and weight data are registered for each of the organ region identification model, the classification model 61, and the segmentation model 62.
[0070] The model structure data indicates the type of machine learning model used and the configuration of the neural network. The model data table 211 may actually store the file name of a file describing such information. The weight data indicates weighting coefficients set between nodes in the neural network of the machine learning model. The model data table 211 may actually store the file name of a file describing such weighting coefficients.
[0071] The model structure data and weight data for the organ region identification model are registered in advance in the model data table 211. The model structure data and weight data for the classification model 61 and the segmentation model 62 are generated by the learning processing device 12 and registered in the model data table 211. In this case, the model structure data and weight data for the classification model 61 correspond to the above-mentioned classification model data 61′, and the model structure data and weight data for the segmentation model 62 correspond to the above-mentioned segmentation model data 62′. However, the model structure data for the classification model 61 and the segmentation model 62 may be registered in advance in the model data table 211.
[0072] The model data storage unit 120 of the learning processing device 12 stores a model data table in which the model structure data and weight data for at least the Classification model 61 and the Segmentation model 62 are registered.
[0073] Next, the processing of the learning processing device 12 will be described with reference to a flowchart. Fig. 10 is a flowchart showing an example of the learning processing by the learning processing device. [Step S11] The classification model generation unit 131 acquires learning data from the learning data table 111.
[0074] [Step S12] The Classification model generation unit 131 performs machine learning using the acquired training data to generate a Classification model 61. In this process, machine learning is performed using each patch image included in the training data as input data for training, and the supervised data corresponding to each patch image as supervised data for training. As a result, Classification model data 61′ is generated and registered in the model data storage unit 120.
[0075] [Step S13] The score map generation unit 132 calculates a score for each patch image by inputting each patch image registered in the learning data table 111 into the generated Classification model 61. The calculated scores are registered in association with the patch images in the learning data table 111. Then, a score map is calculated for each tomographic image based on the registered scores.
[0076] [Step S14] The segmentation model generation unit 133 acquires training data from the training data table 112. [Step S15] The segmentation model generation unit 133 generates a segmentation model 62 by performing machine learning using the acquired training data and the generated score map. In this process, machine learning is performed using each tomographic image included in the training data and the score map corresponding to each tomographic image as input data for training, and the correct answer data corresponding to each tomographic image as correct answer data for training. As a result, segmentation model data 62′ is generated and registered in the model data storage unit 120.
[0077] Fig. 11 is a flowchart showing an example of the score map generation process. The process in Fig. 11 corresponds to the process in step S13 in Fig. 10. [Step S21] The score map generation unit 132 selects one case from the training data table 111.
[0078] [Step S22] The score map generating unit 132 selects one tomographic image included in the selected case. [Step S23] The score map generating unit 132 selects one patch image included in the selected tomographic image.
[0079] [Step S24] The score map generation unit 132 calculates a score by inputting the selected patch image to the Classification model 61. The calculated score is registered in the record in the learning data table 111 that corresponds to the selected patch image.
[0080] [Step S25] The score map generation unit 132 determines whether all patch images included in the selected tomographic image have been selected. If there are any unselected patch images, the process proceeds to step S23, where one unselected patch image is selected. On the other hand, if all applicable patch images have been selected, the process proceeds to step S26.
[0081] [Step S26] The score map generator 132 generates a score map corresponding to the selected tomographic image. In this process, the score calculated for the patch image is mapped to all pixels in the region of the tomographic image that corresponds to the patch image.
[0082] [Step S27] The score map generator 132 determines whether all of the cross-sectional images included in the selected case have been selected. If there are any unselected cross-sectional images, the process proceeds to step S22, where one of the unselected cross-sectional images is selected. On the other hand, if all of the applicable cross-sectional images have been selected, the process proceeds to step S28.
[0083] [Step S28] The score map generation unit 132 determines whether all cases registered in the learning data table 111 have been selected. If there are unselected cases, the process proceeds to step S21, where one of the unselected cases is selected. On the other hand, if all cases have been selected, the score map generation process ends.
[0084] Next, the processing of the lesion identification device 22 will be described using a flowchart. Fig. 12 is a flowchart showing an example of lesion identification processing by the lesion identification device. [Step S31] The lesion identification device 22 acquires from the CT device 21 a set of tomographic images obtained by imaging the subject.
[0085] [Step S32] The organ region identification unit 221 inputs the acquired set of tomographic images into the organ region identification model to identify the organ region from the set of tomographic images. Through this process, tomographic images containing the organ region are extracted from the set of tomographic images, and the organ region is identified from the extracted tomographic images.
[0086] [Step S33] The score map generating unit 222 selects one tomographic image including an organ region from the set of tomographic images. [Step S34] The score map generating unit 222 divides the selected tomographic image into patch images of a predetermined size.
[0087] [Step S35] The score map generation unit 222 selects one of the patch images obtained by the division. In this process, a patch image including an organ region is selected. [Step S36] The score map generation unit 222 inputs the selected patch image into the Classification model 61 to calculate a score.
[0088] [Step S37] The score map generation unit 222 determines whether all patch images included in the selected tomographic image have been selected. If there are unselected patch images, the process proceeds to step S35, where one unselected patch image is selected. On the other hand, if all applicable patch images have been selected, the process proceeds to step S38.
[0089] [Step S38] The score map generating unit 222 generates a score map corresponding to the selected tomographic image. In this process, the score calculated for the patch image is mapped to all pixels in the area of the tomographic image that corresponds to the patch image.
[0090] [Step S39] The lesion identification processor 223 inputs the selected tomographic image and the generated score map into the segmentation model 62, thereby identifying whether or not each pixel in the tomographic image represents a lesion.
[0091] [Step S40] The lesion identification processor 223 determines whether all tomographic images containing organ regions have been selected. If there are unselected tomographic images, the process proceeds to step S33, where one unselected tomographic image is selected. On the other hand, if all applicable tomographic images have been selected, the lesion identification process ends.
[0092] 13 is a diagram showing an example of a screen display of a lesion identification result. The lesion identification device 22 can present the lesion identification result to the user by, for example, displaying an identification result screen 70 as shown in FIG. 13 on the display device 204a. The identification result screen 70 includes a slice selection section 71, a tomographic image display section 72, a lesion type selection section 73, and an identification button 74.
[0093] In the slice selection unit 71, by moving a handle 71a on a slider, the slice plane of the tomographic image to be displayed in the tomographic image display unit 72 can be selected. The tomographic image corresponding to the slice plane selected in the slice selection unit 71 is displayed in the tomographic image display unit 72. A lesion area is displayed on this tomographic image based on the lesion identification result corresponding to the same slice plane. In FIG. 13, as an example, a rectangle 72b surrounding the lesion area 72a is superimposed on the tomographic image. As another display method, the lesion area 72a may be displayed in a specific color. This lesion area 72a is a region of a group of pixels identified as a lesion.
[0094] The lesion type selection unit 73 allows the user to select the type of lesion for which the identification results are to be displayed. Here, for example, it is assumed that a classification model 61 and a segmentation model 62 have been generated for each type of lesion, and in this case, the lesion type selection unit 73 selects the machine learning model to be used for lesion identification. By selecting the type of lesion in the lesion type selection unit 73 and pressing the identification button 74, the identification results using the selected machine learning model are displayed on the tomographic image display unit 72.
[0095] Third Embodiment Next, as a third embodiment, a modified example in which the processing of the diagnosis support system according to the second embodiment is partially modified will be described. In the third embodiment, a plurality of score maps generated using patch images of different sizes are used.
[0096] FIG. 14 is a diagram showing lesion identification processing in the third embodiment. Lesion identification device 22 of this embodiment first divides input tomographic image 51 into patch images of different sizes. In the example of FIG. 14 , tomographic image 51 is divided into patch images 52a of 8 pixels x 8 pixels. Furthermore, tomographic image 51 is divided into patch images 52b of 16 pixels x 16 pixels. Furthermore, tomographic image 51 is divided into patch images 52c of 28 pixels x 28 pixels.
[0097] The lesion identification device 22 inputs each of the divided patch images 52a into a classification model 61a and calculates a score for each patch image 52a. This classification model 61a is a machine learning model trained using 8 pixel x 8 pixel patch images as learning data. The lesion identification device 22 generates a score map 53a using the score calculated for each patch image 52a using the classification model 61a.
[0098] Furthermore, lesion identification device 22 inputs each of divided patch images 52b into classification model 61b and calculates a score for each patch image 52b. This classification model 61b is a machine learning model trained using 16 pixel x 16 pixel patch images as learning data. Lesion identification device 22 generates score map 53b using the score calculated for each patch image 52b using classification model 61b.
[0099] Furthermore, the lesion identification device 22 inputs each of the divided patch images 52c into a classification model 61c and calculates a score for each patch image 52c. This classification model 61c is a machine learning model trained using 28 pixel by 28 pixel patch images as learning data. The lesion identification device 22 generates a score map 53c using the score calculated for each patch image 52c using the classification model 61c.
[0100] The lesion identification device 22 then inputs the input tomographic image 51 and the generated score maps 53a to 53c into the segmentation model 62a and performs lesion identification processing for each pixel of the tomographic image 51. This segmentation model 62a uses, as input data for learning, the tomographic image, a score map generated using the classification model 61a based on an 8-pixel by 8-pixel patch image, a score map generated using the classification model 61b based on a 16-pixel by 16-pixel patch image, and a score map generated using the classification model 61c based on a 28-pixel by 28-pixel patch image, and is generated by machine learning using, as ground truth data, a label indicating whether or not each pixel of the tomographic image is a lesion. The identification result output from the segmentation model 62a for each pixel of the tomographic image 51 becomes the final identification result.
[0101] The above process enables lesion identification based on the features of local regions of various sizes, thereby improving lesion identification accuracy. For example, it becomes possible to accurately identify lesions of various sizes according to the size of the patch image without omission.
[0102] 15 is a flowchart showing an example of a learning process in the third embodiment. [Step S51] The classification model generation unit 131 selects a patch size. [Step S52] The learning data storage unit 110 stores learning data tables 111 corresponding to each patch size. The classification model generation unit 131 identifies the learning data table 111 corresponding to the selected patch size from the learning data storage unit 110, and acquires learning data from the identified learning data table 111.
[0103] [Step S53] The Classification model generation unit 131 performs machine learning using the acquired training data to generate a Classification model 61 corresponding to the selected patch size. In this process, machine learning is performed using each patch image included in the training data as input data for training, and the supervised data corresponding to each patch image as supervised data for training. As a result, Classification model data 61' corresponding to the selected patch size is generated and registered in the model data storage unit 120.
[0104] [Step S54] The score map generation unit 132 inputs each patch image registered in the identified learning data table 111 into the generated Classification model 61, thereby calculating a score for each patch image. The calculated scores are registered in association with the patch images in the learning data table 111. Then, a score map is calculated for each tomographic image based on the registered scores. In this step S54, the process shown in FIG. 11 is executed for each score size.
[0105] [Step S55] The score map generation unit 132 determines whether all patch sizes have been selected. If there are unselected patch sizes, the process proceeds to step S51, where one unselected patch size is selected. On the other hand, if all patch sizes have been selected, the process proceeds to step S56.
[0106] [Step S56] The segmentation model generation unit 133 acquires training data from the training data table 112. [Step S57] The segmentation model generation unit 133 generates a segmentation model 62 by performing machine learning using the acquired training data and the generated score maps. In this process, machine learning is performed using each tomographic image included in the training data and the score maps for each patch size corresponding to each tomographic image as input data for training, and using the correct answer data corresponding to each tomographic image as correct answer data for training. As a result, data indicating the segmentation model 62 is generated and registered in the model data storage unit 120.
[0107] 16 is a flowchart showing an example of the lesion identification process in the third embodiment. [Step S61] The lesion identification device 22 acquires from the CT device 21 a set of tomographic images obtained by imaging the subject.
[0108] [Step S62] The organ region identification unit 221 inputs the acquired set of tomographic images into the organ region identification model to identify the organ region from the set of tomographic images. Through this process, tomographic images containing the organ region are extracted from the set of tomographic images, and the organ region is identified from the extracted tomographic images.
[0109] [Step S63] The score map generating unit 222 selects one tomographic image including an organ region from the set of tomographic images. [Step S64] The score map generating unit 222 selects a patch size.
[0110] [Step S65] The score map generation unit 222 generates a score map corresponding to the selected patch size. This process is executed by applying the selected patch size to the processes of steps S35 to S38 in Fig. 12. In step S36, the Classification model 61 corresponding to the selected patch size is used.
[0111] [Step S66] The score map generation unit 222 determines whether all patch sizes have been selected. If there are unselected patch sizes, the process proceeds to step S64, where one unselected patch size is selected. On the other hand, if all patch sizes have been selected, the process proceeds to step S67.
[0112] [Step S67] The lesion identification processor 223 inputs the selected tomographic image and the generated score map for each patch size into the segmentation model 62a, thereby identifying whether or not each pixel of the tomographic image represents a lesion.
[0113] [Step S68] The lesion identification processor 223 determines whether all tomographic images containing organ regions have been selected. If there are unselected tomographic images, the process proceeds to step S63, where one unselected tomographic image is selected. On the other hand, if all applicable tomographic images have been selected, the lesion identification process ends.
[0114] The processing functions of the devices (e.g., lesion identification devices 1 and 22, learning processing device 12) shown in each of the above embodiments can be realized by a computer. In this case, a program describing the processing content of the functions to be possessed by each device is provided, and the processing functions are realized on the computer by executing the program. The program describing the processing content can be recorded on a computer-readable recording medium. Examples of computer-readable recording media include magnetic storage devices, optical discs, and semiconductor memories. Examples of magnetic storage devices include hard disk drives (HDDs) and magnetic tapes. Examples of optical discs include CDs (Compact Discs), DVDs (Digital Versatile Discs), and Blu-ray Discs (BD, registered trademark).
[0115] When distributing a program, for example, the program is recorded on a portable recording medium such as a DVD or CD and sold. Alternatively, the program can be stored in a storage device of a server computer and transferred from the server computer to other computers via a network.
[0116] A computer that executes a program stores, for example, a program recorded on a portable recording medium or a program transferred from a server computer in its own storage device. The computer then reads the program from its own storage device and executes processing in accordance with the program. Note that the computer can also read the program directly from a portable recording medium and execute processing in accordance with that program. The computer can also execute processing in accordance with the program received each time a program is transferred from a server computer connected via a network.
[0117] The foregoing merely illustrates the principles of the present invention. Further, since numerous modifications and changes will be apparent to those skilled in the art, the present invention is not limited to the exact construction and application shown and described above, and all corresponding modifications and equivalents are deemed to be within the scope of the present invention as defined by the appended claims and their equivalents.
[0118] REFERENCE SIGNS LIST 1 Lesion identification device 1a Memory unit 1b Processing unit 2 Tomographic image 3 Probability data 4a, 4b Machine learning model
Claims
1. A computer acquires a first tomographic image of the inside of a first human body, and inputs the first tomographic image to a first machine learning model that identifies whether or not a specific lesion region exists, with first image blocks obtained by dividing the input tomographic image into blocks of a certain size as units, thereby obtaining a first probability that each of the first image blocks included in the first tomographic image is the specific lesion region. Based on the obtained first probability, the computer generates first probability data including a second probability for each pixel included in the first tomographic image. The computer inputs the first tomographic image and the first probability data to a second machine learning model that identifies whether or not a specific pixel region exists, with pixels included in the input tomographic image as units, thereby identifying whether or not each pixel included in the first tomographic image is the specific lesion region. This is a lesion identification method.
2. The computer further uses, as learning input data, a plurality of tomographic images of the inside of one or more second human bodies, and second probability data including a third probability calculated for each pixel of the plurality of tomographic images based on a second probability calculated for each second image block obtained by dividing each of the plurality of tomographic images into blocks of the certain size using the first machine learning model. The computer generates the second machine learning model by machine learning using, as correct answer data, a label indicating whether or not each pixel of the plurality of tomographic images is the specific lesion region. This is the lesion identification method according to claim 1.
3. In obtaining the first probability, the first tomographic image is input to each of a plurality of the first machine learning models that identify whether or not a specific lesion region exists, with image blocks of different sizes as units, thereby obtaining the first probability for each of the plurality of the first machine learning models. In generating the first probability data, a plurality of the first probability data corresponding to each of the plurality of the first machine learning models are generated based on the first probability generated using the corresponding first machine learning model among the plurality of the first machine learning models. In the identification, the first tomographic image and the plurality of the probability data are input to the second machine learning model, thereby identifying whether or not each pixel included in the first tomographic image is the specific lesion region. This is the lesion identification method according to claim 1.
4. Cause a computer to obtain a first tomographic image obtained by photographing the inside of a first human body, input the obtained tomographic image into a first machine learning model that identifies whether or not a specific lesion area exists in units of first image blocks obtained by dividing the input tomographic image into blocks of a certain size, thereby obtaining a first probability that each of the first image blocks included in the first tomographic image is the specific lesion area, generate probability data including a second probability for each pixel included in the first tomographic image based on the obtained first probability, and input the first tomographic image and the probability data into a second machine learning model that identifies whether or not a specific pixel included in the input tomographic image is the lesion area, thereby causing the computer to execute a process of identifying whether or not each pixel included in the first tomographic image is the specific lesion area.
Citation Information
Patent Citations
Image processing device and image processing method
JP2022173399A
Data processing method, means and system
WO2021202204A1
Focus detection method and device based on lung images
CN111612749A
Control method, information terminal and program
JP2018102916A
Medical image processing device and program
JP2021087729A