Image processing apparatus and image processing method
The image processing apparatus and method use a cascaded neural network to generate attention maps and difference images, addressing the inaccuracy of existing tumor detection methods and enabling precise tumor segmentation for evaluating treatment effects.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2022-05-19
- Publication Date
- 2026-07-29
AI Technical Summary
Existing methods for automatically detecting tumors using deep learning are inaccurate, especially for small tumors, making it difficult to evaluate tumor treatment effects like RECIST, and existing breast cancer MRI segmentation methods fail to accurately detect new tumors compared to previous stages.
An image processing apparatus and method using a cascaded neural network approach, including an attention map generation unit, difference value calculation unit, and first acquisition unit to generate attention maps and difference images from multiple scans, enabling accurate image segmentation and detection of new tumors.
Enables automatic and accurate detection of new tumors by leveraging attention maps and difference images, improving tumor segmentation accuracy and enabling effective evaluation of tumor treatment effects.
Smart Images

Figure 0007897040000001 
Figure 0007897040000002 
Figure 0007897040000003
Abstract
Description
Technical Field
[0001] The embodiments disclosed in this specification and the drawings relate to an image processing apparatus and an image processing method.
Background Art
[0002] In order to grasp the condition of cancer patients, it is necessary to evaluate the treatment effect of the patient's tumor. As an evaluation criterion for the treatment effect of a tumor, RECIST (Response Evaluation Criteria in Solid Tumors) that classifies the patient's target lesion into one of Complete Response, Partial Response, Stable Disease, and Progressive Disease can be mentioned. In RECIST, it is necessary to determine whether a new tumor has appeared in the patient's target lesion compared to the immediately previous stage, and currently this detection is mainly performed by a doctor manually comparing the current scan image and the scan image of the immediately previous stage.
[0003] In the prior art, there is a technique for automatically detecting a tumor based on a scan image of a patient's target lesion using deep learning technology. However, compared with organ detection technology based on deep learning, the accuracy of tumor detection technology based on deep learning is generally low. Especially for small tumors, since the tumor region cannot often be accurately segmented, in the prior art for detecting tumors using deep learning technology, it is difficult to directly apply it to the evaluation of tumor treatment effects such as RECIST.
[0004] To address this challenge, a cascaded neural network segmentation network has been proposed. In this method, the scanned image is first input into a low-resolution first segmentation network to obtain image segmentation results, and these results are then input into a high-resolution second segmentation network along with the scanned image to obtain more accurate image segmentation results. However, the improvement in accuracy obtained by this method is limited and cannot meet the need to evaluate the effectiveness of tumor treatment.
[0005] Furthermore, a breast cancer MRI segmentation method based on a hierarchical convolutional neural network has been disclosed that performs image segmentation on scan images based on pre-enhancement and post-enhancement images to determine whether or not a tumor is present. However, even with this method, it is not possible to automatically and accurately detect whether or not a new tumor has appeared in the patient's target lesion compared to the previous stage. [Prior art documents] [Patent Documents]
[0006] [Patent Document 1] Chinese Patent Application Publication No. 110796672 Specification [Overview of the project] [Problems that the invention aims to solve]
[0007] One of the problems that the embodiments disclosed herein and in the drawings aim to solve is to enable automatic and accurate detection of whether or not a new tumor has appeared in a patient's target lesion. However, the problems that the embodiments disclosed herein and in the drawings aim to solve are not limited to the above problem. Problems corresponding to the effects of each configuration shown in the embodiments described later can also be positioned as other problems. [Means for solving the problem]
[0008] The image processing apparatus according to the embodiment comprises an attention map generation unit, a difference value calculation unit, and a first acquisition unit. The attention map generation unit generates, based on a first image and a second image collected with different modalities for a target lesion of a patient, a first attention map which is the same size as the first image and shows the degree of attention each pixel gives to the pixel at the same position in the first image when performing image segmentation, and a second attention map which is the same size as the second image and shows the degree of attention each pixel gives to the pixel at the same position in the second image when performing image segmentation. The difference value calculation unit generates an attention difference image which shows the difference between the first attention map and the second attention map. The first acquisition unit performs image segmentation on the second image based on at least the first image and the attention difference image and acquires the segmentation result of the second image. [Brief explanation of the drawing]
[0009] [Figure 1] Figure 1 shows an example of the configuration of an image processing apparatus according to the first embodiment. [Figure 2] Figure 2 is a flowchart of the image processing method according to the first embodiment. [Figure 3] Figure 3 shows an example of the first and second images according to the first embodiment. [Figure 4] Figure 4 shows an example of a first attention map and a second attention map according to the first embodiment. [Figure 5] Figure 5 shows an example of an attention difference image according to the first embodiment. [Figure 6] Figure 6 shows an example of the final result of image segmentation of the second image according to the first embodiment. [Figure 7] Figure 7 shows the pixel value difference image and attention difference image of the first image and the second image according to the first embodiment. [Modes for carrying out the invention]
[0010] The image processing apparatus and image processing method of the embodiment will be described below with reference to the drawings.
[0011] (First embodiment) Figure 1 shows an example of the configuration of an image processing apparatus according to the first embodiment.
[0012] The image processing device 1 comprises a communication unit 11, an input unit 12, a display unit 13, a storage unit 14, and an image segmentation unit 15. The communication unit 11, input unit 12, display unit 13, storage unit 14, and image segmentation unit 15 are communicated together, for example, via a bus (not shown).
[0013] The communication unit 11 includes, for example, a communication interface such as a NIC. The communication unit 11 can communicate with external equipment via a network and send and receive information. The communication unit 11 can output the received information to the image segmentation unit 15. The communication unit 11 can transmit information from the image segmentation unit 15 to an external device connected via the network.
[0014] The input unit 12 receives input operations from users such as doctors and specialists, and outputs signals based on the received input operations to the image segmentation unit 15. For example, the input unit 12 can be implemented by a mouse and keyboard, trackball, switch, button, joystick, touch panel, etc. Alternatively, the input unit 12 can be implemented by a user interface that accepts audio input such as a microphone. If the input unit 12 is a touch panel, the display unit 13, described later, can be formed integrally with the input unit 12.
[0015] The display unit 13 displays various information. For example, the display unit 13 displays an image generated by the image segmentation unit 15 and displays a GUI (Graphical User Interface) for accepting user input. For example, the display unit 13 is an LCD (Liquid Crystal Display) or an organic EL (Electroluminescence) display.
[0016] The memory unit 14 is implemented using storage devices such as ROM (Read Only Memory), flash memory, RAM (Random Access Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), and registers. Flash memory, HDD, SSD, etc., are non-transient storage media. These non-transient storage media are implemented using other storage devices connected via a network, such as NAS (Network Attached Storage) and external storage servers. The memory unit 14 stores the data necessary to construct the neural network, which will be described later.
[0017] The image segmentation unit 15 performs image segmentation on the current scan image of the patient's target lesion based on scan images of the patient's target lesion in the immediate preceding stage and the current stage. In this embodiment, image segmentation means classifying each pixel in the scan image and generating an image segmentation result that includes classification information for each pixel. The classification categories include at least three categories: "tumor," "organ," and "background." Hereinafter, the scan image of the patient's target lesion in the immediate preceding stage will be referred to as the first image, and the scan image of the patient's target lesion in the current stage will be referred to as the second image. The image segmentation unit 15 includes an attention map generation unit 151, a difference value calculation unit 152, a first acquisition unit 153, and a post-processing unit 154.
[0018] Each component of the image segmentation unit 15 is realized by a hardware processor such as a CPU and a GPU executing a program (software) stored in the storage unit 14. Some or all of these multiple components may be realized by hardware such as LSI, ASIC, and FPGA, or may be realized by the cooperative operation of software and hardware. The above-mentioned program may be stored in the storage unit 14 in advance, or may be stored in a removable storage medium such as a DVD and a CD-ROM, and installed from the storage medium to the storage unit 14 by attaching the storage medium to the drive device of the image processing apparatus 1.
[0019] The attention map generation unit 151 acquires the first image and the second image from an external device or the storage unit 14 via the communication unit 11, generates an image segmentation result and an attention map regarding the first image based on the first image, and generates an image segmentation result and an attention map regarding the second image based on the second image.
[0020] The attention map generation unit 151 includes a second acquisition unit 1511 and a mapping unit 1512. The second acquisition unit 1511 constructs a second neural network by reading the trained structural parameters of the second neural network stored in the storage unit 14, and performs image segmentation on the input scanned image based on the second neural network. The mapping unit 1512 calculates attention information of each pixel of the scanned image to a specific category through the second neural network based on the scanned image input to the second acquisition unit 1511 and the image segmentation result thereof, and generates an attention map. The attention information of each pixel to a specific category indicates the degree of interest of the neural network in each pixel when determining which pixel should be classified into the specific category in image segmentation. Generally, the higher the value of the attention information of a pixel, the more the neural network is interested in the pixel when classifying the specific category, and the higher the possibility that the pixel is classified into the specific category.
[0021] The difference value calculation unit 152 calculates the difference value between the attention map of the first image and the attention map of the second image generated by the attention map generation unit 151, generates an attention difference image, and outputs the attention difference image to the first acquisition unit 153.
[0022] The first acquisition unit 153 constructs a first neural network by reading the trained structural parameters of the first neural network stored in the storage unit 14, and performs image segmentation based on the input scanned image and the attention difference image calculated by the difference value calculation unit 152 by the first neural network, and outputs the generated image segmentation result to the post-processing unit 154.
[0023] The post-processing unit 154 processes the image segmentation results output from the first acquisition unit 153 and determines whether or not a new tumor exists based on the post-processed image segmentation results.
[0024] Figure 2 is a flowchart of the image processing method according to the first embodiment. The image segmentation method according to this embodiment will be described below with reference to Figure 2. In this embodiment, the first image is a CT scan image of the patient's liver lesion area immediately preceding the lesion, and the second image is a CT scan image of the patient's liver lesion area at the current stage, but the embodiment is not limited thereto. For example, the lesion area may be an organ such as the lungs, stomach, or brain, and the scan image may be acquired by an MRI scan, ultrasound scan, or the like. In this embodiment, both the first and second images are two-dimensional grayscale images that show information of one cross-section of the patient's lesion area, and this information is stored in a matrix of, for example, width W × height H.
[0025] Furthermore, in this embodiment, for the sake of explanation, we will describe a situation where image segmentation is performed using only one set of the first and second images. However, it is also possible to use multiple sets of the first and second images simultaneously and perform the image segmentation process in parallel.
[0026] When the input unit 12 receives a command from the user to start image segmentation processing, the following steps S101 to S107 are executed.
[0027] In step S101, the display unit 13 informs the user that input of information for the first and second images is required. Based on the user's input, the first and second images are acquired and output to the attention map generation unit 151. For example, a dialog box with the text "Please input the first and second images" is displayed on the display unit 13. Based on the user's operation, the first and second images are read from external equipment or the storage unit 14 and output to the attention map generation unit 151.
[0028] Figure 3 shows examples of the first and second images according to the first embodiment. Figure 3(a) shows an example of the first image, and Figure 3(b) shows an example of the second image. In this embodiment, the first and second images are pre-processed images. The pre-processing includes registration of the first and second images, normalization of image pixel values, image resampling, and image noise reduction.
[0029] As can be seen in Figure 3, in this case, the patient already had a tumor (shadow area) in liver region A at the stage immediately preceding Figure 3(a), and at the current stage shown in Figure 3(b), a new tumor (shadow area) appeared in liver region B. In Figures 3(a) and 3(b), the large gray area represents the actual liver tissue, and the large black area at the top represents the background.
[0030] In step 102, the attention map generation unit 151 performs image segmentation on the first and second images input in step S101, classifying each pixel in the first and second images into one of "tumor," "liver parenchyma," or "background." First, in the attention map generation unit 151, the second acquisition unit 1511 reads the parameters of the trained second neural network stored in the memory unit 14 and constructs the second neural network. The second neural network here is a convolutional neural network that performs image segmentation, such as U-net or attention U-net. Subsequently, the second acquisition unit 1511 inputs the first and second images into the second neural network, calculates the image segmentation results of the first and second images by forward propagation, and records the intermediate data generated in each layer of the neural network in order to calculate the image segmentation results. It should be noted that in step S102, image segmentation was performed on the first and second images via the second neural network. However, this image segmentation is performed to generate attention maps for the first and second images, and the image segmentation result obtained for the second image here is not the final image segmentation result for the second image.
[0031] The second neural network and its training process in this embodiment are described below. The second neural network in this embodiment is a multilayer convolutional neural network that takes data representing a scanned image as input and outputs image segmentation results for the scanned image through calculations in each layer. In image segmentation by a convolutional neural network, each pixel is classified based on all image information within a certain range centered on the pixel itself. To train the second neural network, it is necessary to prepare a large number of scanned images and ground truth data for their image segmentation. Ground truth data for image segmentation is data that contains accurate classification information for each pixel in the image. Subsequently, multiple sets of scanned images are input to the second neural network for each training session, the distance between the output image segmentation results and the ground truth data is calculated, and the parameters of the neural network are corrected using methods such as gradient descent and stochastic gradient descent so that this distance is shortened. The training is repeated until the distance between the image segmentation results output from the second neural network and the ground truth data satisfies a predetermined accuracy. Here, the aforementioned distance is, for example, the Manhattan distance or the Euclidean distance.
[0032] In step S103, the mapping unit 1512 in the attention map generation unit 151 uses the image segmentation results and intermediate data of the first and second images calculated in step S102, along with an attention map generation method such as CAM (class activation mapping) or grad-CAM (grad-class activation mapping), to perform backpropagation using a second neural network. This calculates attention information for each pixel of the first and second images corresponding to the category "liver parenchyma" in image segmentation, and generates a first attention map, which is an attention map for the first image, and a second attention map, which is an attention map for the second image.
[0033] Figure 4 shows examples of a first attention map and a second attention map according to the first embodiment. The attention map is an image with the same dimensions as the scanned image, and each pixel in the attention map represents the attention information of the pixel at the same position in the scanned image corresponding to the attention map. Figure 4(a) shows the first attention map, and Figure 4(b) shows the second attention map. In Figure 4, the higher the brightness of a pixel (the closer to white), the higher the value of the attention information for that pixel. That is, the neural network shows a high degree of interest in that pixel when classifying "liver parenchyma," the higher the weight assigned to that pixel when determining whether that pixel and surrounding pixels belong to "liver parenchyma," and the higher the probability that that pixel belongs to "liver parenchyma." On the other hand, the lower the brightness of a pixel (closer to black), the lower the attention information value for that pixel. This means that when the neural network classifies the pixel as "liver parenchyma," it has less interest in that pixel, and when it determines whether that pixel and surrounding pixels belong to "liver parenchyma," it assigns a lower weight to that pixel, indicating a lower probability that the pixel belongs to "liver parenchyma." As can be seen from Figure 4, in Figure 4(b), the neural network's interest in region B where a new tumor appears decreases, and as a result, the neural network recognizes that the possibility of that part becoming "liver parenchyma" is lower than in the previous stage.
[0034] In step S104, the difference value calculation unit 152 generates an attention difference image based on the first attention map and the second attention map generated by the attention map generation unit 151 in step S103. Specifically, the difference value calculation unit 152 first normalizes the first attention map and the second attention map generated in step S103, extracts the region of interest of the liver based on the normalized first attention map and the second attention map, performs weighting on the first attention region and the second attention region based on the region of interest, calculates the distance between each pixel of the weighted first attention map and the second attention map, and generates an attention difference image. Here, the distance is, for example, the Manhattan distance or the Euclidean distance.
[0035] Figure 5 shows an example of an attention difference image according to the first embodiment. As can be seen from Figure 5, there is a clear difference between the first attention map and the second attention map in region B where a new tumor has developed.
[0036] In step S105, the first acquisition unit 153 performs image segmentation on the second image based on the attention difference image and the second image. First, the first acquisition unit 153 reads the parameters of the trained first neural network stored in the memory unit 14 and constructs the first neural network. The first neural network here is a convolutional neural network that performs image segmentation, such as U-net or attention U-net. Subsequently, the first acquisition unit 153 inputs the second image and the attention difference image representing the difference between the first attention map and the second attention map into the first neural network and performs image segmentation on the second image based on the information of the second image and the attention difference image.
[0037] The first neural network of this embodiment and its training process will be described below. The first neural network of this embodiment is a multi-layer convolutional neural network that takes a scanned image and an attention difference image as input and outputs an image segmentation result for the scanned image through calculations in each layer. By using an attention difference image that represents the difference between the attention map of the current scanned image and the attention map of the previously scanned image, the first neural network can grasp which areas have undergone significant changes compared to the previous stage. For example, in this example, the neural network has grasped that a clear change has appeared in area B. In order to train the first neural network, it is necessary to prepare a large number of scanned images, attention difference images corresponding to the scanned images, and ground truth data for image segmentation of the scanned images. Ground truth data for image segmentation is data that contains accurate classification information for each pixel in the image. Subsequently, in each training session, multiple sets of scanned images and their corresponding attention difference images are input to the first neural network. The distance between the output image segmentation result and the ground truth data is calculated, and the neural network parameters are corrected using methods such as gradient descent or stochastic gradient descent to reduce this distance. This training is repeated until the distance between the image segmentation result output by the first neural network and the ground truth data satisfies a predetermined accuracy. Here, the distance is, for example, the Manhattan distance or the Euclidean distance.
[0038] Figure 6 shows an example of the final result of image segmentation of the second image according to the first embodiment. In Figure 6, accurate image segmentation was performed on the newly appeared tumor in region B, and the contour of the tumor was accurately segmented.
[0039] Figure 7 shows the pixel value difference image and attention difference image of the first and second images according to the first embodiment. Figure 7(a) is a pixel value difference image showing the difference in pixel values of each pixel between the first and second images of this embodiment. In the pixel value difference image, the higher the brightness (closer to white), the smaller the difference in pixel values between the first and second images for that pixel, and the lower the brightness (closer to black), the larger the difference in pixel values between the first and second images for that pixel. Figure 7(b) is the same image as the attention difference image in Figure 5. As can be seen by comparing Figure 7(a) and Figure 7(b), the pixel value difference image shows a large difference in the area where a new tumor has developed due to the influence of noise, but also shows a large difference in the area unrelated to the newly developed tumor. On the other hand, the attention difference image shows a large difference only in the area where a new tumor has developed. In other words, the attention difference image contains important information for detecting the development of a new tumor, but contains almost no noise information. Pixel value difference images contain information for detecting the development of new tumors, but they also contain a large amount of noise. If image segmentation is performed using pixel value difference images instead of attention difference images, there is a risk that areas that are not tumors may be classified as tumor areas due to the influence of noise.
[0040] In step S106, the post-processing unit 154 performs post-processing on the image segmentation result of the second image acquired in step S105, such as smoothing and correction of error segmentation areas, and the post-processed image segmentation result is used as the final image segmentation result of the second image.
[0041] In step S107, the post-processing unit 154 determines the tumor regions already present in the first image based on the ground truth data of the image segmentation of the first image, and determines whether or not newly appearing tumor regions exist based on the final result of the image segmentation of the second image. If newly appearing tumor regions exist, it performs processing such as tumor texture analysis and tumor dimension evaluation. Here, for example, the image segmentation result of the first image calculated in step S102 may be used as the ground truth data for the image segmentation of the first image, or the image segmentation result of the first image obtained by this image processing method or other image segmentation methods may be used as the ground truth data for the image segmentation of the first image.
[0042] After image segmentation is complete, the image segmentation unit 15 can display the final result of the image segmentation of the second image on the display unit 13, store it in the storage unit 14, or transmit it to an external device via the communication unit 11.
[0043] In this embodiment, when performing image segmentation on the current scan image of a patient's target lesion, accurate image segmentation can be achieved not only based on the current manipulation image, but also on the attention difference image, which is the difference in the attention map between the current scan image and the immediately preceding scan image. Furthermore, since the image processing method of the present invention can acquire stepwise changes in lesions such as tumors with high accuracy, it can be better applied to the evaluation of tumor treatment effects, such as RECIST.
[0044] (Second embodiment) The second embodiment will now be described. The second embodiment differs from the first embodiment described above in that the first image and the second image are scan images of the same stage of the target lesion but from different modalities. The following description will focus on the differences from the first embodiment, and the points common to the first embodiment will not be explained. In the description of the second embodiment, the same parts as in the first embodiment will be denoted by the same reference numerals.
[0045] In this embodiment, the first and second images are scan images from different modalities, meaning they are generated by different means. For example, the first image is a CT image of the patient's target lesion in its current state, and the second image is an MRI image of the patient's target lesion in its current state. Alternatively, for example, both the first and second images are CT images of the patient's target lesion in its current state, but the contrast parameters or imaging parameters used to generate the first and second images are different.
[0046] If the first and second images are scanned using different scanning methods, the second neural network is a multi-layer convolutional neural network, comprising an input layer corresponding to the first image and an input layer corresponding to the second image. By inputting the scanned image data into the corresponding input layers, image segmentation is performed on the scanned image. Alternatively, two second neural networks may be prepared, and segmentation may be performed on the first and second images respectively.
[0047] In this embodiment, when performing image segmentation on current scan images of a patient's target lesion, image segmentation is performed based on attention difference images, which are the differences in attention maps of scan images from different modalities. This allows for comparison of differences between scan images from different modalities, enabling accurate image segmentation.
[0048] (Third embodiment) The third embodiment will now be described. The third embodiment differs from the first embodiment described above in that the first and second images are scan images of the immediate preceding stage and the current stage of the target lesion using different modalities. The following description will focus on the differences from the first embodiment, and the points common to the first embodiment will not be explained. In the description of the third embodiment, the same parts as in the first embodiment will be denoted by the same reference numerals.
[0049] In this embodiment, the first and second images are scan images from different modalities, meaning they are generated by different means. For example, the first image is a CT image of the step immediately preceding the patient's target lesion, and the second image is an MRI image of the current step of the patient's target lesion. Alternatively, for example, the first image is a CT image of the step immediately preceding the patient's target lesion, and the second image is also a CT image of the current step of the patient's target lesion, but the contrast parameters or imaging parameters used to generate the first and second images are different.
[0050] If the first and second images are scanned using different scanning methods, the second neural network is a multi-layer convolutional neural network, comprising an input layer corresponding to the first image and an input layer corresponding to the second image. By inputting the scanned image data into the corresponding input layers, image segmentation is performed on the scanned image. Alternatively, two second neural networks may be prepared, and segmentation may be performed on the first and second images respectively.
[0051] In this embodiment, when performing image segmentation on a patient's target lesion at its current stage, the image segmentation is performed based on an attention difference image, which is the difference in attention maps between scan images of different modalities at different stages. This allows for comparison of differences between scan images of different modalities, enabling accurate image segmentation. Furthermore, since the image processing method of the present invention can acquire stepwise changes in lesions such as tumors with high accuracy, it can be better applied to the evaluation of tumor treatment efficacy, such as RECIST.
[0052] (modified version) In the embodiments described above, the case in which the first image and the second image are scanned images generated by the same scanning technique has been explained, but the first image and the second image may be obtained from different scanning techniques. In this case, two different neural networks can be provided in the second acquisition unit, and image segmentation can be performed on the first image and the second image, respectively.
[0053] Furthermore, although the above embodiments describe a situation where the first and second images have already been pre-processed, the first and second images may be images before pre-processing. In this case, after inputting the first and second images before pre-processing, the image processing apparatus according to this embodiment performs pre-processing on the first and second images.
[0054] Furthermore, in the embodiments described above, the first neural network was described as receiving a scan image and an attention difference image as input. However, the first neural network can also receive, in addition to these, the scan image from the previous stage, the image segmentation result of the scan image from the previous stage output by the second neural network, and the image segmentation result of the scan image from the current stage output by the second neural network.
[0055] Furthermore, although the above embodiments describe a situation where the scanned image is a two-dimensional image, the scanned image may also be a three-dimensional image containing information on multiple cross-sections of the patient's lesion area. Additionally, the scanned image may be an RGB image.
[0056] According to at least one embodiment described above, it is possible to automatically and accurately detect whether or not a new tumor has appeared in the target lesion of a patient.
[0057] While several embodiments have been described, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These embodiments can be carried out in a variety of other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims and their equivalents. [Explanation of Symbols]
[0058] 1 Image processing device 15 Image Segmentation Section 151 Attention Map Generation Unit 152 Difference Calculation Unit 153 First acquisition part 154 Post-processing 1511 2nd Acquisition Department 1512 Mapping section
Claims
1. An attention map generation unit generates, based on a first image and a second image acquired with different modalities for a patient's target lesion, a first attention map which is the same size as the first image and shows the degree of attention each pixel gives to the pixel at the same position in the first image when performing image segmentation, and a second attention map which is the same size as the second image and shows the degree of attention each pixel gives to the pixel at the same position in the second image when performing image segmentation. A difference value calculation unit that generates an attention difference image showing the difference between the first attention map and the second attention map, A first acquisition unit performs image segmentation on the second image based at least on the first image and the attention difference image, and obtains the segmentation result of the second image. An image processing device equipped with the following features.
2. The attention map generation unit includes a second acquisition unit and a mapping unit, The second acquisition unit performs image segmentation on the first image and the second image, and acquires the segmentation result of the first image, the first intermediate data generated in each layer of the neural network when calculating the segmentation result of the first image, the segmentation result of the second image, and the second intermediate data generated in each layer of the neural network when calculating the segmentation result of the second image. The image processing apparatus according to claim 1, wherein the mapping unit generates the first attention map using the segmentation result of the first image and the first intermediate data by CAM (class activation mapping) or grad-CAM based on the neural network, and generates the second attention map using CAM or grad-CAM based on the segmentation result of the second image and the second intermediate data.
3. The image processing apparatus according to claim 1 or 2, wherein the difference value calculation unit performs pixel value processing on the first attention map and the second attention map, extracts regions of interest, weights the first attention map and the second attention map based on the regions of interest, and calculates the distance between corresponding pixels in the weighted first attention map and the second attention map to generate an attention difference image.
4. The image processing apparatus according to claim 1 or 2, wherein the first image is a scan image of the step immediately preceding the target lesion of the patient, and the second image is a scan image of the current step of the target lesion of the patient.
5. The image processing apparatus according to claim 1 or 2, further comprising a post-processing unit that performs post-processing on the segmentation result of the second image, including smoothing and correction of error segmentation regions.
6. The image processing apparatus according to claim 2, wherein the first acquisition unit further performs image segmentation in the second image based on the image segmentation result of the second image acquired by the second acquisition unit.
7. The image processing apparatus according to claim 2, wherein the first acquisition unit further performs image segmentation of the second image based on the first image.
8. Based on first and second images acquired with different modalities for the target lesion of the patient, a first attention map is generated, which is the same size as the first image and shows the degree of attention each pixel gives to the pixel at the same position in the first image when performing image segmentation, and a second attention map is generated, which is the same size as the second image and shows the degree of attention each pixel gives to the pixel at the same position in the second image when performing image segmentation. An attention difference image is generated that shows the difference between the first attention map and the second attention map. Perform segmentation of the second image based at least on the first image and the attention difference image, and obtain the segmentation result of the second image. An image processing method that includes the following.