Information processing device, information processing method, and program

The information processing device uses two-dimensional tomographic images to generate three-dimensional likelihood maps for CT images, addressing the inefficiency of manual correction in creating ground truth data and enhancing machine learning model training.

JP7851109B2Active Publication Date: 2026-04-24CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CANON KK
Filing Date
2021-11-26
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Creating three-dimensional ground truth image data for tomographic images, such as CT images, requires significant manual correction efforts, making the process laborious and inefficient.

Method used

An information processing device that acquires region information from two-dimensional tomographic images orthogonal to different axes and generates a three-dimensional likelihood map using a multivariate Gaussian function to efficiently create ground truth image data, which is then used to train a 3D U-Net model for segmentation.

Benefits of technology

Efficiently generates three-dimensional ground truth image data, reducing manual correction efforts and enabling accurate training of machine learning models for image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007851109000009
    Figure 0007851109000009
  • Figure 0007851109000010
    Figure 0007851109000010
  • Figure 0007851109000011
    Figure 0007851109000011
Patent Text Reader

Abstract

To efficiently generate three-dimensional correct answer image data indicating an area of an object.SOLUTION: An information processing apparatus 100 of the present invention comprises: a first acquisition unit 110 that acquires first area information on an object on first two-dimensional tomographic image data orthogonal to a first axis of three-dimensional tomographic image data; a second acquisition unit 120 that acquires second area information on the object on second two-dimensional tomographic image data orthogonal to a second axis different from the first axis of the three-dimensional tomographic image data; and a generation unit 130 that generates a three-dimensional likelihood map indicating the likelihood of the object based on the first area information and the second area information as three-dimensional correct answer image data corresponding to the three-dimensional tomographic image data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This specification discloses an information processing device, an information processing method, and a program for creating three-dimensional ground truth image data. [Background technology]

[0002] The accuracy of machine learning-based segmentation depends on the amount of training data, which consists of training image data and ground truth image data; therefore, it is desirable to prepare a large amount of training data. Since creating ground truth image data is a laborious task, techniques for efficiently creating ground truth image data are important. For example, as disclosed in Non-Patent Document 1, semi-automatic segmentation techniques that require user input of region information can be used to create ground truth image data by providing foreground information (information about the region of the object) and background information (information about the region other than the object). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Y. Boykov et al., Graph Cuts and Efficient ND Image Segmentation, int. j. comput. vision, 70(2), 2006. [Overview of the Initiative] [Problems that the invention aims to solve]

[0004] However, when creating ground truth image data using semi-automatic segmentation technology, foreground and background information must be repeatedly corrected manually. Creating three-dimensional ground truth image data for three-dimensional tomographic image data such as CT image data requires a significant amount of repeated correction effort.

[0005] The present invention aims to provide an information processing device that can efficiently generate three-dimensional ground truth image data indicating the region of an object. [Means for solving the problem]

[0006] The information processing device according to the present invention includes: a first acquisition unit that acquires first region information of an object on a first two-dimensional tomographic image data orthogonal to a first axis of three-dimensional tomographic image data; a second acquisition unit that acquires second region information of an object on a second two-dimensional tomographic image data orthogonal to a second axis different from the first axis of three-dimensional tomographic image data; and, based on the first region information and the second region information, determines the object's resemblance as three-dimensional ground truth image data corresponding to the three-dimensional tomographic image data. three A generator that generates a dimensional likelihood map and ,of Preparation The generation unit determines a reference point based on at least one of the first region information and the second region information, and generates the likelihood map such that the likelihood is high for voxels close to the reference point and low for voxels far from the reference point. . [Effects of the Invention]

[0007] According to the disclosures herein, it is possible to efficiently generate three-dimensional ground truth image data that indicates the region of an object. [Brief explanation of the drawing]

[0008] [Figure 1] A diagram showing an example of the functional configuration of an information processing device according to the first embodiment. [Figure 2] A diagram showing an example of the hardware configuration of an information processing device according to the first embodiment. [Figure 3] A diagram showing an example of the processing procedure of an information processing device according to the first embodiment. [Figure 4] A diagram illustrating image data according to the first to third embodiments. [Figure 5] A diagram showing an example of the processing procedure of an information processing device according to the first embodiment. [Modes for carrying out the invention]

[0009] Preferred embodiments of the information processing apparatus disclosed herein will be described below with reference to the drawings. Identical or equivalent components, members, and processes shown in each drawing are denoted by the same reference numerals, and redundant descriptions are omitted as appropriate. Furthermore, some components, members, and processes are omitted from the drawings as appropriate.

[0010] The present invention will be described below using a liver tumor depicted in abdominal CT image data acquired by an X-ray computed tomography (X-ray CT) scanner as an example. Although abdominal CT image data is an example of three-dimensional tomographic image data, the present invention is also applicable to three-dimensional tomographic image data such as tomographic image data acquired by a magnetic resonance imaging (MRI) scanner, a positron emission tomography (PET) scanner, and an ultrasound scanner. Furthermore, the present invention is applicable not only to liver tumors but also to other lesions (e.g., pulmonary nodules, lymph nodes, bone metastases, etc.) and any other structures. Moreover, the embodiments of the present invention are not limited to the embodiments described below. [Examples]

[0011] <First Embodiment> In this embodiment, the target object is a liver tumor, the liver tumor is approximated as an ellipsoid, and a method for generating a three-dimensional likelihood map (hereinafter referred to as the liver tumor likelihood map) that indicates the likelihood of being a liver tumor as three-dimensional ground truth image data is described. In this embodiment, the first axis is, for example, an axial axis, and the first two-dimensional tomographic image data orthogonal to the axial axis is the axial cross-sectional image data of abdominal CT image data. This axial cross-sectional image data is, for example, an axial tomographic image data that best represents the shape of the liver tumor among a plurality of axial cross-sectional image data constituting the abdominal CT image (for example, the axial cross-sectional image data in which the area of ​​the liver tumor is maximized) (hereinafter referred to as the representative axial cross-sectional image data). The first region information is a two-dimensional ground truth image data (hereinafter referred to as the liver tumor axial ground truth image data) that shows the region of the liver tumor depicted in the representative axial cross-sectional image data. Furthermore, the second axis, which is different from the first axis, the axial axis, is, for example, the coronal axis, and the second two-dimensional tomographic image data orthogonal to the coronal axis is the coronal cross-sectional image data of the abdominal CT image data. This coronal cross-sectional image data is, for example, a coronal tomographic image data that best represents the shape of the liver tumor among multiple coronal cross-sectional image data constituting the abdominal CT image (for example, the coronal cross-sectional image data in which the area of ​​the liver tumor is largest) (hereinafter referred to as representative coronal cross-sectional image data). The second region information is a two-dimensional ground truth image data that shows the region of the liver tumor depicted in the representative coronal cross-sectional image data (hereinafter referred to as the coronal ground truth image data of the liver tumor).

[0012] The information processing apparatus 100 according to this embodiment generates a three-dimensional likelihood map of the correct liver tumor so as to spread in an ellipsoidal shape (concentric circles and continuously changing likelihood). More specifically, the information processing apparatus 100 first determines a reference point based on the axial correct image data and the coronal correct image data of the liver tumor. Then, centering on the reference point, a three-dimensional likelihood map of the liver tumor is generated so as to spread in an ellipsoidal shape (concentric circles and continuously changing likelihood). At this time, the likelihood of each voxel included in the three-dimensional likelihood map of the liver tumor is set by a multivariate Gaussian function which is a probability distribution function. That is, the information processing apparatus 100 according to this embodiment generates a three-dimensional likelihood map of the correct liver tumor so as to conform to a multivariate Gaussian distribution centering on the reference point.

[0013] Furthermore, the information processing apparatus 100 according to this embodiment uses the three-dimensional likelihood map of the correct liver tumor as three-dimensional correct image data and trains a learning model by a method based on machine learning. That is, the information processing apparatus 100 trains a learning model using three-dimensional tomographic image data which is three-dimensional abdominal CT image data and the three-dimensional likelihood map of the correct liver tumor corresponding to the abdominal CT image data as teacher data. In this embodiment, among deep learning techniques, 3D U-Net, which is one of Convolutional Neural Networks (CNN), is used as a learning model. By training 3D U-Net using the above-mentioned teacher data, it becomes possible to infer a region likely to be a liver tumor.

[0014] Hereinafter, the axial axis, the coronal axis, and the sagittal axis shall be axes corresponding to the Z axis, the Y axis, and the X axis, respectively, in the image data coordinate system. That is, the representative axial cross-sectional image data and the axial correct image data of the liver tumor are image data on the XY plane, and the representative coronal cross-sectional image data and the coronal correct image data of the liver tumor are image data on the XZ plane.

[0015] Hereinafter, the functional configuration of the information processing apparatus 100 according to the present embodiment will be described with reference to FIG. 1. As shown in the figure, the information processing apparatus 100 includes a generation processing unit 101 for correct image data including a first acquisition unit 110, a second acquisition unit 120, and a generation unit 130, and a learning processing unit 102 including a teacher data acquisition unit 140 and a learning unit 150. Further, the information processing apparatus 100 according to the present embodiment includes a storage device 70. Note that the generation processing unit 101 for correct image data and the learning processing unit 102 may function as an information processing system configured from different devices.

[0016] The storage device 70 is an example of a computer-readable storage medium and is a large-capacity storage device typified by a hard disk drive (HDD) or a solid state drive (SSD). The storage device 70 holds abdominal CT image data, which is three-dimensional tomographic image data in which an object is depicted, axial correct image data of a liver tumor, which is first region information, and coronal correct image data of a liver tumor, which is second region information. The axial correct image data of the liver tumor and the coronal correct image data of the liver tumor are, for example, two-dimensional mask image data annotated by a doctor or a radiological technologist. In the present embodiment, the axial correct image data of the liver tumor and the coronal correct image data of the liver tumor are two-dimensional mask image data for representative two-dimensional tomographic image data (for example, two-dimensional tomographic image data having the largest area of the liver tumor) selected from two-dimensional tomographic image data orthogonal to each axis. The axial correct image data of the liver tumor and the coronal correct image data of the liver tumor are binary image data in which the value of voxels included in the region of the liver tumor is 1 and the value of other voxels is 0. The above-described expression format of the mask image data is an example, and any format that can represent the region of the liver tumor may be used. For example, the first region information and the second region information may be images representing the likelihood of a liver tumor for each voxel with multiple values. Note that the storage device 70 may be configured as a device different from the information processing apparatus 100 described later.

[0017] First, the configuration of the generation processing unit 101 for correct image data in the information processing apparatus 100 will be described.

[0018] The first acquisition unit 110 acquires axial ground truth image data of a liver tumor from the storage device 70 as information of a first region of the object on the first two-dimensional tomographic image data that is orthogonal to the first axis of the three-dimensional tomographic image data, and transmits it to the generation unit 130.

[0019] The second acquisition unit 120 acquires coronal ground truth image data of a liver tumor from the storage device 70 as information about the second region of the object on the second two-dimensional tomographic image data, which is orthogonal to the second axis, which is different from the first axis of the three-dimensional tomographic image data, and transmits it to the generation unit 130.

[0020] The generation unit 130 generates a three-dimensional likelihood map indicating the likelihood of an object as three-dimensional ground truth image data corresponding to three-dimensional tomographic image data, based on the first and second region information. Specifically, it receives axial ground truth image data of the liver tumor, which is the first region information, from the first acquisition unit 110, and coronal ground truth image data of the liver tumor, which is the second region information, from the second acquisition unit 120. Subsequently, the generation unit 130 determines reference points based on the axial ground truth image data and coronal ground truth image data of the liver tumor, and generates a ground truth three-dimensional likelihood map of the liver tumor based on a multivariate Gaussian function. Then, the generation unit 130 stores the ground truth three-dimensional likelihood map of the liver tumor in the storage device 70. The three-dimensional likelihood map of the correct liver tumor generated by the generation unit 130 has the same image data size as the abdominal CT image data, and is a continuous image data where voxels with a high likelihood of liver tumors are represented by values ​​close to 1, and voxels with a low likelihood of liver tumors are represented by values ​​close to 0. The above representation format of the three-dimensional likelihood map of the correct liver tumors is just one example; any format that can represent the high or low likelihood of liver tumors is acceptable.

[0021] In other words, the configuration of the ground truth image data generation processing unit 101 in the information processing device 100 includes: a first acquisition unit 110 that acquires first region information of an object on a first two-dimensional tomographic image data orthogonal to a first axis of three-dimensional tomographic image data; a second acquisition unit 120 that acquires second region information of an object on a second two-dimensional tomographic image data orthogonal to a second axis different from the first axis of three-dimensional tomographic image data; and a generation unit 130 that generates a three-dimensional likelihood map indicating the likelihood of the object being identified as three-dimensional ground truth image data corresponding to the three-dimensional tomographic image data, based on the first region information and the second region information.

[0022] Next, the configuration of the learning processing unit 102 in the information processing device 100 will be described.

[0023] The training data acquisition unit 140 receives multiple abdominal CT image data as three-dimensional training image data and multiple three-dimensional likelihood maps of liver tumors as three-dimensional ground truth image data corresponding to each of the multiple abdominal CT image data from the storage device 70, and transmits them to the learning unit 150.

[0024] The learning unit 150 receives multiple three-dimensional abdominal CT image data and multiple three-dimensional ground truth liver tumor likelihood maps corresponding to each of the abdominal CT image data from the training data acquisition unit 140. Next, it initializes the parameters of the learning model. Subsequently, the learning unit 150 uses the multiple abdominal CT image data and the multiple three-dimensional ground truth liver tumor likelihood maps corresponding to each of the abdominal CT image data as training data to train the learning model using a machine learning-based method. Then, the learning unit 150 saves the parameters of the trained learning model to the storage device 70. In this embodiment, the learning model is a 3D U-Net, which is a type of CNN. That is, the learning unit 150 considers the three-dimensional abdominal CT image data and the three-dimensional ground truth liver tumor likelihood maps corresponding to the abdominal CT image data as a set of training data and trains the 3D U-Net using multiple sets of training data.

[0025] At least a portion of each part of the information processing device 100 shown in Figure 1 may be implemented as an independent device. Alternatively, each function may be implemented as software. In this embodiment, each part is assumed to be implemented by software.

[0026] Figure 2 shows an example of the hardware configuration of the information processing device 100. The information processing device 100 has the configuration of a known computer (information processing device). The information processing device 100 includes, as its hardware configuration, a CPU 201, main memory 202, magnetic disk 203, display memory 204, monitor 205, mouse 206, and keyboard 207.

[0027] The CPU (Central Processing Unit) 201 primarily controls the operation of each component. The main memory 202 stores control programs executed by the CPU 201 and provides a workspace for program execution by the CPU 201. The magnetic disk 203 stores programs for implementing various application software, including the OS (Operating System), device drivers for peripheral devices, and programs for processing described later. By executing programs stored in the main memory 202, magnetic disk 203, etc., the CPU 201 realizes the functions (software) of the information processing device 100 shown in Figure 1 and the processing shown in the flowchart described later.

[0028] The display memory 204 temporarily stores display data. The monitor 205 is, for example, a CRT monitor or an LCD monitor, and displays image data, text, etc., based on the data from the display memory 204. The mouse 206 and keyboard 207 accept pointing input and character input, respectively, from the user. Each of the above components is connected to each other via a common bus 208 so that they can communicate with one another.

[0029] The CPU 201 corresponds to an example of a processor or control unit. In addition to the CPU 201, the information processing device 100 may have at least one of a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array). Alternatively, the CPU 201 may be replaced with at least one of a GPU or an FPGA. The main memory 202 and magnetic disk 203 correspond to an example of memory or storage device.

[0030] Next, the processing procedure of the information processing device 100 according to this embodiment will be explained using Figures 3 and 5. First, the processing procedure of the correct image data generation processing unit 101 in the information processing device 100 will be explained with reference to Figure 3.

[0031] (Step S3100) In step S3100, the first acquisition unit 110 acquires axial ground truth image data of a liver tumor from the storage device 70 as first region information of the object on the first two-dimensional tomographic image data that is orthogonal to the first axis of the three-dimensional tomographic image data, and transmits it to the generation unit 130.

[0032] (Step S3200) In step S3200, the second acquisition unit 120 acquires coronal ground truth image data of a liver tumor from the storage device 70 as second region information of the object on the second two-dimensional tomographic image data, which is orthogonal to the second axis, which is different from the first axis of the three-dimensional tomographic image data, and transmits it to the generation unit 130.

[0033] (Step S3300) In step S3300, the generation unit 130 generates a three-dimensional likelihood map indicating the likelihood of the object as three-dimensional ground truth image data corresponding to three-dimensional tomographic image data, based on the first region information and the second region information. Specifically, the generation unit 130 receives the axial ground truth image data of the liver tumor, which is the first region information, and the coronal ground truth image data of the liver tumor, which is the second region information. Then, based on the axial ground truth image data of the liver tumor and the coronal ground truth image data of the liver tumor, it determines the reference points for generating the three-dimensional ground truth likelihood map of the liver tumor.

[0034] In this embodiment, the generation unit 130 generates a three-dimensional likelihood map of the liver tumor so that it spreads out in an ellipsoidal shape (the likelihood changes concentrically and continuously) around a reference point. Therefore, it is preferable that the reference point is approximately at the center of the liver tumor. For this reason, the method for determining the reference point so that it is approximately at the center of the liver tumor will be explained below using Figure 4. The figure shows the first region information, which is the axial ground truth image data 410 of the liver tumor, and the second region information, which is the coronal ground truth image data 420 of the liver tumor. The axial ground truth image data 410 of the liver tumor holds the ground truth region 411 of the liver tumor (hereinafter referred to as the axial region 411 of the liver tumor) depicted in the representative axial cross-sectional image data corresponding to the axial ground truth image data 410 of the liver tumor. Similarly, the coronal ground truth image data 420 of the liver tumor holds the coronal region 421 of the liver tumor.

[0035] The axial ground truth image data 410 and the coronal ground truth image data 420 of the liver tumor are two-dimensional image data corresponding to different axes, and as shown in the figure, the two two-dimensional ground truth image data intersect in three-dimensional space. As described above, the two two-dimensional ground truth image data hold the region of the liver tumor corresponding to the cross-section with the largest area of ​​the liver tumor in each axis direction, so it is considered that the center point of the liver tumor is near the midpoint of the line segment where the axial region 411 and the coronal region 421 of the liver tumor intersect. Therefore, the generation unit 130 determines the midpoint of this line segment as the reference point 440.

[0036] The method for determining the reference point by the generation unit 130 is not limited to the example described above. For example, the centroid of the axial region 411 of the liver tumor or the centroid of the coronal region 421 of the liver tumor may be used as the reference point, or the midpoint of the two centroids may be used as the reference point. Alternatively, the generation unit 130 may obtain the bounding boxes of the axial region 411 and the coronal region 421 of the liver tumor and use the center coordinates of the bounding boxes as the reference point. In addition, any method of selecting the approximate center coordinates of the liver tumor, which is the target object, is acceptable.

[0037] (Step S3400) In step S3400, the generation unit 130 generates a likelihood map indicating the likelihood of an object such that the likelihood is high for voxels close to the reference point and low for voxels far from the reference point. Specifically, the generation unit 130 generates a three-dimensional likelihood map of the correct liver tumor, spreading out in an ellipsoidal shape (concentric and continuously changing likelihood) around the reference point determined in step S3300. In this embodiment, the generation unit 130 generates the likelihood map of the correct liver tumor using a multivariate Gaussian function, which is a probability distribution function. The generation unit 130 generates the likelihood p of the liver tumor for voxel i in the likelihood map of the correct liver tumor. i This is calculated using the following formula.

[0038]

number

[0039] Here,

[0040]

number

[0041] This is a vector that indicates the coordinates of voxel i in the image data coordinate system.

[0042]

number

[0043] is a vector representing the coordinates of reference point 440 in the image data coordinate system. Furthermore, Σ is the covariance matrix, which functions as a parameter determining the extent of the spread of the likelihood map of the correct liver tumors. In particular, the diagonal elements of the covariance matrix represent the variance in the X-axis (sagittal axis), Y-axis (cornal axis), and Z-axis (axial axis) directions.

[0044]

number

[0045] These determine the degree of spread of the likelihood distribution in each axial direction. That is, the likelihood of the liver tumor in voxel i is determined by the variance in each axial direction. In this embodiment, as an example, in order to make the likelihood of the liver tumor near the contour of the liver tumor 0.5, the variance in the X-axis direction is

[0046]

number

[0047] , the dispersion in the Y-axis direction

[0048]

number

[0049] , the dispersion in the Z-axis direction

[0050]

number

[0051] Set it to r. x ,r y ,r zThese are the radii in the X, Y, and Z directions of the liver tumor, respectively. These radii can be obtained from the first region information, which is the axial ground truth image data of the liver tumor, and the second region information, which is the coronal ground truth image data of the liver tumor. For example, since the axial ground truth image data 410 of the liver tumor is image data on the XY plane, the radius in the X direction and the radius in the Y direction can be obtained from the axial region 411 of the liver tumor. Similarly, since the coronal ground truth image data of the liver tumor is image data on the XZ plane, the radius in the X direction and the radius in the Z direction can be obtained from the coronal region 421 of the liver tumor. Therefore, the generation unit 130 obtains the radius in the X direction and the radius in the Y direction from the axial region 411 of the liver tumor, and the radius in the Z direction from the coronal region 421 of the liver tumor. Using these radii, the generation unit 130 determines the variance in each direction and generates a three-dimensional likelihood map of the liver tumor according to equation 1. Figure 4(b) shows the three-dimensional likelihood map 430 of the correct liver tumor generated by the generation unit 130 in this embodiment, where lighter colors represent voxels with a high likelihood of liver tumor and darker colors represent voxels with a low likelihood of liver tumor. In the three-dimensional likelihood map 430 of the correct liver tumor, the likelihood changes as you move outward from the reference point 440. In other words, the likelihood of the three-dimensional likelihood map 430 of the correct liver tumor generated by the generation unit 130 changes in an ellipsoidal shape around the reference point 440.

[0052] The method used by the generation unit 130 to obtain the radius of the liver tumor can be any method that obtains the radius corresponding to any three axes. For example, the generation unit 130 may obtain the bounding boxes of the axial region 411 and the coronal region 421 of the liver tumor and calculate the radius of each axis based on the length of each side of the bounding boxes. Alternatively, the generation unit 130 may use the distance from the reference point to the voxel that constitutes the contour of the axial region 411 and the coronal region 421 of the liver tumor as the radius in the first axis direction, and obtain the remaining two radii from the two axes orthogonal to the first axis direction. In this case, since the three axes do not coincide with the X, Y, and Z axes, the likelihood map of the liver tumor can be generated in the same way as in the example above, and then the likelihood map of the liver tumor can be rotated in space so that the three axes become the basis. Alternatively, by setting values ​​for elements other than the diagonal components of the covariance matrix, the generation unit 130 can generate the likelihood map of the liver tumor so that the three axes become the basis.

[0053] Furthermore, if there is an inconsistency between the axial region 411 and the coronal region 421 of the liver tumor on the line where they intersect, the generation unit 130 may obtain the radius in the X-axis direction using the average, maximum, or minimum value of the radii obtained from each region. An inconsistency between the axial region 411 and the coronal region 421 of the liver tumor refers, for example, to a case where, when focusing on a certain voxel, that voxel is included in the liver tumor region in one two-dimensional ground truth image data, but not in the liver tumor region in the other.

[0054] After completing the above process, step S3400 is terminated.

[0055] (Step S3500) In step S3500, the generation unit 130 outputs a three-dimensional likelihood map showing the likelihood of a three-dimensional correct liver tumor region and stores it in the storage device 70.

[0056] (Step S3600) In step S3600, the information processing device 100 determines whether or not there is three-dimensional tomographic image data (three-dimensional abdominal CT image data) of the target to be processed. If there is, it returns to step S3100 and performs the ground truth image generation process for the remaining training image data. On the other hand, if there is no training image data of the target to be processed, the three-dimensional ground truth image data generation process is terminated.

[0057] In accordance with the above procedure, the information processing device 100 according to this embodiment generates a likelihood map of the correct liver tumor using the three-dimensional correct image data generation processing unit 101.

[0058] Next, the processing procedure of the learning processing unit 102 in the information processing device 100 will be explained with reference to Figure 5. In the learning processing unit 102, the likelihood map of the correct liver tumors generated by the correct image data generation processing unit 101 is used to train a learning model using a machine learning-based method.

[0059] (Step S5100) In step S5100, the training data acquisition unit 140 acquires multiple three-dimensional abdominal CT image data and three-dimensional likelihood maps of ground truth liver tumors corresponding to each of the multiple abdominal CT image data as training data sets and transmits them to the learning unit 150. In other words, the training data acquisition unit 140 considers abdominal CT image data and the likelihood maps of ground truth liver tumors as a set of training data, acquires multiple training data sets, and transmits them to the learning unit 150.

[0060] (Step S5200) In step S5200, the learning unit 150 initializes the parameters of the learning model, 3D U-Net. More specifically, the learning unit 150 initializes the kernel weights of the convolutional layer, for example, using a known method. In this embodiment, the convolutional layer is initialized using the method for determining initial values ​​based on a normal distribution proposed by He [Kaiming He, et al., “Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification,” ICCV, 2015] et al. The initialization of the convolutional layer is not limited to this, and any known method such as constant values ​​or random values ​​may be used. Furthermore, if the 3D U-Net includes layers with parameters such as batch normalization layers and dropout layers, the parameters of each layer are initialized using a known method.

[0061] (Steps S5300~S5700) In steps S5300 to S5700, the learning unit 150 receives the training dataset from the training data acquisition unit 140 and trains the 3D U-Net, which is the learning model, using the training dataset. The following describes how to train the 3D U-Net.

[0062] In step S5300, the learning unit 150 selects multiple training data (batches) from the training dataset to update the parameters of the 3D U-Net. Each training data consists of a pair of training image data (abdominal CT image data) and a ground truth likelihood map of a liver tumor corresponding to the abdominal CT image data. In this step, one or more training data sets are selected.

[0063] In step S5400, the learning unit 150 inputs abdominal CT image data included in the multiple training data selected in step S5300 into the 3D U-Net and performs forward propagation processing. By performing forward propagation processing, the learning unit 150 obtains multiple estimated likelihood maps of liver tumors for each of the multiple input abdominal CT image data.

[0064] In step S5500, the learning unit 150 calculates a loss value based on the likelihood maps of multiple estimated liver tumors obtained in step S5300 and the corresponding likelihood maps of the correct liver tumors. The loss function used is, for example, Mean Squared Error (MSE). The loss function is not limited to MSE; other known methods such as Mean Absolute Error (MAE) or Huber may also be used.

[0065] In step S5600, the learning unit 150 calculates the gradient using backpropagation based on the loss value calculated in step S5500 and updates the parameters of the 3D U-Net. As an example of an optimization method, the Stochastic Gradient Descent (SGD) method is used. The optimization method is not limited to the SGD method; other known methods such as the Adam method or AdaGrad method may also be used.

[0066] In step S5700, the learning unit 150 determines whether the learning termination condition has been met and decides on the next step. As an example, the learning termination condition is reaching a predetermined maximum number of epochs. That is, if the predetermined maximum number of epochs has been reached, the process proceeds to step S5800; otherwise, it returns to step S5300. Other conditions may be used for the learning termination condition. For example, the condition may be that the difference between the loss value of the previous epoch and the loss value of the current epoch satisfies a specified condition.

[0067] The above describes a typical training method for a 3D U-Net to infer the likelihood map of liver tumors.

[0068] (Step S5800) In step S5800, the learning unit 150 outputs the parameters of the learning model and stores them in the storage device 70.

[0069] In accordance with the above procedure, the information processing device 100 according to this embodiment uses the likelihood map of the correct liver tumor generated by the correct image data generation processing device 101 to train a learning model for inferring the likelihood map of the liver tumor using a machine learning-based method in the learning processing unit 102.

[0070] As described above, the information processing device 100 according to this embodiment generates a likelihood map (ground truth image data) of the object based on two-dimensional ground truth image data (first region information) corresponding to the first axis of the object and two-dimensional ground truth image data (second region information) corresponding to the second axis of the object. This makes it possible to efficiently generate three-dimensional ground truth image data that can be used for training a machine learning-based learning model.

[0071] (modified version) In the example above, a likelihood map of the correct object was generated based on a probability distribution function so that it spreads out in an ellipsoidal shape (the likelihood changes concentrically and continuously). However, the method for generating an ellipsoidal likelihood map is not limited to this. For example, an ellipsoidal likelihood map may be generated using the Euclidean distance from a reference point. In this case, the generation unit 130 calculates the Euclidean distance from the reference point for each voxel and generates the likelihood map by normalizing the distance values ​​so that the Euclidean distance of each voxel falls between 0 and 1. At this time, for example, the distance values ​​are normalized in an arbitrary way so that the likelihood near the contour of the object becomes 0.5. Then, the normalized distance values ​​are used as the likelihood indicating the object's resemblance.

[0072] In the example above, the first axis was the axial axis (Z-axis) and the second axis was the coronal axis (Y-axis), but the first and second axes can be in different directions, or any combination of axes. For example, the sagittal axis (X-axis) can be used, or any axis different from the X, Y, and Z axes can be used. Also, the first and second axes do not have to be orthogonal.

[0073] In the example above, the second region information (or first region information) was the two-dimensional ground truth image data of the object corresponding to the second axis, but any information including at least position and size is acceptable. For example, as shown in Figure 4(c), if two contour points 460 of a liver tumor on coronal cross-sectional image data are used as the second region information, the intersection of the axial ground truth image data 410 of the liver tumor (the first region information) and the line segment connecting the two contour points 460 is set as the reference point 440. Then, the radius in the X-axis direction and the radius in the Y-axis direction are obtained from the axial ground truth image data 410 of the liver tumor, and the radius in the Z-axis direction is obtained from the line segment connecting the two contour points 460, and a likelihood map of the ground truth liver tumor is generated in the same way as in the example above. The same applies when using the center coordinates and radius of an ellipsoid approximating the liver tumor on two-dimensional tomographic image data orthogonal to an arbitrary axis as the second region information.

[0074] In the example described above, an example using two pieces of region information about the object was explained, but a likelihood map of the correct object may be generated using three or more pieces of region information. The generation unit 130 acquires, for example, two or more pieces of second region information orthogonal to the second axis, and generates a three-dimensional likelihood map based on these two or more pieces of second region information. That is, the generation unit 130 can generate a three-dimensional likelihood map with high accuracy if, for example, there are multiple second two-dimensional tomographic image data. Note that not only second two-dimensional tomographic image data, but also first two-dimensional tomographic image data, third two-dimensional tomographic image data, or multiple of each may be used.

[0075] Alternatively, the generation unit 130 may use region information corresponding to each of the three different axes. For example, one case is to use the first two-dimensional tomographic image data, which is the axial ground truth image data of the liver tumor (first region information corresponding to the first axis), the second two-dimensional tomographic image data, which is the coronal ground truth image data of the liver tumor (second region information corresponding to the second axis), and the third two-dimensional tomographic image data, which is the sagittal ground truth image data of the liver tumor (third region information corresponding to the third axis). In this case, by determining the reference point according to the same procedure as in the example above and determining the degree of spread of the likelihood distribution in each axis direction based on the radius in each axis direction, a likelihood map of the correct object can be generated. It is also possible to use region information corresponding to the first axis and two pieces of region information corresponding to the second axis. For example, one case is to use the axial ground truth image data of the liver tumor, the first coronal ground truth image data of the liver tumor, and the second coronal ground truth image data of the object. In this case, a likelihood map of liver tumors can be generated by determining the reference point and the degree of spread of the likelihood distribution in each axial direction, following the same procedure as in the example above.

[0076] In the example above, a deep learning-based method such as 3D U-Net was used as the learning model, but it is not limited to this. For example, a Support Vector Machine (SVM) or a classification tree could also be used as the learning model. In this case, the learning method should be an appropriate one depending on the learning model.

[0077] <Second Embodiment> The information processing device according to the first embodiment generated a likelihood map of the correct liver tumor using a probability distribution function. This example describes how the information processing device according to this embodiment generates mask image data of an ellipsoid that shows the approximate region of the liver tumor using the equation of an ellipsoid, and generates a three-dimensional likelihood map of the correct liver tumor from this mask image data.

[0078] The configuration of the information processing apparatus according to this embodiment is the same as that of the information processing apparatus 100 according to the first embodiment. Hereinafter, with reference to Figure 1, the functional configuration of the information processing apparatus according to this embodiment will be described, omitting any overlap with that of the information processing apparatus according to the first embodiment. In the information processing apparatus 100 according to this embodiment, the storage device 70, the first acquisition unit 110, the second acquisition unit 120, the training data acquisition unit 140, and the learning unit 150 are identical to those in the first embodiment, and therefore their description will be omitted.

[0079] The generation unit 130 receives axial ground truth image data of the liver tumor, which is the first region information, from the first acquisition unit 110, and coronal ground truth image data of the liver tumor, which is the second region information, from the second acquisition unit 120. Next, the generation unit 130 determines reference points based on the axial ground truth image data and the coronal ground truth image data of the liver tumor, and generates mask image data of an ellipsoid that shows the approximate region of the liver tumor based on the equation of the ellipsoid. Subsequently, the generation unit 130 generates a likelihood map of the liver tumor by smoothing the mask image data of the ellipsoid that shows the approximate region of the liver tumor using a Gaussian filter. Then, the likelihood map of the liver tumor is stored in the storage device 70.

[0080] The hardware configuration of the information processing device 100 according to this embodiment is the same as that of the first embodiment, so a description will be omitted.

[0081] Next, using Figure 3, the processing procedure of the three-dimensional ground truth image data generation processing unit 101 in the information processing device 100 of this embodiment will be described. In the following description, parts that overlap with the description of the information processing device 100 according to the first embodiment will be omitted.

[0082] (Steps S3100~S3300) Steps S3100 to S3300 are the same as steps S3100 to S3300 in Embodiment 1, so their explanation will be omitted.

[0083] (Step S3400) In step S3400, the generation unit 130 generates a likelihood map of the correct liver tumor so as to spread in an ellipsoidal shape (likelihood changes concentrically and continuously) centering on the reference point determined in step S3300. The generation unit 130 in the present embodiment first generates mask image data of an ellipsoid indicating a rough region of the liver tumor using the equation of the ellipsoid. Then, the generation unit 130 generates a likelihood map of the correct liver tumor by smoothing the mask image data of the ellipsoid indicating a rough region of the liver tumor using a Gaussian filter or the like.

[0084] Using FIG. 4, a method for generating mask image data of an ellipsoid indicating a rough region of the liver tumor will be described. FIG. 4(d) shows mask image data 470 of an ellipsoid indicating a rough region of the liver tumor. In the present embodiment, voxels corresponding to coordinates (x, y, z) that satisfy the following conditions are regarded as voxels belonging to the ellipsoid indicating a rough region of the liver tumor.

[0085]

Equation

[0086] Here, (c x , c y , c z ) are the coordinates of the reference point 440 in the image data coordinate system determined in step S3300, and (r x , r y , r z ) represent the radii in each axial direction with respect to the reference point 440. The radii in each axial direction with respect to the reference point 440 are obtained by the same method as in the first embodiment. The mask image data 470 of the ellipsoid holds an ellipsoidal region 471 indicating a rough region of the liver tumor. As an example, it is assumed that the value of the voxels included in the ellipsoidal region 471 is 1 (light color) and the value of the other voxels is 0 (dark color). That is, in the mask image data 470 of the ellipsoid, the value of the voxels that satisfy the conditions shown in Equation 2 is 1, and the value of the other voxels is 0.

[0087] Next, the generation unit 130 applies a Gaussian filter to smooth the ellipsoidal mask image data 470, which represents the general area of ​​the liver tumor, thereby generating a ground-level likelihood map 430 of the liver tumor. In the ground-level likelihood map 430 of the liver tumor generated in this way, the voxel values ​​(likelihood of the liver tumor) are high near the reference point 440, and the voxel values ​​decrease as the distance from the reference point 440 increases.

[0088] Furthermore, the method used by the generation unit 130 to generate the likelihood map of the correct liver tumor is not limited to smoothing with a Gaussian filter; any method that generates a likelihood map such that voxels inside the contour of the ellipsoid region have a high likelihood and voxels outside have a low likelihood is acceptable. For example, the generation unit 130 generates a likelihood map indicating the likelihood of an object based on the distance from the contour of the region corresponding to the three-dimensional region image data. Specifically, the generation unit 130 may generate a likelihood map of the correct liver tumor by performing a distance transformation on the mask image data of an ellipsoid representing the general region of the liver tumor to generate a distance map based on the distance from the contour of the ellipsoid region 471, and then normalizing the voxel values ​​(distance values) of the distance map.

[0089] (Step S3500~S3600) Steps S3500 to S3600 are the same as steps S3500 to S3600 in Embodiment 1, so their explanation will be omitted.

[0090] In accordance with the above procedure, the information processing device 100 according to this embodiment generates a likelihood map of the correct liver tumor using the correct image data generation processing unit 101. Then, following the same processing procedure as the learning processing unit 102 in the first embodiment, the information processing device 100 uses the likelihood map of the correct liver tumor as correct image data and learns a learning model for inferring the likelihood map of the liver tumor using a machine learning-based method.

[0091] As described above, the information processing device 100 according to this embodiment generates a three-dimensional likelihood map (ground truth image data) of a ground truth object based on two-dimensional ground truth image data (first region information) corresponding to a first axis of the object and two-dimensional ground truth image data (second region information) corresponding to a second axis of the object. This makes it possible to efficiently generate ground truth image data that can be used for training a machine learning-based learning model. Furthermore, the information processing device 100 may perform inference processing using the trained learning model by its control unit.

[0092] In the example described above, the generation unit 130 approximated the object as an ellipsoid to generate a likelihood map of the three-dimensional object, but it is not limited to this. For example, in the case of an object with a shape close to a cylinder (such as a spine), the generation unit 130 first uses the first and second region information to generate a cylindrical mask image data that shows the approximate region of the object based on the equation of a cylinder. When generating the cylindrical mask image data, the radius and height of the cylinder are necessary, so this information can be obtained from the region information corresponding to each axis. Then, as in the example above, the generation unit 130 generates the likelihood map of the correct object by smoothing the cylindrical mask image data that shows the approximate region of the object. Similarly, if the object is close to a rectangular prism, the generation unit 130 obtains the length of each side of the rectangular prism from the first and second region information to generate a rectangular prism mask image data that shows the approximate region of the object, and then smooths this to generate the likelihood map of the correct three-dimensional object. In addition, the generation unit 130 may generate three-dimensional region image data by approximating the region of the object using a parametric shape model. In this case, the generation unit 130 obtains dependent parameters for representing the parametric shape model from the first region information and the second region information. A parametric shape model is, for example, a function of a closed surface or a statistical shape model. In these cases, the generation unit 130 may use the approximate center of the object as the reference point, or it may use characteristic positions such as the upper or lower end of the object as the reference point.

[0093] <Third Embodiment> In the first and second embodiments, a method for generating a ground-level likelihood map of a liver tumor was described by approximating the region of the liver tumor with an ellipsoid. In this embodiment, a method for generating a ground-level likelihood map of a liver tumor is described by interpolating the first region information and / or the second region information. In this embodiment, the first region information is axial ground-level image data of the liver tumor, and the second region information is two contour points of the liver tumor on representative coronal cross-sectional image data.

[0094] The information processing device according to this embodiment generates a likelihood map of the liver tumor by linearly interpolating (morphing) the axial ground truth image data of the liver tumor (first region information corresponding to the first axis) in the direction of the line segment connecting two contour points in the liver tumor (second axis). At this time, the information processing device 100 generates the likelihood map of the liver tumor such that the likelihood of the liver tumor is high for voxels that are close to the axial region of the liver tumor in the axial ground truth image data of the liver tumor, and the likelihood of the liver tumor is low for voxels that are far from the axial region of the liver tumor.

[0095] The configuration of the information processing apparatus according to this embodiment is the same as that of the information processing apparatus 100 according to the first embodiment. Hereinafter, with reference to Figure 1, the functional configuration of the information processing apparatus according to this embodiment will be described, omitting any overlap with that of the information processing apparatus according to the first embodiment. In the information processing apparatus 100 according to this embodiment, the first area information acquisition unit 110, the training data acquisition unit 140, and the learning unit 150 are identical to those in the first embodiment, and therefore their description will be omitted.

[0096] The storage device 70 stores abdominal CT image data, which is three-dimensional tomographic image data depicting the object; axial ground truth image data of the liver tumor, which is first region information; and two contour points of the liver tumor on representative coronal cross-sectional image data, which is second region information. The two contour points of the liver tumor on representative coronal cross-sectional image data are, for example, coordinate values ​​in the image coordinate system.

[0097] The second region information acquisition unit 120 acquires two contour points in the liver tumor on the representative coronal cross-sectional image data, which is the second region information, from the storage device 70 and transmits them to the generation unit 130.

[0098] The generation unit 130 receives axial ground truth image data of the liver tumor, which is the first region information, from the first region information acquisition unit 110, and receives two contour points of the liver tumor on the representative coronal cross-sectional image data, which is the second region information, from the second region information acquisition unit 120. Next, the generation unit 130 generates a ground truth likelihood map of the liver tumor by linearly interpolating the axial ground truth image data of the liver tumor in the coronal axis direction based on the two contour points of the liver tumor on the representative coronal cross-sectional image data. Then, the ground truth likelihood map of the liver tumor is stored in the storage device 70.

[0099] The hardware configuration of the information processing device 100 according to this embodiment is the same as that of the first embodiment, so a description will be omitted.

[0100] Next, the processing procedure of the ground truth image data generation processing unit 101 in the information processing device 100 in this embodiment will be described using Figure 3. In the following description, parts that overlap with the description of the information processing device 100 according to the first embodiment will be omitted.

[0101] (Step S3100) Step S3100 is the same as step S3100 in Embodiment 1, so its description is omitted.

[0102] (Step S3200) In step S3200, the second region information acquisition unit 120 acquires two contour points of the liver tumor on the representative coronal cross-sectional image data, which is the second region information, from the storage device 70 and transmits them to the generation unit 130.

[0103] (Step S3300) Step S3300 is the same as step S3300 in Embodiment 1, so its description is omitted.

[0104] (Step S3400) In step S3400, the system receives the first region information, which is the axial ground truth image data of the liver tumor, and the second region information, which is two contour points of the liver tumor on the representative coronal section image data. Based on the two contour points of the liver tumor on the representative coronal section image data, the system linearly interpolates the axial ground truth image data of the liver tumor in the coronal axis direction to generate a likelihood map of the ground truth liver tumor.

[0105] As shown in Figure 4(c), suppose that the first region information is given as axial ground truth image data 410 of a liver tumor, and the second region information is given as two contour points 460 of the liver tumor on representative coronal cross-sectional image data. In this case, the information processing device 100 generates a likelihood map of the ground truth liver tumor by linearly interpolating the axial region 411 of the liver tumor along the line segment connecting the two contour points 460, with the two contour points 460 as endpoints. Figure 4(e) is a conceptual diagram of the method for generating a likelihood map of the ground truth liver tumor by linear interpolation. For example, the generation unit 130 generates a likelihood map of the ground truth liver tumor such that the likelihood of the liver tumor is higher (lighter color) closer to the location where the axial region 411 of the liver tumor exists, and decreases (darker color) as it moves towards the axial region 412, axial region 413, and contour point 460. In other words, the generation unit 130 generates a likelihood map indicating the likelihood of the object such that voxels close to the region corresponding to the first region information and / or the second region information have a high likelihood, and voxels farther away have a low likelihood. In this way, the information processing device generates a likelihood map in three-dimensional space such that the likelihood of a liver tumor is high near the two-dimensional ground truth image data given as the first region information, and the likelihood of a liver tumor decreases as the distance from the two-dimensional ground truth image data increases. This makes it possible to efficiently generate three-dimensional ground truth image data that can be used for training a machine learning-based learning model.

[0106] (Step S3500~S3600) Steps S3500 to S3600 are the same as steps S3500 to S3600 in Embodiment 1, so their explanation will be omitted.

[0107] As described above, the information processing device 100 according to this embodiment generates a three-dimensional likelihood map (ground truth image data) of a ground truth object based on two-dimensional ground truth image data (first region information) corresponding to the first axis of the object and two-dimensional ground truth image data (second region information) corresponding to the second axis of the object. This makes it possible to efficiently generate ground truth image data that can be used for training a machine learning-based learning model. The information processing device 100 may also perform inference processing using the trained learning model.

[0108] In accordance with the above procedure, the information processing device 100 according to this embodiment generates a likelihood map of the correct liver tumor using the correct image data generation processing unit 101. Then, following the same processing procedure as the learning processing unit 102 in the first embodiment, the information processing device 100 uses the likelihood map of the correct liver tumor as correct image data and learns a learning model for inferring the likelihood map of the liver tumor using a machine learning-based method.

[0109] (modified version) In the example above, linear interpolation was used to generate the likelihood map of the correct liver tumor, but the interpolation method is not limited to linear interpolation. For example, suppose that the first region information is axial ground truth image data of the liver tumor, and the second region information is coronal ground truth image data of the liver tumor. In this case, the region of the liver tumor held in each two-dimensional ground truth image data can be rotated around an axis shared by the two cross-sections (a type of interpolation method) to generate a solid of revolution, and the likelihood map of the correct liver tumor can be generated by integrating the two solids of revolution. The generation of the solid of revolution can be based on any known method, but for example, the upper half of the solid of revolution can be generated using the contour above the axis, and the lower half of the solid of revolution can be generated using the contour below the axis. Furthermore, the solids of revolution can be integrated using the weighted mean, maximum, or minimum value for each voxel. Furthermore, when using weighted averaging, it is desirable to set the weight of the rotational body generated from the first region information to be larger for angles closer to the first representative two-dimensional tomographic image data, and to set the weight of the rotational body generated from the second region information to be larger as the angle approaches the second representative two-dimensional tomographic image data. According to the above, it is possible to generate three-dimensional region information that approximates the region information of the two cross-sections as closely as possible. [Explanation of Symbols]

[0110] 70 Storage device 101 Processing unit for generating correct image data 102 Learning Processing Unit 110 First acquisition section 120 Second acquisition section 130 Generation part 140 Training Data Acquisition Unit 150 Learning Department

Claims

1. A first acquisition unit that acquires first regional information of an object on a first two-dimensional tomographic image data that is orthogonal to the first axis of the three-dimensional tomographic image data, A second acquisition unit that acquires second region information of the object on a second two-dimensional tomographic image data which is orthogonal to a second axis different from the first axis of the three-dimensional tomographic image data, A generation unit generates a three-dimensional likelihood map indicating the object's resemblance as three-dimensional ground truth image data corresponding to the three-dimensional tomographic image data, based on the first region information and the second region information. It has, The information processing device is characterized in that the generation unit determines a reference point based on at least one of the first region information and the second region information, and generates the likelihood map such that the likelihood is high in voxels close to the reference point and low in voxels far from the reference point.

2. A first acquisition unit that acquires first regional information of an object on a first two-dimensional tomographic image data that is orthogonal to the first axis of the three-dimensional tomographic image data, A second acquisition unit that acquires second region information of the object on a second two-dimensional tomographic image data which is orthogonal to a second axis different from the first axis of the three-dimensional tomographic image data, A generation unit generates a three-dimensional likelihood map indicating the object's resemblance as three-dimensional ground truth image data corresponding to the three-dimensional tomographic image data, based on the first region information and the second region information. It has, At least one of the first region information and the second region information is two-dimensional ground truth image data of the object corresponding to each axis, The generation unit is characterized by generating the likelihood map based on a probability distribution function.

3. A first acquisition unit that acquires first regional information of an object on a first two-dimensional tomographic image data that is orthogonal to the first axis of the three-dimensional tomographic image data, A second acquisition unit that acquires second region information of the object on a second two-dimensional tomographic image data which is orthogonal to a second axis different from the first axis of the three-dimensional tomographic image data, A generation unit generates a three-dimensional likelihood map indicating the object's resemblance as three-dimensional ground truth image data corresponding to the three-dimensional tomographic image data, based on the first region information and the second region information. It has, The information processing device is characterized in that the generation unit generates three-dimensional region image data based on the first region information and the second region information, and generates a likelihood map such that the likelihood is high for voxels located inside the contour of the three-dimensional region image data and low for voxels located outside the contour.

4. The information processing apparatus according to claim 3, characterized in that the generation unit generates the three-dimensional region image data based on a parametric shape model.

5. The information processing apparatus according to claim 4, characterized in that the generation unit generates the likelihood map by smoothing the three-dimensional region image data.

6. The information processing apparatus according to claim 4, characterized in that the generation unit generates the likelihood map based on the distance from the contour of the region corresponding to the three-dimensional region image data.

7. A first acquisition unit that acquires first regional information of an object on a first two-dimensional tomographic image data that is orthogonal to the first axis of the three-dimensional tomographic image data, A second acquisition unit that acquires second region information of the object on a second two-dimensional tomographic image data which is orthogonal to a second axis different from the first axis of the three-dimensional tomographic image data, The system includes a generation unit that generates a three-dimensional likelihood map indicating the object's resemblance as three-dimensional ground truth image data corresponding to the three-dimensional tomographic image data, based on the first region information and the second region information, The information processing device is characterized in that the generation unit generates the likelihood map such that the likelihood is high for voxels that are close to the region corresponding to at least one of the first region information and the second region information, and low for voxels that are far away.

Citation Information

Patent Citations

  • Image processing apparatus, computer program, recording medium, and image processing method

    JP2014035597A

  • Image segmentation using neural network methods

    JP2019526863A